A method and system for optimizing allocation of neural network storage and computing resources based on FPGA
By establishing a resource relationship model between layers and within layers on FPGA, and combining the optimized allocation method, the rationality problem of FPGA in the allocation of resources at each layer of neural network is solved, and efficient calculation and low-latency application of convolutional neural networks in edge scenarios is realized.
Patent Information
- Application Number
- CN202510018312.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The existing technology is difficult to effectively solve the rationality problem of FPGA in the resource allocation of each layer of neural network, resulting in the excessive delay and resource utilization of convolutional neural networks in edge scenarios, limiting their application.
A method of optimization allocation of neural network memory resources based on FPGA is proposed. By establishing an inter-layer computing resource relationship model and an in-layer memory resource relationship model, combining the computing resource allocation method and memory resource and delay balance optimization method, the parallelism of each layer is reasonably allocated to achieve efficient resource allocation and delay optimization.
It realizes the reduction of data flow blocking under the premise of minimum resource utilization, ensures that the module delays within each layer and the delays between each layer are consistent, thereby achieving the effect of minimum total delay, and improving the computing speed and application performance of convolutional neural networks in edge scenarios.
Smart Images

Figure CN119440853B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method and system for optimizing allocation of neural network storage and computing resources based on FPGA. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Convolutional neural networks are being applied to all walks of life. Most algorithms such as image classification, target detection and tracking, image segmentation, and posture estimation require convolutional neural networks as feature extractors. The characteristics of convolutional neural networks enable them to obtain deeper features, thereby making the algorithm more accurate. However, convolutional neural networks also bring high computing power and storage requirements, which in turn brings higher energy consumption and resource usage.
[0004] In edge scenarios, the computing power, storage, and energy consumption of devices are all limited, but at the same time, edge scenarios are also very sensitive to latency. Therefore, due to the high computing power and storage requirements of convolutional neural networks, the application of convolutional neural networks in edge scenarios is greatly restricted.
[0005] In order to optimize the performance of convolutional neural networks in edge scenarios, the industry widely uses FPGA (Field Programmable Gate Array) to accelerate neural networks, and uses a series of optimization technologies to reduce latency and resource usage and improve throughput.
[0006] Existing research on using FPGA to accelerate neural networks mainly focuses on the following aspects: network lightweighting, computing architecture optimization, improving computing parallelism, batch processing, etc. However, FPGA resources are limited, and the parallelism of each layer is highly correlated with the overall delay and resource usage. Therefore, whether the resources allocated to each layer of the neural network are reasonable will have a great impact on the delay and total resource usage of the final accelerator.
[0007] However, there is little research on how to allocate computing resources and storage resources at each layer, and it is impossible to solve the problem of whether the resources allocated by FPGA to each layer of the neural network are reasonable. Summary of the invention
[0008] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for optimizing the allocation of storage and computing resources of a neural network based on FPGA, and proposes an inter-layer resource efficient allocation module including an inter-layer computing resource relationship model and a computing resource allocation method, and an intra-layer memory resource and delay balance optimization module including an intra-layer memory resource relationship model and a memory resource and delay balance optimization method. By reasonably allocating the parallelism of each layer, reducing data flow blockage under the premise of minimum resource occupation, and aiming at consistent delays of modules within each layer and consistent delays between layers, the effect of minimizing the total delay is finally achieved, the computing speed of the convolutional neural network in edge scenarios is improved, the application performance of the convolutional neural network in edge scenarios is optimized, and the performance of the FPGA accelerator in edge scenarios is greatly improved.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] The first aspect of the present invention provides a method for optimizing the allocation of neural network storage and computing resources based on FPGA.
[0011] A method for optimizing allocation of storage and computing resources of a neural network based on FPGA, comprising:
[0012] Analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers;
[0013] Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method;
[0014] Analyze the latency, parallelism, and memory resource consumption of each module within the neural network layer, and establish a memory resource relationship model within the layer;
[0015] Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factors of each layer of the neural network after allocation, the balance optimization of intra-layer memory resources and delay is achieved.
[0016] The second aspect of the present invention provides a neural network storage and computing resource optimization allocation system based on FPGA.
[0017] A neural network storage and computing resource optimization allocation system based on FPGA, comprising:
[0018] The computing resource allocation module is configured to: analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers;
[0019] Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method;
[0020] The memory resource and delay balance optimization module is configured to: analyze the delay and parallelism of each module in the neural network layer, and the memory resource consumption, and establish a memory resource relationship model within the layer;
[0021] Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factors of each layer of the neural network after allocation, the balance optimization of intra-layer memory resources and delay is achieved.
[0022] The third aspect of the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a method as described in the first aspect of the present invention are implemented.
[0023] A fourth aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in a method as described in the first aspect of the present invention.
[0024] A fifth aspect of the present invention provides a computer program product comprising instructions, which, when running on a computer, enables the computer program to implement the steps in a method as described in the first aspect of the present invention when executed by a processor.
[0025] One or more of the above technical solutions have the following beneficial effects:
[0026] The present invention proposes an inter-layer resource efficient allocation module including an inter-layer computing resource relationship model and a computing resource allocation method, and an intra-layer memory resource and delay balance optimization module including an intra-layer memory resource relationship model and an intra-layer memory resource and delay balance optimization method. By reasonably allocating the parallelism of each layer, data flow blockage is reduced under the premise of minimum resource occupancy, and the consistency of module delays within each layer and the consistency of delays between layers are taken as the goals, ultimately achieving the effect of minimizing the total delay.
[0027] The inter-layer computing resource efficient allocation module proposed in the present invention can take the delay consistency between each layer of the neural network as the optimization direction, and combine the resource conditions of the target platform to be deployed to reasonably allocate the parallel factors of each layer of the neural network. and , achieving the ultimate goal of maximizing computing resource utilization and minimizing overall latency. By improving the efficient allocation of computing resources, the utilization of computing resources in FPGA can be greatly improved, the idle time of computing resources can be reduced, and tasks can be completed in a faster time and with lower energy consumption.
[0028] The intra-layer memory resource and delay balance optimization module proposed in the present invention mainly focuses on the balance optimization of BRAM resources and delays used as row buffers in the dual sliding window module. The model takes the consistency of delays of each module in the layer as the optimization direction, and can minimize the number of BRAM ports required while avoiding idle computing units, and ultimately minimize BRAM resource consumption. By balancing memory resources and delays, the consumption of on-chip memory resources can be significantly reduced, allowing all parameters to be placed in on-chip memory, thereby effectively reducing unnecessary off-chip memory accesses. Therefore, through efficient optimization of computing resources and memory resources, power consumption can be significantly reduced, battery life can be extended, or dependence on external power supplies can be reduced.
[0029] The accelerator optimized by the resource allocation method proposed in the present invention can complete the reasoning process of the neural network model in a shorter time, thereby providing near real-time reasoning response. It is more suitable for application scenarios that are sensitive to delay, such as intelligent video surveillance and autonomous driving, and greatly improves the performance of FPGA accelerators in edge scenarios.
[0030] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0032] Figure 1 This is a flow chart of a method for optimizing allocation of storage and computing resources of a neural network based on FPGA in Embodiment 1 of the present invention;
[0033] Figure 2 This is a flow chart of a method for allocating computing resources between layers in Embodiment 1 of the present invention;
[0034] Figure 3 This is a flow chart of a method for optimizing memory resources and delay balance in a middle layer in Embodiment 1 of the present invention;
[0035] Figure 4 The diagram is a comparison of the effects of the inter-layer computing resource allocation method in the first embodiment of the present invention; wherein (a) shows the comparison of the delay before and after the AllocSCR optimization; (b) shows the comparison of the DSP consumption before and after the AllocSCR optimization;
[0036] Figure 51 is a diagram comparing the effects of the in-layer memory resources and delay balance optimization method in the first embodiment of the present invention; wherein (a) shows the comparison of the delay ratios of MVM and DWM before and after BalanceBL optimization; (b) shows the comparison of the BRAM consumption of DWM before and after BalanceBL optimization;
[0037] Figure 6 This is a timeline diagram of the dual sliding window module and matrix-vector multiplication module used for convolution calculation before and after balanced optimization in Example 1 of the present invention. DETAILED DESCRIPTION
[0038] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0039] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0040] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0041] Terminology explanation:
[0042] 1. FPGA: Field Programmable Gate Array, field programmable gate array;
[0043] 2. BRAM: Block Random Memory, block random memory;
[0044] 3. SWM: Sliding Window Module, which is used to output the input feature map to MVM in the order required by the convolution calculation;
[0045] 4. MVM: Matrix Vector multiplication Module, matrix vector multiplication module, responsible for convolution calculation and BnReLU (batch normalization and activation) operation;
[0046] 5. Row buffer: used to cache three rows of data in the double sliding window module, in preparation for subsequent output in the order of data required by convolution;
[0047] 6. Latency: delay, the time from the start of the module to the end of the module execution;
[0048] 7. DSP: Digital Signal Processor, refers to the computing resources in FPGA - digital processing unit;
[0049] 8. BnReLU: Batch normalization and ReLU activation, which are commonly used operations in convolutional neural networks.
[0050] Embodiment 1
[0051] This embodiment discloses a method for optimizing the allocation of storage and computing resources of a neural network based on FPGA, and proposes an inter-layer resource efficient allocation module including an inter-layer computing resource relationship model and a computing resource allocation method, and an intra-layer memory resource and delay balance optimization module including an intra-layer memory resource relationship model and a memory resource and delay balance optimization method. Figure 1 As shown, a method for optimizing allocation of storage and computing resources of a neural network based on FPGA includes:
[0052] Step 1: Analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers;
[0053] Step 2: Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method;
[0054] Step 3: Analyze the delay, parallelism, and memory resource consumption of each module in the neural network layer, and establish a memory resource relationship model within the layer;
[0055] Step 4: Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factor of each layer of the neural network after allocation, the balance optimization of the intra-layer memory resources and latency is achieved.
[0056] In order to more clearly illustrate this embodiment, a method for optimizing the allocation of storage and computing resources of a neural network based on FPGA is provided. Figure 1 As shown, the implementation process can be described in detail as follows:
[0057] Step 1: Analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers.
[0058] In this embodiment, first, delay and computing resource modeling is performed, that is, two parallel factors are established: and The relationship model between FPGA computing resources and each layer delay.
[0059] For the parallel factor and The specific instructions are as follows:
[0060] Assume that the shape of the output feature map of the i-th convolutional layer is (1920, 1080, 3), corresponding to ( ), the number of channels for each pixel is 3, there are a total of 1920*1080 pixels, and the delay required to output a complete single pixel is .
[0061] In order to obtain The output data of the channel needs clock cycle, the specific formula is as follows:
[0062] ;
[0063] in, Represents the size of the convolution kernel of the i-th layer; and Respectively represent the number of input and output channels of the i-th convolutional layer; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously. Refers to the delay required to output a single channel of a single pixel, which is then multiplied by Refers to the output of a single pixel The delay required for each channel.
[0064] Based on the above parallel factors, the specific steps to establish the inter-layer computing resource relationship model are as follows:
[0065] Step 1-1: Establish a Dual Sliding Windows Module (DWM) delay model:
[0066] The operation logic of DWM: First, the initialization loading phase, it receives The first layer receives the input data and stores it in the buffer. Then it enters the operation phase and receives the input data from the previous layer. Line input data, overwriting the buffer that is no longer needed The data at this position is no longer needed for the next convolution calculation, and then the sliding output is performed in the order required for the convolution calculation. This process will be repeated. Second-rate.
[0067] According to the above operation logic, the delay equation of DWM is established as follows:
[0068] ;
[0069] ;
[0070] ;
[0071] in, Represents the size of the convolution kernel of the i-th layer; Represents the step size of the i-th layer convolution; , Respectively represent the output width and height of the i-th convolutional layer; , , They represent the input width, number of channels, and number of paddings of the i-th convolutional layer respectively; Indicates the number of inputs packed into a single DSP.
[0072] Represents the delay of DWM receiving a new line of data. represents the parallelism of the input channels, Represents the start interval of the loop that receives each data when receiving a new row of data.
[0073] Represents each The delay of row data output to MVM, Represents the start interval of the output cycle.
[0074] ;
[0075] ;
[0076] in, Represents the input height of the i-th convolutional layer.
[0077] In this step, by establishing a double sliding window module delay model, the delay of the double sliding window module can be calculated.
[0078] Step 1-2: Establish the Matrix Vector multiplication Module (MVM) delay model:
[0079] ;
[0080] ;
[0081] in, Represents the number of multiply-accumulate operations (MACs) of the i-th convolution layer, Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously. Represents the number of inputs packed into a single DSP, which is the computational resource in the FPGA occupied by the convolution calculation.
[0082] , , Respectively represent the output width, height, and number of channels of the i-th convolutional layer; Represents the size of the convolution kernel of the i-th layer; Represents the number of input channels of the i-th convolutional layer.
[0083] In this step, by establishing a matrix-vector multiplication module delay model, the delay of the matrix-vector multiplication module can be calculated.
[0084] Step 1-3: Establish DSP resource model based on MACs:
[0085] In this embodiment, the DSP consumed by the calculation is divided into two types: the DSP consumed by the MACs in Conv (convolution) and the DSP consumed by the MACs in Conv (convolution). and MACs consumed by BnReLU (batch normalization and ReLU activation) , the specific expression is as follows:
[0086] ;
[0087] Because the packed weights come from different convolution kernels. At the same time, PE represents the number of convolution kernels participating in the calculation at the same time. So PE needs to be divided by .
[0088] Establish The equation is as follows:
[0089] ;
[0090] in, Indicates the number of weights packed into a DSP; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously.
[0091] After each convolution, a BnReLU operation is required. Therefore, after completing the calculation of PE convolution kernels in parallel, PE BnReLU calculation operations must also be completed in parallel.
[0092] Establish The equation is as follows:
[0093] ;
[0094] The DSP used for BnReLU calculation here cannot improve the throughput of the accelerator. Therefore, in subsequent resource allocation, the priority of SIMD is increased and the priority of PE is decreased to use all DSPs for improving throughput as much as possible (corresponding to when sorting the list in the parallel factor product separation algorithm before step 2-1 in the following, when the products of two parallel factors are the same, the integer factor pairs list is sorted in descending order according to the size in the integer factor pair).
[0095] In this step, based on the multiply-accumulate operation numbers (MACs), a resource model of the FPGA resources occupied by convolutional calculation is established, and the consumption of DSP resources can be calculated through the parallelism.
[0096] By establishing the double sliding window module delay model, the matrix-vector multiplication module delay model, and the DSP resource model based on MACs, it is to accurately calculate the delay or resource occupancy required by the corresponding module, facilitate the subsequent accurate allocation of computing resources, and improve the optimization effect of resources.
[0097] Step 2: Based on the inter-layer computing resource relationship model, through the computing resource allocation method, obtain the optimal parallel factors of each layer of the neural network after allocation.
[0098] In this embodiment, as Figure 2 shown, in order to more reasonably allocate the computing resources of each layer, a novel inter-layer computing resource allocation method, Allocation-SIMDPE-Computing-Resource (AllocSCR), is proposed.
[0099] Since the delay of CNN mainly comes from the multiply-accumulate (MAC) operation in the convolutional layer, in the data flow architecture adopted in this embodiment, the convolutional calculation in the convolutional layer is mapped to MVM. Therefore, in this embodiment, the resources are mainly allocated according to the delay of the matrix-vector multiplication module in step 1 ( ), and then the delay of the double sliding window module in step 1 ( ) is used for auxiliary comparison to adjust the allocated amount of computing resources for each layer.
[0100] Among them, the delay of MVM is mainly affected by the following factors:
[0101] a) The MACs of the current layer.
[0102] b) The product of the current layer : .
[0103] c) The number of inputs of the current layer that are packed into a single DSP is expressed as: .
[0104] In this embodiment, the final allocation of SIMD and PE is determined through three steps:
[0105] Step 2-1: Initial allocation.
[0106] Step 2-2, first fine-tuning: increase the parallel factor of the layer with latency greater than the average ( ).
[0107] Step 2-3, second fine-tuning: reduce the parallel factor of the layer with latency less than the maximum value ( ).
[0108] As shown in Table 1, the proposed inter-layer computing resource allocation method is introduced in detail in Algorithm 1.
[0109] Table 1 Algorithm examples of inter-layer computing resource allocation methods
[0110]
[0111] Before making specific allocations, a parallel factor product separation algorithm is designed, that is, the best parallel factor pair is obtained according to the product of two given parallel factors and the scaling factor. As shown in Table 2, the proposed parallel factor product separation algorithm is introduced in detail in Algorithm 2.
[0112] Table 2 Example of parallel factor product separation algorithm
[0113]
[0114] Calculate from 1 to scale times later All integer factor pairs of , ), form a list of integer factor pairs; then sort the list in descending order according to the product of the integer factor pairs. When the products are the same, sort them according to the integer factor pairs. Sort the list of integer factor pairs in descending order, and then traverse the entire list, and take the factor pair in the list that satisfies the current layer constraints and has the smallest element subscript as the best parallel factor pair.
[0115] Combining Table 1 and Table 2, the detailed allocation is as follows:
[0116] Step 2-1: Perform initial allocation based on the inter-layer computing resource relationship model:
[0117] (1) Since the latency of each layer is proportional to MAC and inversely proportional to the product of SIMD and PE, we first need to determine the latency of each layer based on the proportion of MAC in each layer to the total MAC of CNN. Proportion.
[0118] (2) The layers determined according to (1) Proportion, after allocating each layer After that, you need to determine and The value of .
[0119] Considering that there may not be two integers in the current layer and The constraints can be satisfied, so a list of integer factor pairs needs to be constructed , The elements in are scale factors , by assigning a scaling factor to each layer Come to Scale it down.
[0120] in, The initial value of is determined empirically. In this embodiment, it is initialized to 1.2.
[0121] (3) Traverse each layer in turn and obtain the best parallel factor pair as the two parallel factors of the current layer according to the parallel factor product separation algorithm and .
[0122] Step 2-2: Increase the parallel factor of the layer with latency greater than the average and perform the first fine-tuning:
[0123] Since the initial allocation is performed under ideal conditions, it is assumed that We can find exactly the factor pairs that meet the requirements. and , to achieve consistency in latency across all layers. However, after amplification It may still be impossible to find a factor pair that satisfies the constraints, resulting in excessive latency in some layers. Therefore, we need to fine-tune the parallel factors of layers whose latency is greater than the average:
[0124] (1) To avoid the impact of extreme data, we first remove the layers with too small latency (the specific selection of the layers to be removed is based on empirical values) and calculate the overall latency average.
[0125] (2) Traverse each layer. If the latency of a layer is greater than the average latency, the layer is The scaling factor Increase, and then repeat the parallel factor product separation algorithm until the latency of this layer is less than or equal to the average latency, or no SIMD and PE combination that satisfies the constraints can be found.
[0126] Step 2-3: Reduce the parallel factor of the layer whose delay is less than the maximum value and perform a second fine-tuning:
[0127] In this embodiment, it is also necessary to tighten the computing resources of each layer according to the maximum delay.
[0128] When the following situation occurs: the delay of a certain layer is already at the minimum value allowed by the constraints, but it is still large compared to other layers. Due to the characteristics of the pipeline, the delay of the entire accelerator is limited by the layer with the largest delay. Therefore, no matter how small the delay of other layers is, it will no longer affect the overall delay of the accelerator, which will lead to a waste of computing resources in other layers. Therefore, we need to tighten the computing resources of other layers based on the maximum delay.
[0129] At the same time, for each layer, the following may happen: the input feature map is too large, resulting in too large a delay for DWM, while the number of channels in this layer is shallow, so the MVM delay is smaller than that of DWM. In this case, the amount of computing resources allocated to this layer also needs to be reduced.
[0130] The specific approach to tightening computing resources is to reduce the amount of computing resources allocated to this layer. First, find the layer with the largest delay, then traverse each layer in turn and gradually reduce , recalculate the corresponding delay until the delay of this layer is less than or equal to the maximum value.
[0131] The resource allocation method established through the above steps reasonably sets the parallel factors of each layer and , thereby maximizing resource utilization and minimizing overall latency.
[0132] Step 3: Analyze the delay, parallelism, and memory resource consumption of each module in the neural network layer, and establish a memory resource relationship model within the layer.
[0133] This embodiment takes a DWM (Dual Sliding Window Module) using a BRAM (Block Random Access Memory) as a buffer as an example. Since the number of BRAM ports is limited, each BRAM has at most two ports for reading data. When:
[0134] ;
[0135] That is, when the number of ports required is greater than 2, in order to increase the number of ports, the common practice is to use ArrayPartition, Array Reshape, and make copies. Here, this embodiment takes making copies as an example to model the memory resources within the layer. We will use it in the cache sliding window module Row input data Copy to In the copy.
[0136] in, Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the number of input channels of the i-th convolutional layer; Indicates the number of buffer copies required according to the parallelism.
[0137] In order to accurately balance and optimize BRAM resources and latency in the future, we establish an intra-layer memory (BRAM) resource relationship model based on the module latency and memory resources within the neural network layer.
[0138] Specifically, the steps to establish the intra-layer memory resource relationship model are as follows:
[0139] Step 3-1: Model the buffer BRAM consumption in DWM.
[0140] The calculation formula for the BRAM resources consumed by the DWM buffer is as follows:
[0141] ;
[0142] in, Represents the bit width of the DWM buffer. Represents the depth of the DWM buffer. Indicates the number of buffer copies required according to the parallelism. and is the configuration of BRAM, is the width of the BRAM, is the depth of the BRAM, and the valid value range is shown in the following formula:
[0143] ;
[0144] ;
[0145] in, Represents the bit width of the BRAM18K in simple dual-port (SDP) configuration mode. Represents the depth of the BRAM18K in simple dual-port (SDP) configuration mode. Represents the bit width of BRAM18K in true dual-port (TDP) configuration mode. Represents the depth of BRAM18K in truedual-port (TDP) configuration mode.
[0146] Step 3-2: Model the bit width, depth, and number of buffer copies of the DWM buffer.
[0147] The specific calculation formula is as follows:
[0148] ;
[0149] ;
[0150] ;
[0151] in, Indicates the depth of the DWM buffer; Indicates the bit width of the DWM buffer; The value depends on the number of ports required ; Indicates the configuration mode of BRAM. When , it means that the BRAM is configured in SDP mode. Indicates that the BRAM is configured in TDP mode.
[0152] Step 3-3: Model the relationship between the amount of data read in parallel from the DWM buffer in each loop and the degree of parallelism.
[0153] ;
[0154] in, Indicates the amount of data that needs to be read from the buffer in each loop. The read data can be used by MVM to complete a single parallel calculation; Represents the size of the convolution kernel of the i-th layer; Represents the overall width of two adjacent convolution windows. The specific calculation formula is as follows:
[0155] ;
[0156] in, Represents the output of two adjacent convolution windows. The number of overlapping parts of the rows is subtracted (for example, for two adjacent convolution windows with a kernel size of 3 and a stride of 1, when When the value is 5, we need to take 5 numbers from the first window and 5 numbers from the second window. Then there will be overlap between these two numbers. The number of overlaps between the two rows is 1. ), the specific calculation formula is as follows:
[0157] ;
[0158] ;
[0159] (Single convolution Window out) indicates the number of data required for each parallel calculation of MVM for a single convolution window. Represents the number of input channels of the i-th convolutional layer; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution.
[0160] Step 3-4: Model the latency and parallelism of the operations required for each module in the layer to complete two convolutions in parallel.
[0161] ;
[0162] ;
[0163] in, Indicates the start interval of the output data pipeline of the dual sliding window module; Represents the size of the convolution kernel of the i-th layer; and Respectively represent the number of input and output channels of the i-th convolutional layer; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously.
[0164] Step 3-5: Establish a relationship model between BRAM resource consumption, parallelism, and latency.
[0165] ;
[0166] in, The total number of ports for the DWM buffer, through Corresponding BRAM resource consumption; is the amount of data read in parallel from the DWM buffer according to the degree of parallelism, corresponding to the degree of parallelism ; The start interval of the DWM output data pipeline, corresponding to the delay of DWM .
[0167] Through the relationship model established above, we can accurately calculate the amount of data that needs to be read from the DWM buffer and the BRAM resources required according to the degree of parallelism. , , To accurately balance and optimize BRAM resources and latency.
[0168] Step 4: Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factor of each layer of the neural network after allocation, the balanced optimization of intra-layer memory resources and delay is achieved.
[0169] Ideally, The number of buffers should be such that they can be output within one clock cycle. data, that is, to ensure that the buffer meets the peak bandwidth requirements, but this method will occupy too many BRAM resources. Since the number of BRAM is limited, this embodiment proposes a method for optimizing the balance of memory resources and latency within the layer, Balance BRAM and Latency. (BalanceBL), by determining the maximum startup interval of the DWM buffer output data pipeline , and then get the minimum number of ports required , and then determine , . This ensures timely data delivery while minimizing BRAM consumption, ultimately optimizing the balance between memory resources and latency within the layer.
[0170] By analyzing the behavior pattern of convolution calculations under the pipeline:
[0171] Figure 6 Indicates that in the optimization of memory resources and latency balance within the layer, ( ) = (1, 6, 6) and ( ) = (4, 3, 3) weights for convolution, and the parallelism is configured as For example, here is a timeline diagram of memory read operations (R0, R1, …) and computation operations (C0, C1, …) in the original case.
[0172] like Figure 6 As shown, R0 represents the number of times DWM reads from the Buffer during the first convolution. The operation of data, C0 represents the parallel calculation of MVM each time during the first convolution The operation of data. R1 and C1 represent the second convolution, and C2 and R2 are the same. Ideally, each DWM loop needs to read =4 data (by formula derive), which requires consumption = 4 ports. However, we found that R0 is idle in the three clock cycles after executing three times, that is, the utilization rate of the BRAM port in the time dimension is only 1 / 2. So at this time, the BRAM port is not fully utilized. This is because in a single convolution calculation, DWM only needs to output data, and MVM needs to complete The delay of DWM in a single convolution calculation is , the delay of MVM is ,when < When the delay difference between the two will always exist, which will lead to the waste of BRAM ports in the time dimension. Therefore, we need to balance the memory resources and delay within the layer.
[0173] Specifically, Figure 3 As shown in the figure, by comparing the delay of completing a convolution calculation of the dual sliding window module and the matrix-vector multiplication module, the startup interval of the dual sliding window module output pipeline is determined, and then the minimum number of ports of the buffer in the dual sliding window module is determined according to the startup interval and the amount of parallel read data, so as to minimize the BRAM resource consumption of the buffer while ensuring that the delay of the overall pipeline remains unchanged, that is, the balance optimization of memory resources and delay is achieved. The specific steps are as follows:
[0174] Step 4-1: Calculate the maximum startup interval of the DWM (Dual Sliding Window Module) output without affecting the total pipeline delay .
[0175] To ensure timely data delivery, we need to ensure the latency of a single convolution window output to the computational unit. Less than the computational delay of the single convolution window , that is, to ensure that:
[0176] ;
[0177] in, The specific calculation formula has been given in step 3.
[0178] That is to ensure that: ;
[0179] Then we can determine The maximum value of , the maximum value is as follows:
[0180] ;
[0181] in, Represents the number of output channels of the i-th convolutional layer; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously.
[0182] Step 4-2: According to the maximum value of the startup interval Calculate the number of DWM buffer copies that need to be made .
[0183] and The calculation method has been given in step 3. Substitute the maximum value Right now:
[0184] ;
[0185] Since each BRAM has at most two ports, the number of ports in the buffer needs to be increased. For the convenience of analysis, this embodiment increases the number of ports by duplicating the buffer. The number of DWM buffer copies required is As shown below:
[0186] C ;
[0187] Step 4-3: According to the number of DWM buffer copies Calculate the BRAM consumption under different BRAM configurations and output the BRAM configuration with the minimum BRAM consumption (the configuration content includes: , , ).
[0188] The specific calculation method has been modeled in step 3, as shown below:
[0189] ;
[0190] Step 4-4: According to step 4-1 4-2 The number of DWM buffer copies obtained With the BRAM configuration obtained in 4-3, we can correctly configure the BRAM of DWM to minimize BRAM consumption while ensuring timely data provision, thus achieving a balanced optimization of memory resources and latency.
[0191] By building an efficient resource allocation model between layers and an optimization model for balancing memory resources and latency within a layer, the optimal parallel factors of each layer of the neural network after allocation can be used to optimize the balance between memory resources and latency within the layer, and the following beneficial effects can be achieved:
[0192] 1. Reduce energy consumption:
[0193] By improving the efficient allocation of computing resources, the utilization rate of computing resources in FPGA can be greatly improved, the idle time of computing resources can be reduced, and tasks can be completed in a faster time and with lower energy consumption; by balancing memory resources and delays, the consumption of on-chip memory resources can be significantly reduced, allowing all parameters to be placed in on-chip memory, thereby effectively reducing unnecessary off-chip memory access. Therefore, through efficient computing resources and memory resource optimization, power consumption can be significantly reduced, battery life can be extended, or dependence on external power can be reduced.
[0194] 2. Improve computing performance:
[0195] The accelerator optimized by all the methods of this embodiment can complete the reasoning process of the neural network model in a shorter time, thereby providing near real-time reasoning response, and is more suitable for application scenarios that are sensitive to delays, such as intelligent video surveillance and autonomous driving.
[0196] Through the optimization method of this embodiment, the experimental results are obtained: on the FPGA development board named PYNQ-Z2, the FPS is increased to 642 frames at a power consumption of 5.05W, and the computing power can reach 269.09GOPS. This greatly improves the performance of FPGA accelerators in edge scenarios.
[0197] In this embodiment, experiments have proved that by introducing AllocSCR, corresponding computing resource allocation according to the computing power required by each layer is achieved.
[0198] After the introduction of BalanceBL, the delay of DWM is determined according to the bandwidth of the data required by MVM, and then the configuration of the buffer in DWM is determined, thereby reducing the waste of BRAM and achieving timely data supply with the minimum number of BRAMs. Ultimately, the balance optimization of BRAM resources and delay is achieved.
[0199] The optimization effects of AllocSCR and BalanceBL are specifically shown as follows:
[0200] 1. AllocSCR optimization effect:
[0201] like Figure 4 As shown in the figure, compared with the original situation, by reducing the SIMD × PE of the sixth layer, this layer achieves the same latency as other layers. At the same time, for the specific SIMD and PE of each layer, under the premise of keeping the allocated SIMD × PE basically unchanged (that is, the computing resources consumed by this layer are basically unchanged), by increasing SIMD and reducing PE, the number of BnReLU operators required is reduced, thereby reducing the number of DSPs required for this layer.
[0202] From the DSP line of AllocSCR in the figure, we can see that the total consumption of DSP has decreased by 21.7% compared to when AllocSCR was not introduced. By optimizing the configuration of SIMD and PE, AllocSCR also optimizes the data supply module DWM of each layer. As can be seen in the figure, the DWM delay of conv0 with AllocSCR introduced is reduced by 72.73% compared to when AllocSCR was not introduced, breaking the bottleneck of data supply and ensuring timely data supply.
[0203] 2. BalanceBL optimization effect:
[0204] like Figure 5 As shown in the figure, after the introduction of BalanceBL, the delay of DWM is determined according to the bandwidth of the data required by MVM, and then the configuration of the buffer in DWM is determined, thereby reducing the waste of BRAM and achieving timely data supply with the minimum number of BRAM. Figure 5 The two bar graphs on the left show that compared with the case without the introduction of BalanceBL, the ratio of the data output delay of MVM to DWM has decreased significantly. The ratio of the first layer has decreased by 83.3%, and all subsequent layers have decreased by 50%, which significantly balances the delay of DWM and MVM. Figure 5 In the area chart on the middle right, DWM BRAM consumption is reduced by 61.25%, achieving optimized allocation of DWM BRAM resources.
[0205] Here, the variables in this embodiment are summarized and described:
[0206] Represents the size of the convolution kernel of the i-th layer; and Respectively represent the number of input and output channels of the i-th convolutional layer; , , , Respectively represent the input width, height, number of channels, and number of paddings of the i-th convolutional layer; , , Respectively represent the output width, height, and number of channels of the i-th convolutional layer; and are the two parallel factors of the ith layer of CNN; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously.
[0207] Embodiment 2
[0208] The purpose of this embodiment is to provide a neural network storage and computing resource optimization allocation system based on FPGA, including:
[0209] The computing resource allocation module is configured to: analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers;
[0210] Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method;
[0211] The memory resource and delay balance optimization module is configured to: analyze the delay and parallelism of each module in the neural network layer, and the memory resource consumption, and establish a memory resource relationship model within the layer.
[0212] Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the parallel factors of each layer of the neural network after allocation, the balance optimization of the intra-layer memory resources and delay is achieved.
[0213] Embodiment 3
[0214] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0215] Embodiment 4
[0216] The purpose of this embodiment is to provide a computer-readable storage medium.
[0217] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.
[0218] Embodiment 5
[0219] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments.
[0220] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0221] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0222] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A method for optimizing the allocation of storage and computing resources of a neural network based on FPGA, characterized in that: include: Analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers; Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method. The specific steps are: Perform initial allocation based on the inter-layer computing resource relationship model: determine the ratio and value of the parallel factors of each layer, calculate all integer factor pairs from 1 to the product of the proportionally enlarged parallel factors, and form a list of integer factor pairs; Expand the parallel factors of the layers with delays greater than the average value and perform the first fine-tuning: calculate the average value of the overall delay, traverse each layer, and when the delay is greater than the average, increase the parallel factor of the layer by a multiple, and then repeat the parallel factor product separation algorithm until the delay of the layer is less than or equal to the average delay, or no parallel factor combination that meets the constraints can be found; Reduce the parallelism factor of the layers whose latency is less than the maximum value, and perform secondary fine-tuning: find the layer with the largest latency, then traverse each layer in turn, gradually reduce the scale, and recalculate the corresponding latency until the latency of the layer is less than or equal to the maximum value; Analyze the latency, parallelism, and memory resource consumption of each module within the neural network layer, and establish a memory resource relationship model within the layer; Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factors of each layer of the neural network after allocation, the balance optimization of intra-layer memory resources and delay is achieved.
2. The method for optimizing the allocation of storage and computing resources of a neural network based on FPGA as claimed in claim 1, characterized in that: The analysis of the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and the establishment of the inter-layer computing resource relationship model specifically includes: Establish a double sliding window module delay model, and calculate the delay of the double sliding window module based on the size of the convolution kernel, the input width of the convolution layer, the number of channels, the number of paddings, and the output width and height of the convolution layer; By building a matrix-vector multiplication module delay model, the delay of the matrix-vector multiplication module is calculated based on the size of the convolution kernel, the input width, height, number of channels, the amount of padding, the parallelism within the convolution kernel and across the convolution kernels, and the number of inputs packed into a single DSP for calculation; Establish a DSP resource model, calculate the consumption of DSP resources based on the parallelism within the convolution kernel, the parallelism across the convolution kernel, the number of inputs packed into a single DSP, and the number of weights packed into a single DSP, and establish the priority of allocating the two parallelisms; Among them, the parallelism within the convolution kernel and the parallelism across the convolution kernel are the parallelism factor.
3. The method for optimizing the allocation of storage and computing resources of a neural network based on FPGA as claimed in claim 1, characterized in that: The establishing of the intra-layer memory resource relationship model specifically includes: Establish the relationship between the BRAM resource consumption of the dual sliding window module buffer and the BRAM configuration, buffer configuration, and the number of buffer copies; Calculate the width and depth of the double sliding window module buffer and the number of buffer copies; Establish the relationship between the amount of data read in parallel and the degree of parallelism in the buffer of the dual sliding window module; The relationship between the latency and parallelism of the operations required for the dual sliding window module and the intra-layer matrix-vector multiplication module to complete two convolutions in parallel is established respectively; Establish the relationship between the consumption of BRAM resources of the dual sliding window module buffer and parallelism and delay.
4. The method for optimizing the allocation of storage and computing resources of a neural network based on FPGA as claimed in claim 1, characterized in that: The optimal parallel factor of each layer of the neural network after allocation based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model, and the balance optimization of the intra-layer memory resources and delay is achieved, specifically: Calculate the maximum value of the start interval of the output data pipeline of the double sliding window module under the constraints; Calculate the number of required dual sliding window module buffer copies according to the maximum value of the startup interval; According to the number of buffer copies of the dual sliding window module, the consumption under different buffer configurations is calculated and the buffer configuration with the minimum consumption is output.
5. The method for optimizing the allocation of storage and computing resources of a neural network based on FPGA as claimed in claim 4, characterized in that: By comparing the delay of completing a convolution calculation in the dual sliding window module and the matrix-vector multiplication module, the startup interval of the dual sliding window module output pipeline is determined. Then, the minimum number of ports of the buffer in the dual sliding window module is determined according to the startup interval and the amount of parallel read data. The BRAM resource consumption of the buffer is minimized while ensuring that the delay of the overall pipeline remains unchanged, that is, the balance optimization of memory resources and delay is achieved. The specific formula is: in, represents the delay of the double sliding window module outputting the data of a single convolution window from the buffer to the computing unit, Indicates the delay of the matrix-vector multiplication module to complete a single convolution calculation, Indicates that the constraints The maximum start interval of the output data pipeline of the lower double sliding window module, Represents the size of the convolution kernel of the i-th layer; and Represent the number of input channels and output channels of the i-th convolutional layer respectively; Represents the degree of parallelism within a single convolution kernel, indicating the number of elements calculated simultaneously in a convolution; Represents the parallelism across convolution kernels, indicating the number of convolution kernels calculated simultaneously.
6. A neural network storage and computing resource optimization allocation system based on FPGA, characterized in that: include: The computing resource allocation module is configured to: analyze the parallel factor and the delay of each layer of the neural network, the parallel factor and the computing resource consumption of the FPGA, and establish a computing resource relationship model between layers; Based on the inter-layer computing resource relationship model, the optimal parallel factor of each layer of the neural network after allocation is obtained through the computing resource allocation method. The specific steps are: Perform initial allocation based on the inter-layer computing resource relationship model: determine the ratio and value of the parallel factors of each layer, calculate all integer factor pairs from 1 to the product of the proportionally enlarged parallel factors, and form a list of integer factor pairs; Expand the parallel factors of the layers with delays greater than the average value and perform the first fine-tuning: calculate the average value of the overall delay, traverse each layer, and when the delay is greater than the average, increase the parallel factor of the layer by a multiple, and then repeat the parallel factor product separation algorithm until the delay of the layer is less than or equal to the average delay, or no parallel factor combination that meets the constraints can be found; Reduce the parallelism factor of the layers whose latency is less than the maximum value, and perform secondary fine-tuning: find the layer with the largest latency, then traverse each layer in turn, gradually reduce the scale, and recalculate the corresponding latency until the latency of the layer is less than or equal to the maximum value; The memory resource and delay balance optimization module is configured to: analyze the delay and parallelism of each module in the neural network layer, and the memory resource consumption, and establish a memory resource relationship model within the layer; Based on the inter-layer computing resource relationship model, the intra-layer memory resource relationship model and the optimal parallel factors of each layer of the neural network after allocation, the balance optimization of intra-layer memory resources and delay is achieved.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are performed.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Binary neural network acceleration method and system based on FPGA
CN110458279A
Fully homomorphic encryption neural network reasoning acceleration method and system based on resource reuse
CN116048811A