Accelerator multi-objective optimization method driven by adaptive reinforcement learning
By combining iterative optimization techniques with adaptive reinforcement learning, the resource and configuration issues of deploying large-scale deep neural networks on FPGA platforms were resolved, achieving efficient convolutional accelerator parameter configuration and improving deployment efficiency and adaptability.
Patent Information
- Application Number
- CN202511449492.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-11
AI Technical Summary
The deployment of large-scale deep neural networks on FPGA platforms currently faces challenges such as limited computing resources, bottlenecks in off-chip communication bandwidth, and a large parameter configuration space, resulting in low deployment efficiency, poor hardware adaptability, and difficulty in completing the global optimal configuration within a reasonable time using traditional methods.
By combining iterative optimization techniques with adaptive reinforcement learning, a joint objective function is designed by constructing a state-action-reward interaction mechanism, and a lightweight search module is used to optimize parameter configuration, thereby achieving efficient deployment of the convolutional accelerator.
It significantly improves the efficiency and intelligence of accelerator deployment, optimizes resource utilization and adaptability, and is particularly suitable for resource-constrained embedded SoC platforms.
Smart Images

Figure CN120911543A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of optimized design of accelerators, and relates to an adaptive reinforcement learning driven accelerator multi-objective optimization method, in particular to a high-efficiency parameter configuration search and hardware mapping method of a convolution operation accelerator realized through a loop optimization technology and an adaptive reinforcement learning driven accelerator multi-objective optimization framework. BACKGROUND
[0002] In recent years, a convolutional neural network (CNN) accelerator based on an FPGA (Field-Programmable Gate Array) has shown wide application potential in the fields of image recognition, target detection, edge computing and the like due to its excellent energy efficiency ratio, flexible reconfigurability and ability to adapt to diverse task scenarios. In particular, in embedded and low-power devices, the FPGA provides a solution that takes into account both performance and energy efficiency for model deployment.
[0003] A convolution operation is a core operator for image and feature processing in deep learning, which mainly performs local calculation on an input feature map through a sliding convolution kernel to extract spatial features. Figure 1 The basic process of a convolution operation is shown. The convolution operation involves six-dimensional calculation: the number of input feature map channels (N if ), the number of output feature map channels (N of ), the height of the output feature map (N iy ), the width of the input feature map (N ix ), the height of the convolution kernel (N ky ), and the width of the convolution kernel (N kx ).
[0004] In actual implementation, the convolution operation is usually completed by four layers of nested loops, which are named Loop-1, Loop-2, Loop-3 and Loop-4. Among them, the horizontal direction (height and width) of the input image is processed by the nested loop Loop-3, which corresponds to the sliding window operation of the convolution kernel on the input image; the horizontal direction (height and width) of the convolution kernel is processed by the nested loop Loop-1, which represents the multiply-accumulate (MAC) operation of the convolution kernel under a single channel; the input channel number and the output channel number are processed by two independent nested loops Loop-2 and Loop-4 respectively, because each output channel needs to traverse all input channels. In this way, the original six-dimensional convolution calculation is effectively simplified into four layers of nested loops, which not only improves the calculation efficiency, but also optimizes the utilization rate of hardware resources.
[0005] Currently, deploying large-scale deep neural networks (e.g., VGG, ResNet) on FPGA platforms still faces many challenges, mainly including limited computing resources, off-chip communication bandwidth bottlenecks, and large deployment parameter configuration space. These factors together lead to low deployment efficiency, poor hardware adaptability, and large manual parameter tuning workload. To address these challenges, researchers widely use loop optimization techniques to efficiently map convolution operations. Common methods include loop unrolling, loop tiling, and loop interchange. Loop unrolling increases parallel processing units to improve computing throughput, loop tiling divides feature maps and weights into blocks to adapt to on-chip cache and reduce off-chip access, and loop interchange adjusts the calculation order to optimize bandwidth usage and data reuse strategies.
[0006] However, current research still has obvious deficiencies in loop optimization. For example, most designs only focus on the inner nested loops (e.g., Loop-1 and Loop-2), ignoring the differences in output spatial dimensions (Loop-3 and Loop-4) between convolution layers and the uneven impact of resource usage, which directly limits the accelerator's adaptability and scalability across network layers. In addition, on resource-constrained FPGA platforms, if the parallelism is fixed at the outer loop, it can easily lead to on-chip resource waste and runtime bottlenecks, especially when dealing with 1x1 convolution or residual structures.
[0007] Current research attempts to optimize performance on specific convolution layers through data reuse mechanisms and automated structure search tools. For example, some accelerators customized for 5x5 convolution kernels can lead to resource waste when facing 1x1 convolution layers; and designs that concentrate all parallelism in Loop-4 have efficiency problems when dealing with shallow networks with small spatial receptive fields. Therefore, another key difficulty in CNN deployment is the high dimension and strong coupling of the parameter configuration space. Hardware implementation parameters (e.g., P if , P of , K b , F bThere is a complex interaction between them, and traditional manual experience method or static search algorithm (such as greedy search, genetic algorithm) is difficult to complete global optimal configuration within a reasonable time. To solve this problem, reinforcement learning (Reinforcement Learning, RL) is introduced into the accelerator design space exploration (Design Space Exploration, DSE) task. Among them, the deep deterministic policy gradient (Deep Deterministic Policy Gradient, DDPG) algorithm, that is, the reinforcement learning agent, because of its ability to handle continuous action space, has become a powerful tool for adapting CNN accelerator tuning tasks. SUMMARY
[0008] The application proposes an accelerator multi-objective optimization method driven by adaptive reinforcement learning, which combines cyclic optimization technology and reinforcement learning strategy, and can automatically optimize the parameter configuration combination of convolution accelerator under the conditions of platform resource constraints and network model structure. The method builds a state-action-reward interaction mechanism, designs a joint objective function, and combines a lightweight search module to decouple and optimize some independent factors, effectively improving the efficiency and intelligence level of accelerator deployment, especially suitable for resource-constrained embedded SoC (System-on-a-Chip) platforms. Specifically, the application first analyzes the mechanism of convolution loop unrolling factor, quantization precision, and cache partition key design variables, establishes a system-level state modeling method, and designs a DDPG framework suitable for continuous action space optimization as a configuration strategy learner; at the same time, a lightweight searcher is implemented by running a Python script on the CPU to search and allocate the block parameters of Loop-3 and Loop-4 of convolution operation in real time, thereby reducing the computational burden of the reinforcement learning agent. Finally, an adaptive configuration method driven by reinforcement learning and a matching search process are proposed to realize the joint optimization of unrolling scale, quantization precision and block parameters.
[0009] The technical scheme of the application is as follows:
[0010] An adaptive reinforcement learning driven accelerator multi-objective optimization method, the steps are as follows:
[0011] Step 1: Build a search framework for reinforcement learning model (RL).
[0012] Step 1.1: Define the state space: establish a state vector that can describe the accelerator system, including the structure parameters of convolutional neural network (CNN), the available resources of the current FPGA, and the action selected by the reinforcement learning agent in the current state.
[0013] The state vector is represented as: , ;
[0014] wherein, and denote the width and height of the convolution kernel, respectively, and denote the width and height of the input feature map, respectively, and denote the number of channels of the input feature map and the output feature map, respectively, and denote the width and height of the output feature map, respectively, and denote the step size of the convolution operation in the horizontal direction and the vertical direction, respectively, denotes the number of convolution kernel groups or parallel computing clusters, used to distinguish between grouped convolution or parallel scale, denotes the number of on-chip cache resources (BRAM / URAM resources), denotes the number of available DSP computing units, denotes the first action selected by the reinforcement learning agent, denotes the second action selected by the reinforcement learning agent, or .
[0015] Step 1.2: Define the continuous action space: The designed continuous action space includes two types of factors: one is the hardware-related loop unrolling factor, and the other is the convolutional neural network-related quantization precision factor. To ensure that the generated actions meet the resource constraints of the FPGA, the continuous actions need to be mapped to the preset search interval after outputting to obtain legal configuration parameters through the discretization function.
[0016] In each training iteration, the reinforcement learning model generates 4 actions. The 4 actions are divided into two groups: Action-1 and Action-2. Among them, Action-1 contains three actions: , , 2 or 3; corresponding to the loop unrolling factor , where represents the loop unrolling factor in the input channel direction, which determines the input channel parallelism; represents the loop unrolling factor in the horizontal direction of the output feature map, which determines the spatial parallelism; represents the loop unrolling factor in the output channel direction, which determines the output channel parallelism. Action-2 contains one action , ; used to select the quantization precision factor of the convolutional neural network , where represents the quantization bit width of the convolution weight, Quantization bit-width of input feature map. The convolution operation is completed by four layers of nested loops, which are named Loop-1, Loop-2, Loop-3, and Loop-4, respectively. The horizontal direction (height and width) of the input image is processed by the nested loop Loop-3. The horizontal direction (height and width) of the convolution kernel is processed by the nested loop Loop-1. The input channel number and the output channel number are processed by two independent nested loops Loop-2 and Loop-4, respectively. The four actions are generated as follows:
[0017] For the generation process of the action in Action-1, the reinforcement learning agent first outputs a continuous value in the interval [0, 1] , which is mapped to an integer value of the loop-unfolding factor by the discretization function . The mapping process is scaled according to the search interval of each loop-unfolding factor, so as to ensure that the output result meets the resource constraints of FPGA. The formula is as follows:
[0018]
[0019] For the generation process of the action in Action-2, the reinforcement learning agent outputs a continuous number , and maps it to a discrete bit-width by the following formula:
[0020]
[0021] where and represent the minimum and maximum values of the optional quantization bit-width, respectively.
[0022] Step 1.3: Constructing a joint reward function: according to the comprehensive constraints of inference delay and precision, a joint reward function is designed to evaluate the pros and cons of each action selection and guide the reinforcement learning model to gradually learn the optimal configuration strategy.
[0023] In the reinforcement learning model, in each round of iterative training, the convolutional neural network is quantized by the selected unified quantization bit-width , and then the quantized inference prediction delay is obtained by calling the delay model. Then is compared with the target delay upper limit . If , the convolutional neural network quantization model is fine-tuned for one epoch to restore the precision, and the precision index is obtained. If , the fine-tuning is skipped and a penalty is given. The joint reward function is defined in a piecewise form:
[0024]
[0025] wherein, is the validation set accuracy of the quantized convolutional neural network model, is the benchmark accuracy of the quantized convolutional neural network model, is a scaling factor, the reinforcement learning model explores E episodes under each experimental setting, calculates the reward according to the above formula and updates the policy accordingly until convergence.
[0026] Step 2: define the parameter optimization process and hardware mapping strategy;
[0027] Step 2.1: experience replay and input sampling: store the data generated by the interaction between the reinforcement learning agent and the environment through the experience replay mechanism, and randomly sample in the training process, so as to break the time correlation and ensure the diversity and independence of the training samples.
[0028] The reinforcement learning model first stores the interaction data of the reinforcement learning agent in the environment in the experience replay area, and then randomly samples a batch of data in the training process, so as to contact more diverse and independent samples in the training to improve the stability and convergence speed of the training; the interaction data includes the current state, the current action, the current reward and the next state.
[0029] Step 2.2: reinforcement learning model training: use the data obtained by random sampling to train the reinforcement learning model, which predicts Q values according to the input state and action, and continuously improves the policy quality through parameter update to ensure the gradual convergence of the learning process.
[0030] The data obtained by random sampling is used for the training of the reinforcement learning model, which will predict a value, i.e. Q value, according to the current state and action. Q value represents the long-term comprehensive income that can be obtained after taking a certain action in a given state, which is the core index to measure the quality of the action. In the training process, the reinforcement learning model constantly updates the parameters by referring to the Q value to ensure the stability of the policy learning and avoid violent shocks.
[0031] Step 2.3: policy update and action selection: adopt The greedy strategy balances between exploration and utilization. When the exploration condition is met, a random action is selected, otherwise the action with the highest Q value is selected. The final output action is the optimal configuration scheme of the accelerator system.
[0032] When the reinforcement learning model is trained, the accelerator system selects actions according to the greedy strategy. The specific operation is as follows: first, the accelerator system generates a random number between 0 and 1 and compared with a pre-set threshold value When the accelerator system randomly selects an action to ensure the exploration of the policy; when the accelerator system selects an action with the highest predicted Q value to ensure the utilization of existing experience. The selected action corresponds to the optimal parameter configuration of the accelerator system, including the loop unrolling factor and the quantization precision factor.
[0033] Step 3: Introduce a lightweight searcher to complete the decoupled parameter optimization
[0034] Each network layer L of the convolutional neural network has an independent block configuration , i.e., the block size in Loop-3 and Loop-4 directions, which determines the data partition scale of each network layer. A lightweight searcher is implemented on the CPU side to run Python scripts in real time, which is used to search for the optimal loop blocking parameters in real time, thereby achieving efficient blocking optimization without additional time overhead. During the data reading process of each network layer L, the accelerator system will run the search process on the CPU to dynamically find the optimal block configuration .
[0035] Step 3.1: Preliminary optimization path selection. Determine whether each loop is tiled, and select paths that meet the conditions and can enter subsequent optimization to avoid resource waste or bandwidth bottlenecks caused by unreasonable configurations.
[0036] First, check whether Loop-1 and Loop-2 are not tiled, which may cause additional transmission overhead due to computation. Then, determine whether Loop-3 and Loop-4 are tiled, which may have the possibility of further optimization if they are tiled, and then enter the subsequent data flow analysis and mapping strategy. If the conditions are not met, it may cause computation idle or bandwidth bottleneck.
[0037] Step 3.2: Data transmission optimization and resource allocation. Under the premise of meeting resource constraints, try multiple allocation attempts for on-chip cache and DSP computation units, calculate the data transmission amount of each configuration scheme, and select the optimal configuration scheme to achieve the balance between storage and computation.
[0038] In the case where Loop-3 and Loop-4 are tiled, first determine whether and If the condition is met, the pixel buffer and the weight buffer are allocated in different proportions, and the data transmission BW1 and BW2 under different configuration schemes are calculated, and the optimal configuration scheme is screened out. Among them, BW1 and BW2 respectively represent the total data transmission when Loop-3 and Loop-4 are used as inner loop, which is used to compare the bandwidth consumption under different loop unrolling orders, so as to screen the optimal configuration scheme. representing the loop unrolling factor , representing the block configuration , representing the scale of the convolutional neural network . representing the loop unrolling factor not more than the block configuration , otherwise the hardware parallelism is greater than the single block data volume, which will cause idle and waste; the block configuration cannot exceed the total scale of the convolutional neural network , otherwise the block configuration loses its meaning. In the condition, represents the time required to transfer the feature block into the on-chip cache, is the data processing time under the current design. The calculation formula is as follows:
[0039]
[0040] Among them, represents the transmission frequency of off-chip input data; represents the data transmission bit width of off-chip bus; , , respectively represent the block size of the width of the input feature map, the height of the input feature map, and the number of channels of the input feature map; , , respectively represent the block size of the width of the convolution kernel, the height of the convolution kernel, and the number of channels of the output feature map, which is used to describe the size of the data block that can be processed in the on-chip cache at a time.
[0041] Step 3.3: Strategy selection and loop unrolling: efficiency comparison is performed on the candidate schemes, the optimal loop unrolling order is finally determined, and the result is output as the hardware configuration, realizing efficient execution of convolution operation and data transmission optimization.
[0042] The configuration scheme with the smallest data transfer volume is selected, and BW1 and BW2 are compared: if BW1 is greater than BW2, Loop-3 is selected as the inner loop; otherwise, Loop-4 is selected. When Loop-3 is used as the inner loop, the weighted regroups of the input pixel blocks after each segment are read sequentially to improve the utilization of the on-chip pixel buffer and weight buffer. When Loop-4 is selected as the inner loop, the accelerator system loads pixels into the pixel buffer and weight buffer sequentially to support parallel reading of weight groups. Finally, based on the evaluation results, the optimal loop unrolling order and corresponding cache management strategy are output, achieving efficient execution of convolution operations and optimal data transfer configuration.
[0043] The beneficial effects of this invention are as follows: This invention combines reinforcement learning with lightweight search to achieve intelligent automatic optimization of convolutional accelerator parameters. While meeting latency constraints, it also takes into account accuracy and resource utilization, significantly improving deployment efficiency and adaptability. Attached Figure Description
[0044] Figure 1 This describes the basic process of convolution operations. Among them, and These represent the step size in the horizontal and vertical directions, respectively.
[0045] Figure 2 A search framework for reinforcement learning models.
[0046] Figure 3 This is a process for determining and initially optimizing the path selection for cyclic tiling.
[0047] Figure 4 To evaluate the experimental results of an adaptive reinforcement learning-driven accelerator multi-objective optimization method, where (a) is the consistency between the predicted and observed delays; and (b) is the prediction error. Detailed Implementation
[0048] The present invention will now be described in further detail with reference to the accompanying drawings and technical solutions.
[0049] An adaptive reinforcement learning-driven multi-objective optimization method for accelerators, comprising the following steps:
[0050] Step 1: Construct the search framework for the reinforcement learning (RL) model, the details of which are as follows: Figure 2 As shown.
[0051] Step 1.1: Define the state space: Establish a state vector that can describe the accelerator system. The state vector includes the structural parameters of the convolutional neural network (CNN), the available resources of the current FPGA, and the action selected by the reinforcement learning agent in the current state, which is used to comprehensively characterize the state of the accelerator system.
[0052] The state vector is represented as: , ;
[0053] wherein, and represent the width and height of the convolution kernel, and represent the width and height of the input feature map, and represent the number of channels of the input feature map and the output feature map, and represent the width and height of the output feature map, and represent the step size of the convolution operation in the horizontal direction and the vertical direction, represents the number of convolution kernel groups or parallel computing clusters, used to distinguish between grouped convolution or parallel scale, represents the number of on-chip cache resources (BRAM / URAM resources), represents the number of available DSP computing units, represents the action selected by the reinforcement learning agent, or . In order to ensure the comparability of parameters of different dimensions in the training process, the state variables in all the above state vectors are mapped to the interval [0, 1] through normalization, in order to improve the stability and convergence speed of the reinforcement learning training. Figure 1 is the basic process of the convolution operation.
[0054] Step 1.2: Define the continuous action space: The designed continuous action space includes two types of factors: one is the hardware-related loop-unrolling factor, and the other is the convolutional neural network-related quantization precision factor. In order to ensure that the generated actions meet the resource constraints of the FPGA, the continuous actions need to be mapped to the preset search interval through a discretization function after output, in order to obtain legal configuration parameters.
[0055] In each training iteration, the reinforcement learning model generates 4 actions. The 4 actions are divided into two groups: Action-1 and Action-2. Among them, Action-1 contains three actions: , = 1, 2 or 3; respectively corresponding to the loop-unrolling factor , wherein represents the loop-unrolling factor in the input channel direction (Loop-2), which determines the input channel parallelism; represents the loop-unrolling factor in the horizontal direction of the output feature map (Loop-3), which determines the spatial parallelism; Loop-4, which represents the output channel direction, determines the output channel parallelism. Action-2 contains an action , ; for selecting a quantization precision factor of a convolutional neural network , wherein represents a quantization bit width of a convolutional weight, represents a quantization bit width of an input feature map. In the present application and are kept consistent, for ensuring uniform data format in the convolutional calculation process and reducing hardware overhead. The convolutional operation is completed by using four layers of nested loops, which are named Loop-1, Loop-2, Loop-3, and Loop-4, respectively. The horizontal direction (height and width) of the input image is processed by the nested loop Loop-3. The horizontal direction (height and width) of the convolutional kernel is processed by the nested loop Loop-1. The input channel number and the output channel number are processed by two independent nested loops Loop-2 and Loop-4, respectively. The four actions are generated as follows:
[0056] For the generation process of the action in Action-1, the reinforcement learning agent first outputs a continuous value in the interval [0, 1], which is mapped to an integer value of the loop-unfold factor by a discretization function. The mapping process is scaled according to the search interval of each loop-unfold factor, so as to ensure that the output result meets the resource constraints of the FPGA. The formula is as follows:
[0057]
[0058] For the generation process of the action in Action-2, the reinforcement learning agent outputs a continuous number , and maps it to a discrete bit width by the following formula:
[0059]
[0060] wherein and represent the minimum value and the maximum value of the selectable quantization bit width, respectively. Assuming that the selectable bit width in the example of the present application is {4, 8, 16}, then , , the span is 16−4+1=13, and the “−0.5” in the formula is used to cooperate with the operation, so that the output value can be correctly rounded to the nearest legal bit width. The finally obtained This refers to the integer bit width of the quantization precision factor (e.g., 4, 8, 16). Action-2 only needs to determine one quantization bit width, and only needs to output one action in a single reinforcement learning process.
[0061] Step 1.3: Construct a joint reward function: Based on the combined constraints of inference latency and accuracy, design a joint reward function to evaluate the merits of each action selection and guide the reinforcement learning model to gradually learn the optimal configuration strategy.
[0062] In reinforcement learning models, in each round of training iterations, a convolutional neural network quantization model is first used to quantize the model according to the selected uniform quantization bit width. The convolutional neural network is quantized, and then a delay model is called to obtain the quantized inference prediction delay. ,Will With target latency limit Compared to. If When the inference prediction delay meets the delay constraint, the convolutional neural network quantization model is fine-tuned by one epoch to restore accuracy, thus obtaining the accuracy index. .like When the inference prediction delay exceeds the limit, skip the fine-tuning training and apply a penalty directly. The joint reward function is defined in piecewise form:
[0063]
[0064] in, To quantize the validation set accuracy of the convolutional neural network model, To quantize the baseline accuracy of a convolutional neural network model, such as the accuracy of an unquantized or established reference model on the same validation set; The scaling factor, preferably 0.01, is used to map the accuracy difference to a reasonable range of (-1, 1). This design prioritizes satisfying the latency constraint (exceeding the constraint results in a negative reward, and the magnitude of the negative reward increases with the excess rate), and then pursues higher accuracy after satisfying the latency constraint. Simultaneously, the "latency-first, fine-tuning" process avoids unnecessary training on unsuitable configurations, reducing search overhead. The reinforcement learning model explores E episodes in each experimental setting, calculates the reward according to the above formula, and updates the policy accordingly until convergence.
[0065] Step 2: Define the parameter optimization process and hardware mapping strategy;
[0066] Step 2.1: Experience Replay and Input Sampling: The experience replay mechanism stores the data generated by the interaction between the reinforcement learning agent and the environment, and randomly samples it during the training process, thereby breaking the temporal correlation and ensuring the diversity and independence of training samples.
[0067] The reinforcement learning model first stores interaction data of a reinforcement learning agent in an environment in an experience replay area, and then randomly samples a batch of data in the experience replay area during training, so that more diverse and independent samples are contacted during training to improve the stability and convergence speed of training. The interaction data includes a current state, a current action, a current reward, and a next state.
[0068] Step 2.2: Reinforcement learning model training: The data obtained by random sampling is used to train the reinforcement learning model. The reinforcement learning model predicts a Q value according to the input state and action, and continuously improves the policy quality through parameter updating to ensure gradual convergence of the learning process.
[0069] The data obtained by random sampling is used for training of the reinforcement learning model. The reinforcement learning model will predict a value, i.e., a Q value, according to the current state and the current action. The Q value represents the long-term comprehensive income that can be obtained after taking a certain action in a given state, and is a core indicator for measuring the quality of the action. During the training process, the reinforcement learning model continuously updates the parameters by referring to the Q value to ensure the stability of policy learning and avoid violent shocks.
[0070] Step 2.3: Policy update and action selection: the greedy policy is adopted The greedy policy balances exploration and utilization. When the exploration condition is met, a random action is selected, otherwise the action with the highest current Q value is selected. The final output action is the optimal configuration scheme of the accelerator system.
[0071] When the reinforcement learning model is trained, the accelerator system selects an action according to the greedy policy. The specific operation is as follows: first, the accelerator system generates a random number between 0 and 1 , and compares it with a preset threshold ; when , the accelerator system randomly selects an action to ensure the exploratory nature of the policy; when , the accelerator system selects the action with the highest predicted Q value to ensure the use of existing experience. It should be noted that is a randomly generated number used to trigger the exploration event; is a manually set exploration probability threshold, which is usually set to a larger value at the beginning of training to increase exploration, and then gradually decays to enhance utilization. The comparison between the two is essentially a trade-off between "exploring new actions" and "utilizing optimal actions" through probability, so as to ensure that the policy can avoid falling into a local optimum and gradually converge to an optimal solution. The selected action corresponds to the optimal parameter configuration of the accelerator system, including the loop unrolling factor and the quantization precision factor.
[0072] Step 3: Introduce a lightweight searcher to complete decoupled parameter optimization;
[0073] Each network layer L of each convolutional neural network has an independent block configuration , i.e., the block size in Loop-3 and Loop-4 directions, which determines the data partition scale of each network layer. In order to reduce the computational burden of the reinforcement learning agent, the present application implements a lightweight searcher on the CPU side in real time by running a Python script, which is used to search the loop block parameters in real time, so as to realize efficient block optimization without additional time overhead. During the data reading process of each network layer L, the accelerator system will run the search process as shown in Figure 3 on the CPU to dynamically find the optimal block configuration .
[0074] Step 3.1: preliminary optimization path selection. Determine whether each loop is tiled, and select the path that meets the conditions and can enter the subsequent optimization to avoid resource waste or bandwidth bottleneck caused by unreasonable configuration. The specific content is shown in Figure 3 .
[0075] First, check whether Loop-1 and Loop-2 are not tiled. If not tiled, it may cause additional transmission overhead caused by calculation. Then, determine whether Loop-3 and Loop-4 are tiled. If tiled, it has the possibility of further optimization, and then enter the subsequent data flow analysis and mapping strategy. If the conditions are not met, it may cause calculation idle or bandwidth bottleneck.
[0076] Step 3.2: data transmission optimization and resource allocation. Under the premise of meeting the resource constraints, try multiple allocation of on-chip cache and DSP calculation unit, calculate the data transmission amount of each configuration scheme, and select the optimal configuration scheme to realize the balance of storage and calculation.
[0077] In the case that Loop-3 and Loop-4 are tiled, first determine whether and are met. If the conditions are met, allocate the pixel buffer and the weight buffer according to different proportions in turn, calculate the data transmission amount BW1, BW2 under different configuration schemes, and then select the optimal configuration scheme. Among them, BW1 and BW2 represent the total data transmission amount when Loop-3 and Loop-4 are inner loops, which are used to compare the bandwidth consumption under different loop unrolling orders, so as to select the optimal configuration scheme. represent the loop unrolling factor , represent the block configuration , BW1 and BW2 represent the total data transmission amount when Loop-3 and Loop-4 are inner loops, which are used to compare the bandwidth consumption under different loop unrolling orders, so as to select the optimal configuration scheme. . represent the loop unrolling factor The block configuration should not exceed the total CNN scale Otherwise, the hardware parallelism is greater than the single-block data volume, which will cause idle and waste. The block configuration should not exceed the total CNN scale Otherwise, the block configuration loses its meaning. In the condition, represents the time required to transfer the feature block into the on-chip cache, is the data processing time under the current design. The calculation formula is as follows:
[0078]
[0079] Among them, represents the transmission frequency of off-chip input data; represents the data transmission bit width of the off-chip bus; , , respectively represent the block size of the input feature map width, the input feature map height, and the input feature map channel number. , , respectively represent the block size of the convolution kernel width, the convolution kernel height, and the output feature map channel number, which are used to describe the size of the data block that can be processed in the on-chip cache at a time.
[0080] Step 3.3: Strategy selection and loop unrolling: compare the efficiency of the candidate schemes, finally determine the optimal loop unrolling order, and output the result as the hardware configuration to realize efficient execution of convolution operation and data transmission optimization.
[0081] Select the configuration scheme with the smallest data transmission amount, compare BW1 and BW2: if BW1 is greater than BW2, select Loop-3 as the inner loop; otherwise, select Loop-4 as the inner loop. When Loop-3 is used as the inner loop, the weight groups of each block of input pixel blocks are read in sequence to improve the utilization rate of the on-chip pixel buffer and the weight buffer. When Loop-4 is selected as the inner loop, the accelerator system loads the pixels into the pixel buffer and the weight buffer in sequence to support parallel reading of the weight groups. Finally, according to the evaluation results, the optimal loop unrolling order and the corresponding cache management strategy are output to realize efficient execution of convolution operation and optimal configuration of data transmission.
[0082] The method is verified by experimental verification on the Xilinx ZCU102 FPGA platform. The search space formed on this platform contains about 2430 accelerator configuration combinations, and these configurations are tested to evaluate the prediction accuracy of the reinforcement learning model for neural network inference delay. The experimental results are as follows: Figure 4As shown: Figure 4 (a) in the figure shows the distribution of prediction error as a function of actual delay; Figure 4 Figure (b) shows the alignment between the predicted latency and the actual observed latency. Most sample points are distributed near the diagonal, indicating that the model's predictions are highly consistent with the measured values. It can be seen that the error fluctuation is relatively large when the workload is small and the latency is low, while the prediction error gradually decreases and tends to stabilize as the workload and latency increase. Overall, the prediction error of most samples is controlled below 3%, verifying that the constructed reinforcement learning model has high prediction accuracy under different configurations.
Claims
1. An accelerator multi-objective optimization method driven by adaptive reinforcement learning, characterized in that, The steps are as follows: Step 1: Construct a search framework for the reinforcement learning model; Step 1.1: Define the state space: Establish a state vector that can describe the state of the accelerator system, including the structural parameters of the convolutional neural network, the available resources of the current FPGA, and the action selected by the reinforcement learning agent in the current state; The state vector is represented as: ; in, and These represent the width and height of the convolution kernel, respectively. and These represent the width and height of the input feature map, respectively. and These represent the number of channels in the input feature map and the output feature map, respectively. and These represent the width and height of the output feature map, respectively. and These represent the stride of the convolution operation in the horizontal and vertical directions, respectively. This indicates the number of convolution kernel groups or the number of clusters for parallel computation, used to distinguish between grouped convolution and parallel scale. Indicates the amount of on-chip cache resources (BRAM / URAM resources). Indicates the number of available DSP computing units. This indicates the first step chosen by the reinforcement learning agent. One action, or ; Step 1.2: Define the continuous action space: The designed continuous action space includes two types of factors, one is the hardware-related loop unrolling factor, and the other is the convolutional neural network-related quantization precision factor; To ensure that the generated action meets the resource constraints of the FPGA, the continuous action needs to be mapped to the preset search interval through a discretization function after output to obtain the legal configuration parameters; In each training iteration, the reinforcement learning model generates 4 actions; the 4 actions are divided into two groups: Action-1 and Action-2; wherein Action-1 contains three actions: , , 2 or 3; respectively corresponding to the loop-unfold factor , wherein represents the loop-unfold factor of the input channel direction, which determines the input channel parallelism; represents the loop-unfold factor of the output feature map horizontal direction, which determines the spatial parallelism; represents the loop-unfold factor of the output channel direction, which determines the output channel parallelism; Action-2 contains one action , ; used to select the quantization precision factor of the convolutional neural network , wherein represents the quantization bit width of the convolution weight, represents the quantization bit width of the input feature map; the convolution operation is completed using four levels of nested loops, which are named Loop-1, Loop-2, Loop-3, and Loop-4; wherein the horizontal direction of the input image is processed by the nested loop Loop-3; the horizontal direction of the convolution kernel is processed by the nested loop Loop-1; the input channel number and the output channel number are processed by two independent nested loops Loop-2 and Loop-4, respectively; the specific generation process of the 4 actions is as follows: For the generation process of the action in Action-1, the reinforcement learning agent first outputs a continuous value in the interval [0, 1] , and maps the continuous value to an integer value of the loop unrolling factor by a discretization function ; the mapping process is scaled according to the search interval of each loop unrolling factor , so as to ensure that the output result meets the resource constraints of the FPGA; the formula is as follows: For the generation process of actions in Action-2, the reinforcement learning agent outputs a continuous number and is mapped to a discrete bit-width by the following equation: wherein and respectively represent the minimum and maximum values of the optional quantization bit-width; Step 1.3: Construct a joint reward function: According to the comprehensive constraints of inference delay and precision, design a joint reward function to evaluate the pros and cons of each action selection and guide the reinforcement learning model to gradually learn the optimal configuration strategy; In the reinforcement learning model, in each round of iteration training, the convolutional neural network quantization model is used first to select a uniform quantization bit width The convolutional neural network is quantized, and then the quantized inference prediction delay is obtained by calling the delay model , The target delay upper limit is compared; if , the convolutional neural network quantization model is fine-tuned for one epoch to restore accuracy, and the accuracy index is obtained; if , the fine-tuning is skipped and a penalty is directly given; the joint reward function is defined in a segmented form: wherein, is the validation set accuracy of the quantized convolutional neural network model, is the benchmark accuracy of the quantized convolutional neural network model, is the scaling factor, the reinforcement learning model explores E episodes under each experimental setting, the reward is calculated according to the above formula and the policy is updated accordingly until convergence. Step 2: Define the parameter optimization process and hardware mapping strategy; Step 2.1: Experience replay and input sampling: Store the data generated by the interaction between the reinforcement learning agent and the environment in the experience replay area, and randomly sample in the training process to break the time correlation and ensure the diversity and independence of the training samples; The reinforcement learning model first stores the interaction data of the reinforcement learning agent in the environment in the experience replay area, and then randomly samples a batch of data in the training process, so that more diverse and independent samples are contacted during training to improve the stability and convergence speed of training; The interaction data includes the current state, the current action, the current reward, and the next state; Step 2.2: Reinforcement learning model training: Use the data obtained by random sampling to train the reinforcement learning model, which predicts Q values based on the input state and action, and continuously improves the policy quality through parameter updates to ensure gradual convergence of the learning process; The data obtained by random sampling is used for training of the reinforcement learning model, which will predict a value, i.e. Q value, based on the current state and action; Q value represents the long-term comprehensive income obtained after taking a certain action in a given state, which is the core indicator for measuring the quality of the action; During training, the reinforcement learning model continuously updates the parameters by referring to the Q value to ensure the stability of policy learning and avoid drastic fluctuations; Step 2.3: Policy update and action selection: adopt The greedy policy balances exploration and exploitation; when the exploration condition is met, a random action is selected, otherwise the action with the highest current Q-value is selected; the final output action is the optimal configuration scheme of the accelerator system; Once the reinforcement learning model is trained, the accelerator system... The greedy strategy selects actions; specifically, the accelerator system first generates a random number between 0 and 1. and with preset threshold Compare; when At one time, the accelerator system randomly selects an action to ensure the exploratory nature of the strategy; when At that time, the accelerator system selects the action with the highest current predicted Q value to ensure that existing experience is utilized; the selected action corresponds to the optimal parameter configuration of the accelerator system, including the cycle expansion factor and the quantization accuracy factor; Step 3: Introduce a lightweight searcher to complete decoupled parameter optimization; Each network layer L of the convolutional neural network has an independent block configuration , i.e., the block size in the Loop-3 and Loop-4 directions, which determines the data partition scale of each network layer; a lightweight searcher is implemented by running a Python script on the CPU in real time, which is used to search the loop block parameters in real time, thereby achieving efficient block optimization without additional time overhead; during the data reading process of each network layer L, the accelerator system runs the search process on the CPU to dynamically find the optimal block configuration ; Step 3.1: Preliminary optimization path selection; Determine whether each loop has been tiled, and select paths that meet the conditions and can enter subsequent optimization to avoid resource waste or bandwidth bottlenecks caused by unreasonable configuration; First, check whether Loop-1 and Loop-2 are not tiled, if not, it may cause additional transmission overhead due to calculation; Then determine whether Loop-3 and Loop-4 have been tiled, if they have been tiled, they have the possibility of further optimization, then enter the subsequent data flow analysis and mapping strategy; If the conditions are not met, it may cause idle calculation or bandwidth bottleneck; Step 3.2: Data transmission optimization and resource allocation; under the premise of meeting resource constraints, multiple allocation attempts are made to on-chip cache and DSP computing unit, the data transmission amount of each configuration scheme is calculated, and the optimal configuration scheme is selected to balance storage and calculation; In the case that Loop-3 and Loop-4 have been tiled, firstly, it is judged whether the condition is met and ; if the condition is met, the pixel buffer and the weight buffer are allocated in different proportions, respectively, the data transmission BW1 and BW2 under different configuration schemes are calculated, and the optimal configuration scheme is screened out; wherein, BW1 and BW2 respectively represent the total data transmission when Loop-3 and Loop-4 are taken as inner loops, which are used to compare the bandwidth consumption under different loop unrolling orders, so as to screen the optimal configuration scheme; represents the loop unrolling factor , represents the block configuration , represents the scale of the convolutional neural network ; represents the loop unrolling factor cannot exceed the block configuration , otherwise, the hardware parallelism is greater than the single block data amount, which will cause idle and waste; the block configuration cannot exceed the total scale of the convolutional neural network , otherwise, the block configuration loses its meaning; in the condition , represents the time required for transmitting the feature block into the on-chip cache, is the data processing time under the current design; The calculation formula is as follows: wherein, represents the transmission frequency of the off-chip input data; W represents the data transmission bit width of the off-chip bus; , , respectively represent the block size of the width of the input feature map, the height of the input feature map, and the channel number of the input feature map; , , respectively represent the block size of the width of the convolution kernel, the height of the convolution kernel, and the channel number of the output feature map, used to describe the size of the data block that can be processed at a time in the on-chip cache. Step 3.3: Strategy selection and loop unrolling: efficiency comparison is made on the candidate schemes, the optimal loop unrolling sequence is finally determined, and the result is output as the hardware configuration to realize efficient execution of convolution operation and data transmission optimization; The configuration scheme with the smallest data transmission amount is selected, and BW1 and BW2 are compared: if BW1 is greater than BW2, Loop-3 is selected as the inner loop; otherwise, Loop-4 is selected as the inner loop; when Loop-3 is selected as the inner loop, the weight groups of each divided input pixel block are read in sequence to improve the utilization rate of on-chip pixel buffer and weight buffer; when Loop-4 is selected as the inner loop, the accelerator system loads pixels to the pixel buffer and the weight buffer in sequence to support parallel reading of weight groups; finally, according to the evaluation result, the optimal loop unrolling sequence and the corresponding cache management strategy are output to realize efficient execution of convolution operation and optimal configuration of data transmission.
Citation Information
Patent Citations
Design and implementation method of general convolution operation accelerator architecture based on loop optimization technology
CN118332239A
Intelligent customer service method for adaptively adjusting interaction strategy by using reinforcement learning
CN119311805A
Cyclic partitioning and resource allocation method for reducing CPU empty equations of convolutional neural network of embedded device
CN119988033A
Reinforcement learning method
EP3543918A1
Method and apparatus for state-adaptive reinforcement learning
WO2024212212A1