DDPG reinforcement learning-based FPGA data mining resource dynamic scheduling method and system
Patent Information
- Application Number
- CN202611037800.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-18
AI Technical Summary
[0007]为此,本发明实施例提供一种基于DDPG强化学习的FPGA数据挖掘资源动态调度方法及系统,以解决现有技术因任务动态多变与资源约束复杂导致的资源分配不合理、利用率低下及重配置频繁的技术问题
[0048] This invention collects real-time FPGA resource status and data mining task characteristics, quantizing them into a multi-dimensional state vector as input to the DDPG network. A lightweight Actor-Critic network is designed, where the Actor outputs a resource allocation strategy, and the Critic evaluates long-term value based on the iterative characteristics of data mining. Real-time monitoring of task execution progress and resource status is performed, and a reward is generated and fed back to the DDPG based on multiple objectives: early task completion, efficient resource utilization, and reduced reconfiguration. The strategy output by the Actor is mapped to the hardware layer via the FPGA resource manager, allocating logic units for parallel tasks, BRAM for iterative data, and bus bandwidth for high-bandwidth tasks, while also performing conflict detection. Each cycle, the state, action, and reward are stored in an experience pool. The Critic is sampled and trained to minimize the value error, and the Actor is updated, with a soft update to the target network. A new strategy is generated upon encountering a new task or resource mutation. This invention can adaptively adjust resource allocation, improve FPGA resource utilization, reduce task time and reconfiguration frequency, and effectively cope with the dynamic changes in data mining tasks.
Smart Images

Figure CN122777313A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of FPGA resource dynamic scheduling and data mining technology, specifically to an FPGA data mining resource dynamic scheduling method and system based on DDPG reinforcement learning. Background Technology
[0002] FPGA (Field-Programmable Gate Array), as a programmable logic device, combines the high efficiency of hardware with the reconfigurability of software, and has been widely deployed in computationally intensive applications such as data mining, artificial intelligence inference, and edge computing. FPGAs offer advantages such as low latency, high parallelism, and reconfigurability, allowing for flexible adjustments to the hardware structure to improve computational efficiency based on different tasks. However, the on-chip resources of FPGAs (such as logic units (LUTs), block RAM (BRAM), digital signal processing units (DSPs), and bus bandwidth) are relatively limited. In complex and ever-changing data mining scenarios, efficiently scheduling these scarce resources to ensure the parallel execution of multiple tasks has become a key challenge in the FPGA application field.
[0003] Extensive research has been conducted in academia and industry regarding FPGA resource scheduling. From a technical perspective, existing dynamic task scheduling methods for FPGA hardware can be mainly categorized as follows:
[0004] (1) Static scheduling method. Static scheduling determines all resource allocation schemes before the task starts execution, without considering resource status changes and task characteristic differences during task execution. This type of method is simple to implement and has low control overhead, but it has poor flexibility and is difficult to adapt to the diverse, dynamic, and unpredictable arrival time characteristics of data mining tasks. In scenarios with large fluctuations in task load, it is easy to lead to low resource utilization.
[0005] (2) Scheduling methods based on heuristic rules. These methods allocate resources based on preset heuristic rules (such as first-come, first-served, shortest job first, etc.). In recent years, researchers have further proposed scheduling strategies based on evolutionary computation such as genetic algorithms and particle swarm clustering. However, heuristic rules often lack adaptability when faced with new task types and resource conditions, and rule updates are lagging; evolutionary computation methods have problems such as slow convergence speed and easy getting trapped in local optima, making it difficult to meet the real-time requirements of FPGA data mining scenarios.
[0006] (3) Scheduling methods based on traditional reinforcement learning. In recent years, reinforcement learning has attracted attention for its ability to learn optimization strategies autonomously in dynamic environments. For example, DPUConfig proposed a runtime management framework based on a customized reinforcement learning agent, which dynamically selects the optimal DPU configuration through real-time telemetry data; ReinConfig used deep reinforcement learning to realize concurrent scheduling and asynchronous learning on the edge platform of FPGA multi-accelerator. In addition, some studies have proposed a high-level synthesis scheduling method for FPGA based on deep reinforcement learning, which makes scheduling decisions through self-learning. However, most of the above methods are geared towards specific scenarios such as deep learning inference or high-level synthesis, and their network structures are relatively complex and training overhead is large. When applied to tasks with obvious iterative characteristics such as data mining, traditional reinforcement learning methods still have shortcomings in the following aspects: First, the state space design does not fully reflect the joint constraints of multiple types of FPGA resources (logic units, storage, bandwidth); second, the reward function design is mostly oriented towards a single optimization objective, lacking a comprehensive trade-off between multiple objectives such as early task completion, efficient resource utilization, and reduced reconfiguration times; third, the impact of the iterative characteristics of data mining tasks on long-term value assessment is not fully considered. Summary of the Invention
[0007] To address this, this invention provides a method and system for dynamic scheduling of FPGA data mining resources based on DDPG reinforcement learning, in order to solve the technical problems of unreasonable resource allocation, low utilization and frequent reconfiguration caused by the dynamic and complex resource constraints of existing technologies.
[0008] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0009] According to a first aspect of the present invention, a dynamic scheduling method for FPGA data mining resources based on DDPG reinforcement learning is provided, the method comprising:
[0010] The real-time resource status of the FPGA and the characteristics of the data mining task are collected and quantized into a multi-dimensional state vector as the input of the DDPG network. The real-time resource status includes logic unit utilization, on-chip memory utilization and bus bandwidth utilization. The task characteristics include task type, data volume and computational complexity.
[0011] Design a lightweight DDPGActor-Critic network. The Actor network outputs a resource allocation strategy, which includes the allocation ratio of logical units, the allocation ratio of storage blocks, and the allocation ratio of bandwidth. The Critic network designs a long-term value evaluation function for the iterative characteristics of data mining tasks, and evaluates the long-term impact of the strategy on task time, resource utilization, and reconfiguration frequency.
[0012] Real-time monitoring of task execution progress and FPGA resource status; based on multiple objectives such as early task completion, efficient resource utilization, and reduced reconfiguration, weighted instant rewards are generated and fed back to the DDPG network.
[0013] The Actor network outputs a resource allocation strategy based on the current state, and maps the strategy to the hardware layer through the FPGA resource manager: allocating logic units for the parallel part of the task, allocating block storage (BRAM) for the intermediate data of the iteration, allocating bus bandwidth for high-bandwidth tasks, and performing resource conflict detection to avoid resource contention.
[0014] In each learning cycle, the current state, action, and reward are stored as experience samples in the experience pool. Batch samples are sampled from the experience pool to train the Critic network to minimize the value error. The Actor network is updated based on the evaluation of the Critic network to maximize the expected cumulative reward. The target network is updated using a soft update method. When a new task is detected or the FPGA resource state changes abruptly, the resource allocation strategy is regenerated.
[0015] Furthermore, the quantization into a multidimensional state vector specifically includes:
[0016] Calculate the utilization rate of logic units, on-chip memory, and bus bandwidth respectively, and collect the task type encoding, data volume, and computational complexity of data mining tasks to form an initial state vector.
[0017] The initial state vector is normalized using historical datasets to obtain a normalized state vector, and a resource constraint violation index is calculated. The resource constraint violation index is the sum of the amount by which the current utilization rate of each type of resource exceeds its respective preset threshold.
[0018] The normalized state vector is combined with the resource constraint violation index to obtain the final state vector after expanding the dimensions, which is then used as the input to the DDPG network.
[0019] Furthermore, the Actor network adopts a lightweight fully connected structure, the input layer dimension matches the state vector dimension, the hidden layer contains 128 neurons, and the output layer uses the Softmax activation function to generate a resource allocation strategy vector. The resource allocation strategy vector includes the logical unit allocation ratio, the storage block allocation ratio, and the bandwidth allocation ratio, and the sum of the three is normalized to 1.
[0020] The weight parameters of the Actor network are quantized using INT8 to reduce storage overhead, and matrix multiplication calculation is accelerated by the FPGA's digital signal processing (DSP) block. The weight matrix is stored in BRAM to optimize access latency, and the ReLU function is selected as the activation function.
[0021] Furthermore, the Critic network is designed as a dual-input structure, receiving the current state and the action output by the Actor network respectively. After feature fusion through the hidden layer, it outputs a Q-value, which is used to evaluate the long-term value of the current strategy.
[0022] The long-term value assessment is based on cumulative discount reward calculation, where the immediate reward at each time step is defined as a weighted combination of negative task time, positive resource utilization, and negative reconfiguration times. The discount factor is used to adjust the degree of influence of future rewards on the current decision.
[0023] Furthermore, the weighted generation of instant rewards specifically includes:
[0024] The reward for completing the task ahead of schedule is calculated by multiplying the ratio of the iteration increment at the current time step to the total estimated number of iterations by the first weighting coefficient.
[0025] The reward for efficient utilization of computing resources is calculated by multiplying the current resource utilization rate by the second weighting coefficient and subtracting the bus congestion quantification value by the third weighting coefficient.
[0026] The reconfiguration reduction penalty is calculated by multiplying the negative value of the binary indicator variable indicating whether reconfiguration has occurred by the fourth weighting coefficient.
[0027] The total reward value is obtained by adding the three reward components mentioned above. The weight coefficients are determined by grid search or empirical tuning. The total reward value is limited to a preset range to facilitate DDPG convergence.
[0028] Furthermore, the allocation of logic units for the parallel task portion specifically includes: identifying a set of parallel tasks, obtaining the number of logic units required for each parallel task, calculating the total demand for all parallel tasks, ensuring that it does not exceed the total number of available logic units on the FPGA, taking the smaller value between the demand and the available quantity as the final allocation amount, and mapping the tasks to the free lookup table (LUT) area through a layout algorithm to avoid hotspots.
[0029] The specific steps for allocating BRAM for iterative intermediate data include: extracting the data block size from the action vector, calculating the required number of BRAM blocks based on the capacity of each BRAM block, prioritizing the allocation of contiguous address blocks to reduce memory fragmentation, and verifying that the required number of blocks does not exceed the remaining BRAM quantity.
[0030] The specific steps for allocating bus bandwidth for high-bandwidth tasks include: identifying a set of high-bandwidth tasks, obtaining the bandwidth requirements of each high-bandwidth task, allocating the total bus bandwidth according to the proportion of each task's requirements to the total requirements of all high-bandwidth tasks, and using a weighted fair queue scheduling algorithm to prioritize high-demand tasks.
[0031] Furthermore, the execution resource conflict detection specifically includes:
[0032] Check whether the address ranges of logic units overlap, whether the BRAM block addresses overlap, and whether the bus channels overlap, to ensure that the allocation ranges of various resources do not overlap.
[0033] When a conflict is detected, a bipartite graph matching algorithm is used to resolve the conflict in resource allocation, ensuring that there is no competition for any type of resource before completing the hardware mapping.
[0034] Furthermore, the step of sampling batches of samples from the experience pool to train the Critic network to minimize the value error specifically includes:
[0035] A preset number of samples are uniformly and randomly sampled from the experience replay buffer. For each sample, the updated target value is calculated based on the current instant reward, discount factor, and target Q value output by the target network.
[0036] The loss function of the Critic network is constructed based on the mean square error between the current Q value and the updated target value. The gradient is calculated through backpropagation and the Critic network parameters are updated using an adaptive moment estimator.
[0037] The process of updating the Actor network to maximize the expected cumulative return specifically includes: calculating the policy gradient estimate based on the Q-value gradient of the Critic network for the action and the policy gradient of the Actor network, and updating the Actor network parameters through gradient ascent;
[0038] The soft update method specifically involves updating the Critic target network parameters and the Actor target network parameters by weighting them according to preset soft update coefficients, so that the target network parameters slowly track the online network parameters.
[0039] According to a second aspect of the present invention, a dynamic scheduling system for FPGA data mining resources based on DDPG reinforcement learning is provided, the system comprising:
[0040] The state awareness module is used to collect the real-time resource status of the FPGA and the characteristics of data mining tasks, and quantize them into a multi-dimensional state vector.
[0041] The lightweight DDPG decision module includes an Actor network and a Critic network. The Actor network is used to output a resource allocation strategy based on the current state vector, and the Critic network is used to evaluate the long-term value of the resource allocation strategy.
[0042] The reward generation module is used to monitor the task execution progress and FPGA resource status in real time, generate an instant reward based on multi-objective weighting, and feed it back to the lightweight DDPG decision module.
[0043] The hardware resource mapping module is used to map the resource allocation strategy output by the Actor network to the FPGA hardware layer. The mapping includes allocating logic units for parallel tasks, allocating BRAM for iterative intermediate data, allocating bus bandwidth for high-bandwidth tasks, and performing conflict detection and conflict resolution.
[0044] The experience replay and training module is used to store experience samples, sample batch samples to train the Critic and Actor networks, perform soft updates on the target network, and detect environmental changes and trigger policy regeneration.
[0045] Furthermore, the state awareness module collects the logic unit utilization and on-chip memory utilization in real time through the FPGA's built-in monitoring interface or a dedicated hardware counter, and obtains the bus bandwidth utilization through sampling via an AXI bus monitor. The sampling frequency is at the millisecond level.
[0046] The hardware resource mapping module has a built-in conflict detection submodule and a conflict resolution submodule. The conflict detection submodule is used to check whether there is overlap between the logic unit address range, BRAM block address and bus channel. The conflict resolution submodule uses a bipartite graph matching algorithm to reallocate overlapping resources.
[0047] The embodiments of the present invention have the following advantages:
[0048] This invention collects real-time FPGA resource status and data mining task characteristics, quantizing them into a multi-dimensional state vector as input to the DDPG network. A lightweight Actor-Critic network is designed, where the Actor outputs a resource allocation strategy, and the Critic evaluates long-term value based on the iterative characteristics of data mining. Real-time monitoring of task execution progress and resource status is performed, and a reward is generated and fed back to the DDPG based on multiple objectives: early task completion, efficient resource utilization, and reduced reconfiguration. The strategy output by the Actor is mapped to the hardware layer via the FPGA resource manager, allocating logic units for parallel tasks, BRAM for iterative data, and bus bandwidth for high-bandwidth tasks, while also performing conflict detection. Each cycle, the state, action, and reward are stored in an experience pool. The Critic is sampled and trained to minimize the value error, and the Actor is updated, with a soft update to the target network. A new strategy is generated upon encountering a new task or resource mutation. This invention can adaptively adjust resource allocation, improve FPGA resource utilization, reduce task time and reconfiguration frequency, and effectively cope with the dynamic changes in data mining tasks. Attached Figure Description
[0049] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0050] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0051] Figure 1 A schematic diagram of the logical structure of an FPGA data mining resource dynamic scheduling system based on DDPG reinforcement learning provided in an embodiment of the present invention;
[0052] Figure 2 This is a flowchart illustrating a dynamic scheduling method for FPGA data mining resources based on DDPG reinforcement learning, provided in an embodiment of the present invention. Detailed Implementation
[0053] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] refer to Figure 1 This invention discloses an FPGA data mining resource dynamic scheduling system based on DDPG reinforcement learning. The system includes: a state awareness module 1; a lightweight DDPG decision module 2; a reward generation module 3; a hardware resource mapping module 4; and an experience playback and training module 5.
[0055] Corresponding to the aforementioned FPGA data mining resource dynamic scheduling system based on DDPG reinforcement learning, this invention also discloses an FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning. The following details the FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning disclosed in this invention, in conjunction with the aforementioned FPGA data mining resource dynamic scheduling system based on DDPG reinforcement learning.
[0056] refer to Figure 2 This invention discloses a dynamic scheduling method for FPGA data mining resources based on DDPG reinforcement learning, comprising:
[0057] Step 1: Collect real-time FPGA resource status (logic units, on-chip memory, bus bandwidth utilization) and data mining task characteristics (task type, data volume, computational complexity), quantize them into a multi-dimensional state vector as input to the DDPG network, accurately reflecting FPGA resource constraints. The specific implementation method for this step is as follows:
[0058] Resource status is collected in real time via the FPGA's built-in monitoring interface or a dedicated hardware counter. Logic unit utilization. Defined as the ratio of the number of currently used lookup tables (LUTs) to the total number of available LUTs, calculated using the following formula:
[0059]
[0060] in Obtained by reading the configuration register. From device specifications. On-chip memory (e.g., BRAM) utilization. The calculation is the used storage bit width divided by the total storage capacity, expressed as:
[0061]
[0062] Retrieved from the memory controller status register. Bus bandwidth utilization. Based on the ratio of actual data transmission rate to peak bandwidth, the formula is:
[0063]
[0064] Obtained through sampling by the AXI bus monitor. Depending on the clock frequency and bit width, these metrics are updated every millisecond to ensure real-time performance.
[0065] Collecting data mining task features involves parsing task configuration files or runtime API calls. Task Types Integer encoding mapping is used: 0 represents classification, 1 represents clustering, and 2 represents regression, with values directly assigned based on task metadata. Data volume. The calculation formula, derived from the size of the input dataset and measured in bytes, is as follows: Computational complexity Quantized into floating-point operands (FLOPs), based on the task algorithm model, such as for matrix multiplication tasks.
[0066]
[0067] Historical statistics are used for calibration.
[0068] Quantized into a multidimensional state vector, initial vector Combining resource status and task characteristics:
[0069]
[0070] To accurately reflect resource constraints, normalization and constraint indicators are introduced. Normalization uses historical datasets to calculate the mean and standard deviation:
[0071]
[0072]
[0073] in This is the number of samples. The normalized vector is:
[0074]
[0075] Constraints on Violation of Indicators Defined as resource exceeding threshold penalty:
[0076]
[0077] threshold According to FPGA specifications (e.g.) Final state vector:
[0078]
[0079] The dimensions have been expanded to 7.
[0080] This vector serves as input to the DDPG network, directly fed into the fully connected layers of the actor and critic networks. During training, Ensuring that resource constraints are encoded as continuous values facilitates the policy network's optimization of resource allocation decisions and avoids the loss of accuracy due to discretization.
[0081] Step 2: Design a lightweight DDPGActor-Critic network: The Actor outputs a resource allocation strategy (logic unit, memory block, bandwidth ratio) adapted to hardware such as FPGA, DSP, and BRAM; the Critic evaluates the value of the strategy in terms of task time, resource utilization, and reconfiguration times, and designs a long-term value assessment for the iterative nature of data mining. The specific implementation method for this step is as follows:
[0082] The Actor network employs a lightweight fully connected structure. Its input dimension matches the system state vector (such as task queue length, current resource utilization, and data mining iteration progress), and the output layer uses the softmax activation function to generate a resource allocation strategy vector. ,in This represents the normalized proportion of logic units, memory blocks, and bandwidth. The network is limited to two layers (128 hidden units and 3 output units). Weight parameters are quantized using INT8 to reduce storage overhead, and matrix multiplication calculations are accelerated using the FPGA's DSP block. BRAM is used to store the weight matrix to optimize access latency. The ReLU activation function is chosen to avoid saturation and adapt to hardware resource constraints. The Critic network is designed with a dual-input structure, receiving the state... and Actor output The hidden layer (256 units) outputs the Q-value after fusing features. Evaluate the long-term value of the strategy; integrate the value function to determine the task time. (Negative rewards), resource utilization rate (Positive rewards) and number of reconfigurations (Negative reward) To address the iterative nature of data mining, a discount factor is introduced. Calculate cumulative rewards:
[0083]
[0084] The instant reward is defined as a weighted combination:
[0085]
[0086] Weighting coefficient By optimizing the hyperparameters to balance the metrics, we ensure that the Critic-predicted Q-value closely approximates the target value. The training process utilizes the DDPG algorithm framework, and the experience replay cache stores the samples. Calculate the target Q value after sampling batch data:
[0087]
[0088] Update the Critic parameters to minimize the mean squared error loss:
[0089]
[0090] The Actor parameters are updated via policy gradients to maximize the Q-value:
[0091]
[0092] Target network soft update usage ( To ensure training stability, network lightweighting is achieved by reducing parameter size (e.g., compressing the number of hidden layer units) and FPGA-friendly optimizations (e.g., fixed-point computation), while monitoring resource utilization. A dynamic adjustment strategy is adopted.
[0093] Step 3: Monitor task execution progress (iteration count, intermediate results) and FPGA resource status (utilization, bus congestion) in real time. Based on the goals of early task completion, efficient resource utilization, and reduced reconfiguration, generate a weighted reward and feed it back to the DDPG. The specific implementation method for this step is as follows:
[0094] The real-time monitoring mechanism is implemented through embedded sensors and system APIs. Task execution progress is automatically collected by an iteration counter, which updates the progress after each calculation iteration. Value, of which This indicates the current time step; intermediate results such as output precision or convergence error are read through a predefined interface. This reflects task quality. FPGA resource status monitoring uses hardware performance counters: utilization rate. Bus congestion is calculated as a percentage of available logic units. Based on packet latency and packet loss rate quantization, the calculation formula is as follows:
[0095]
[0096] in This is the current bus latency. This is the maximum tolerance threshold. All data is sampled at millisecond-level frequencies to ensure real-time performance and is transmitted to the central processing unit via a buffer queue.
[0097] Define the Reward component to map the target. Early completion reward. Incentivize rapid progress based on iterative increments Iteration with the total estimate The ratio is given by the formula:
[0098]
[0099] in This is a weighting factor to ensure earlier completion yields higher returns. Resource efficiency utilization rewards. Combined with occupancy rate and congestion ,formula:
[0100]
[0101] and Strive to balance high utilization with low congestion. Reconfiguration reduces penalties. Event-driven, The formula is used to indicate whether a reconfiguration has occurred (0 or 1) in binary format. , Frequent weighting penalties necessitate reconfiguration.
[0102] Weighted reward generation achieves total return through linear combination. The components are normalized and then weighted and summed:
[0103]
[0104] Substituting, we get:
[0105]
[0106] The weighting coefficients are optimized through grid search or empirical tuning, for example, by setting... , , , This ensures tasks are prioritized and managed in advance, while optimizing resource utilization and minimizing reconfiguration. (After calculation) The value range is between [-1, 1], which facilitates the convergence of DDPG.
[0107] Feedback is integrated into the DDPG within the reinforcement learning loop. At each time step... calculate Subsequently, the Q-value function of the Critic network, which is input to the DDPG as an immediate reward, is updated. Action selection is based on the Actor network output, and policy optimization is achieved through empirical replay and gradient descent. This is expressed as the Critic loss function:
[0108]
[0109] in As a discount factor, The state vector incorporates monitoring data to ensure that the reward-guided strategy is optimized for multiple objectives. The entire process is executed in parallel on the FPGA coprocessor, reducing latency.
[0110] Step 4: Based on the status output strategy, the Actor maps hardware through the FPGA resource manager: allocating logic units for parallel tasks, BRAM for iterative intermediate data, and bus bandwidth for high-bandwidth tasks to avoid resource conflicts. The specific implementation method for this step is as follows:
[0111] The Actor module receives the current system state vector. This state includes information such as task queue, resource utilization, and real-time performance metrics. The Actor outputs actions based on the policy network. The policy function is defined as follows:
[0112]
[0113] in These are network weight parameters. It is a state-action feature mapping function, and the probability distribution is ensured through a softmax operation. (Action selection...) Then, it is passed to the FPGA resource manager for parsing. The resource manager decodes the data. Specific resource requirement parameters include parallel task identifier, iteration data size, and bandwidth requirement threshold.
[0114] Allocate logical units to the parallel parts of the tasks and identify sets of parallel tasks. Each task Number of logic units required Calculate total demand during allocation. And ensure that it does not exceed the available FPGA resources. To achieve optimized allocation:
[0115]
[0116] At the same time, the layout algorithm maps tasks to idle LUT areas to avoid hot spots.
[0117] Allocate BRAM for iterative intermediate data, based on action. Extract data block size Each BRAM block has a capacity of [missing information]. Calculate the number of blocks required The allocation process prioritizes contiguous address blocks to reduce memory fragmentation and verifies... ,in This refers to the remaining BRAM quantity, ensuring efficient data storage.
[0118] Allocate bus bandwidth to high-bandwidth tasks and identify high-bandwidth task sets. Each task Bandwidth requirements The total bus bandwidth is Allocation ratio:
[0119]
[0120] satisfy It also implements a weighted fair queue scheduling algorithm to prioritize high-demand tasks.
[0121] To avoid resource conflicts, a conflict detection mechanism is implemented to check for overlaps in logical unit address ranges, BRAM block addresses, and bus channels. Conflict constraints are defined as follows:
[0122]
[0123] A bipartite graph matching algorithm is used to resolve conflicts, ensuring no resource contention. The entire process involves real-time monitoring and feedback to dynamically adjust strategies, thereby improving hardware utilization and system throughput.
[0124] Step 5: In each cycle, store the state, action, and reward into the experience pool, sample and train the Critic to minimize the value error, update the Actor to maximize the Critic evaluation, and softly update the target network; if a new task or resource mutation occurs, repeat steps 1-4 to regenerate the policy. The specific implementation method of this step is as follows:
[0125] At the end of each learning cycle, the agent bases its decisions on the current environment state. Select and execute an action Receive instant rewards and transition to the new state. These elements form a quadruple. Stored in a fixed-capacity experience replay buffer. The buffer uses a first-in, first-out (FIFO) strategy to prevent outdated data from affecting learning stability. The buffer capacity is... The settings should be adjusted according to the complexity of the task to avoid memory overflow.
[0126] From the experience pool A batch of samples is uniformly and randomly sampled from the middle, and the batch size is... Typically 128 or 256. For each sample Calculate the target Q value:
[0127]
[0128] in It is a discount factor (e.g., 0.99). and Represents the target network. Critic network. The loss function is defined as:
[0129]
[0130] The gradient is calculated using backpropagation. And update the Critic parameters using an optimizer (such as Adam). To minimize the mean square error.
[0131] Critic-based evaluation update of Actor policy network To maximize the expected cumulative return, the policy gradient is approximated as:
[0132]
[0133] This gradient indicates the direction of Q-value improvement; gradient ascent is used to optimize the Actor parameters. Learning rate Careful configuration is required to avoid strategy oscillations.
[0134] Perform a soft update of the target network to ensure stability. Using the Polyak averaging strategy: Critic target network parameters are updated as follows:
[0135]
[0136] Meanwhile, the Actor target network is updated as follows:
[0137]
[0138] Among them, the soft update coefficient The value is typically 0.995. This operation is performed synchronously after each Critic / Actor update to keep the target network slowly tracking the online network.
[0139] At the start of each cycle, environmental changes are detected. External signals (such as sudden changes in task ID or resource utilization) are monitored to determine if a new task or resource change has occurred. If a change is detected, the experience pool is cleared. Reinitialize policy network parameters and Then, return to the initial training phase and re-execute the policy generation process. The change threshold can be based on the resource difference rate:
[0140]
[0141] Settings, for example Triggered a reset.
[0142] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A dynamic resource scheduling method for FPGA data mining based on DDPG reinforcement learning, characterized in that, The method includes: The real-time resource status of the FPGA and the characteristics of the data mining task are collected and quantized into a multi-dimensional state vector as the input of the DDPG network. The real-time resource status includes logic unit utilization, on-chip memory utilization and bus bandwidth utilization. The task characteristics include task type, data volume and computational complexity. Design a lightweight DDPGActor-Critic network. The Actor network outputs a resource allocation strategy, which includes the allocation ratio of logical units, the allocation ratio of storage blocks, and the allocation ratio of bandwidth. The Critic network designs a long-term value evaluation function for the iterative characteristics of data mining tasks, and evaluates the long-term impact of the strategy on task time, resource utilization, and reconfiguration frequency. Real-time monitoring of task execution progress and FPGA resource status; based on multiple objectives such as early task completion, efficient resource utilization, and reduced reconfiguration, weighted instant rewards are generated and fed back to the DDPG network. The Actor network outputs a resource allocation strategy based on the current state, and maps the strategy to the hardware layer through the FPGA resource manager: allocating logic units for the parallel part of the task, allocating block storage (BRAM) for the intermediate data of the iteration, allocating bus bandwidth for high-bandwidth tasks, and performing resource conflict detection to avoid resource contention. In each learning cycle, the current state, action, and reward are stored as experience samples in the experience pool. Batch samples are sampled from the experience pool to train the Critic network to minimize the value error. The Actor network is updated based on the evaluation of the Critic network to maximize the expected cumulative reward. The target network is updated using a soft update method. When a new task is detected or the FPGA resource state changes abruptly, the resource allocation strategy is regenerated.
2. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The quantization into a multidimensional state vector specifically includes: Calculate the utilization rate of logic units, on-chip memory, and bus bandwidth respectively, and collect the task type encoding, data volume, and computational complexity of data mining tasks to form an initial state vector. The initial state vector is normalized using historical datasets to obtain a normalized state vector, and a resource constraint violation index is calculated. The resource constraint violation index is the sum of the amount by which the current utilization rate of each type of resource exceeds its respective preset threshold. The normalized state vector is combined with the resource constraint violation index to obtain the final state vector after expanding the dimensions, which is then used as the input to the DDPG network.
3. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The Actor network adopts a lightweight fully connected structure. The dimension of the input layer matches the dimension of the state vector. The hidden layer contains 128 neurons. The output layer uses the Softmax activation function to generate a resource allocation strategy vector. The resource allocation strategy vector includes the allocation ratio of logical units, the allocation ratio of storage blocks, and the allocation ratio of bandwidth, and the sum of the three is normalized to 1. The weight parameters of the Actor network are quantized using INT8 to reduce storage overhead, and matrix multiplication calculation is accelerated by the FPGA's digital signal processing (DSP) block. The weight matrix is stored in BRAM to optimize access latency, and the ReLU function is selected as the activation function.
4. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The Critic network is designed with a dual-input structure, receiving the current state and the action output by the Actor network respectively. After feature fusion through the hidden layer, it outputs a Q-value, which is used to evaluate the long-term value of the current strategy. The long-term value assessment is based on cumulative discount reward calculation, where the immediate reward at each time step is defined as a weighted combination of negative task time, positive resource utilization, and negative reconfiguration times. The discount factor is used to adjust the degree of influence of future rewards on the current decision.
5. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The weighted generation of instant rewards specifically includes: The reward for completing the task ahead of schedule is calculated by multiplying the ratio of the iteration increment at the current time step to the total estimated number of iterations by the first weighting coefficient. The reward for efficient utilization of computing resources is calculated by multiplying the current resource utilization rate by the second weighting coefficient and subtracting the bus congestion quantification value by the third weighting coefficient. The reconfiguration reduction penalty is calculated by multiplying the negative value of the binary indicator variable indicating whether reconfiguration has occurred by the fourth weighting coefficient. The total reward value is obtained by adding the three reward components mentioned above. The weight coefficients are determined by grid search or empirical tuning. The total reward value is limited to a preset range to facilitate DDPG convergence.
6. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The specific steps for allocating logic units for the parallel part of the task include: identifying the set of parallel tasks, obtaining the number of logic units required for each parallel task, calculating the total demand of all parallel tasks and ensuring that it does not exceed the total number of available logic units on the FPGA, taking the smaller value between the demand and the available quantity as the final allocation amount, and mapping the tasks to the free lookup table (LUT) area through a layout algorithm to avoid hotspots. The specific steps for allocating BRAM for iterative intermediate data include: extracting the data block size from the action vector, calculating the required number of BRAM blocks based on the capacity of each BRAM block, prioritizing the allocation of contiguous address blocks to reduce memory fragmentation, and verifying that the required number of blocks does not exceed the remaining BRAM quantity. The specific steps for allocating bus bandwidth for high-bandwidth tasks include: identifying a set of high-bandwidth tasks, obtaining the bandwidth requirements of each high-bandwidth task, allocating the total bus bandwidth according to the proportion of each task's requirements to the total requirements of all high-bandwidth tasks, and using a weighted fair queue scheduling algorithm to prioritize high-demand tasks.
7. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The execution resource conflict detection specifically includes: Check whether the address ranges of logic units overlap, whether the BRAM block addresses overlap, and whether the bus channels overlap, to ensure that the allocation ranges of various resources do not overlap. When a conflict is detected, a bipartite graph matching algorithm is used to resolve the conflict in resource allocation, ensuring that there is no competition for any type of resource before completing the hardware mapping.
8. The FPGA data mining resource dynamic scheduling method based on DDPG reinforcement learning as described in claim 1, characterized in that, The step of sampling batches of samples from the experience pool to train the Critic network to minimize the value error specifically includes: A preset number of samples are uniformly and randomly sampled from the experience replay buffer. For each sample, the updated target value is calculated based on the current instant reward, discount factor, and target Q value output by the target network. The loss function of the Critic network is constructed based on the mean square error between the current Q value and the updated target value. The gradient is calculated through backpropagation and the Critic network parameters are updated using an adaptive moment estimator. The process of updating the Actor network to maximize the expected cumulative return specifically includes: calculating the policy gradient estimate based on the Q-value gradient of the Critic network for the action and the policy gradient of the Actor network, and updating the Actor network parameters through gradient ascent; The soft update method specifically involves updating the Critic target network parameters and the Actor target network parameters by weighting them according to preset soft update coefficients, so that the target network parameters slowly track the online network parameters.
9. A dynamic scheduling system for FPGA data mining resources based on DDPG reinforcement learning, characterized in that, The system includes: The state awareness module is used to collect the real-time resource status of the FPGA and the characteristics of data mining tasks, and quantize them into a multi-dimensional state vector. The lightweight DDPG decision module includes an Actor network and a Critic network. The Actor network is used to output a resource allocation strategy based on the current state vector, and the Critic network is used to evaluate the long-term value of the resource allocation strategy. The reward generation module is used to monitor the task execution progress and FPGA resource status in real time, generate an instant reward based on multi-objective weighting, and feed it back to the lightweight DDPG decision module. The hardware resource mapping module is used to map the resource allocation strategy output by the Actor network to the FPGA hardware layer. The mapping includes allocating logic units for parallel tasks, allocating BRAM for iterative intermediate data, allocating bus bandwidth for high-bandwidth tasks, and performing conflict detection and conflict resolution. The experience replay and training module is used to store experience samples, sample batch samples to train the Critic and Actor networks, perform soft updates on the target network, and detect environmental changes and trigger policy regeneration.
10. The FPGA data mining resource dynamic scheduling system based on DDPG reinforcement learning as described in claim 9, characterized in that, The state awareness module collects the logic unit utilization and on-chip memory utilization in real time through the FPGA built-in monitoring interface or a dedicated hardware counter, and obtains the bus bandwidth utilization through the AXI bus monitor. The sampling frequency is at the millisecond level. The hardware resource mapping module has a built-in conflict detection submodule and a conflict resolution submodule. The conflict detection submodule is used to check whether there is overlap between the logic unit address range, BRAM block address and bus channel. The conflict resolution submodule uses a bipartite graph matching algorithm to reallocate overlapping resources.