CGRA multi-task dynamic resource allocation method and hardware circuit

By introducing modules such as task attribute tables, management modules and finite state machines in CGRA, dynamic resource allocation over hundreds of clock cycles is realized, which solves the problem of low resource utilization in CGRA multi-task scenarios, and improves throughput and processing unit utilization.

CN120276843APending Publication Date: 2025-07-08HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345450.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing CGRA multi-task dynamic resource allocation method is difficult to effectively manage randomly created and destroyed tasks in the face of complex modern computing scenarios, resulting in low resource utilization and thus limiting the multi-task throughput.

Method used

A CGRA multi-task dynamic resource allocation method and hardware circuit are adopted to realize dynamic resource allocation through task attribute table, task management module, processing unit allocation quantity generation module, processing unit shape generation module and configuration preload module, combined with a finite state machine, dynamic resource allocation can be achieved, and resource allocation of tasks can be completed within hundreds of clock cycles.

Benefits of technology

It improves CGRA resource utilization, improves multi-task throughput, and reduces hardware latency and processing unit utilization, and is suitable for CGRA applications of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276843A_ABST
    Figure CN120276843A_ABST
Patent Text Reader

Abstract

The invention discloses a CGRA multi-task dynamic resource allocation method and a hardware circuit, and belongs to the technical field of dynamic resource allocation. The method specifically comprises the following steps: recording the attribute of each task by using a task attribute table; updating a task attribute table according to a central processing unit message, and generating a task queue and a Boolean value for triggering dynamic resource allocation; the number of processing units is calculated and allocated according to task priorities and the number of data flow diagram nodes; determining the shape of the distributed processing unit according to the number of the distributed processing units; pre-loading the configuration information into a configuration memory of the processing unit so as to match the shape of the allocated processing unit; and dynamic resource allocation is realized by using a finite-state machine. According to the hardware circuit, on the basis of an original CGRA circuit architecture, a dynamic resource distributor, a configuration transmission bus between processing units and configuration receiving circuits in the processing units are added. According to the method, dynamic resource changes caused by task creation and destruction can be flexibly handled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a resource allocation method and circuit, and more particularly to a CGRA multi-task dynamic resource allocation method and hardware circuit, belonging to the technical field of dynamic resource allocation. Background Art

[0002] In modern integrated circuit design, the growing demand for computing performance and energy efficiency has highlighted the limitations of traditional fixed-architecture processors. Coarse-grained reconfigurable arrays (CGRAs) have emerged to meet these needs. A CGRA combines the high performance of application-specific integrated circuits (ASICs) with the flexibility of field-programmable gate arrays (FPGAs). Compared with ASICs, which involve a complex process from design to tape-out, CGRAs stand out for their simplicity. It can be easily configured for various computing applications without repeating such a complex process. Compared with FPGAs, CGRAs have higher computing performance and energy efficiency. These characteristics make CGRAs have great application potential in fields such as artificial intelligence acceleration, digital signal processing, and Internet of Things edge computing.

[0003] To fully unleash the application potential of CGRAs in these fields, it has become imperative to understand the nature of concurrent computing tasks. In modern computing scenarios, executing multiple tasks simultaneously is not only a convenience but also a necessity. These tasks are characterized by a complex set of requirements, with constantly changing resource demands due to the random creation and termination of tasks brought about by unpredictable external environmental changes. For example, in autonomous driving, object detection and path planning tasks run simultaneously. New objects or traffic conditions create new tasks, and operations such as the user changing the destination also affect the tasks. In streaming media, video and audio processing occur simultaneously, and changes in the video quality requested by the user affect the resources. In Internet of Things monitoring, sensor-related tasks change randomly due to faults or user-adjusted parameters. Therefore, dynamic resource allocation in CGRAs is the key to improving resource utilization and multi-task throughput.

[0004] However, although CGRAs have great potential in multi-task scenarios, the existing research on their dynamic resource allocation is insufficient. Existing offline methods rely too much on the offline analysis of each individual task. This involves carefully analyzing the possible dynamic resource requirements and how different tasks are combined. For example, in a system with numerous built-in tasks, such as a comprehensive IoT edge device with hundreds of sensor-related tasks, this offline analysis becomes extremely labor-intensive. For modern fast-paced computing scenarios, this is simply impractical. As for online methods, they face their own set of problems. Some are too simple to manage the intense resource contention and release that occur during the random creation and destruction of multiple tasks. Other algorithms can only be applied to tasks where changes in input features are crucial, such as sparse matrix multiplication, which is sensitive to the number of zeros in the streaming input, limiting their scope of application.

[0005] In summary, both offline and online methods are difficult to face complex CGRA multi-task scenarios, resulting in low resource utilization, which in turn limits the multi-task throughput. Therefore, a CGRA multi-task dynamic resource allocation method and hardware circuit are proposed. Summary of the Invention

[0006] The present invention aims to solve the problem of dynamic resource changes caused by task creation or destruction, and thus proposes a CGRA multi-task dynamic resource allocation method and hardware circuit. It can complete dynamic resource allocation within hundreds of clock cycles, thereby improving the CGRA resource utilization rate and further enhancing the multi-task throughput.

[0007] To achieve the above objectives, the present invention adopts the following technical solutions:

[0008] A CGRA multi-task dynamic resource allocation method, which is implemented through the following steps:

[0009] Step 1: Use a task attribute table to record the attributes of each task;

[0010] Step 2: Through a task management module, update the task attribute table according to the central processor message, and generate a task queue and a boolean value that triggers dynamic resource allocation;

[0011] Step 3: When dynamic resource allocation is triggered, through a processing unit allocation quantity generation module, calculate and allocate the number of processing units according to the task priority and the number of data flow graph nodes;

[0012] Step 4: Through a processing unit shape generation module, determine the shape of the allocated processing units according to the number of allocated processing units;

[0013] Step 5: Preload the configuration information into the configuration memory of the processing unit by configuring the preloading module and the bus to match the allocated processing unit shape;

[0014] Step 6: Use a finite state machine to control the working order of the modules in Steps 2 - 5 to achieve dynamic resource allocation.

[0015] Furthermore, the task attribute table in Step 1 includes task number, number of data flow graph nodes, status, task priority, and the number of the processing unit assigned to each task.

[0016] Furthermore, the task attribute table in Step 1 is a RAM or a register file.

[0017] Furthermore, the processing unit allocation quantity generation module in Step 3 calculates the allocated number of processing units through weighted average, and uses the ratio obtained from the number of nodes of the task data flow graph after priority weighting or the total number of nodes of all task data flow graphs as the weight.

[0018] Furthermore, the processing unit shape generation module in Step 4 uses a coordinate - based heuristic method to allocate regular - shaped processing units to each task.

[0019] Furthermore, the configuration preloading module and bus design in Step 5 includes a dual - buffer configuration memory to hide the latency of reloading the configuration and transmits the configuration information to the processing unit through a specified bus.

[0020] A hardware circuit for implementing a CGRA multi - task dynamic resource allocation method includes: a CGRA circuit architecture, a dynamic resource allocator, a configuration transmission bus between processing units, and a configuration acceptance circuit inside the processing unit. The dynamic resource allocator is integrated in the CGRA circuit architecture and includes a task attribute table, a task management module, a processing unit allocation quantity generation module, a processing unit shape generation module, a finite state machine, and a configuration preloading module; the configuration transmission bus between processing units is used to transmit configuration information; the configuration acceptance circuit inside the processing unit is used to receive and store configuration information.

[0021] Furthermore, in the dynamic resource allocator,

[0022] The task attribute table is used to record the attributes of each task, including task number, number of data flow graph nodes, task status, priority, and the number of the allocated processing unit, and is accessed by other modules;

[0023] The task management module is used to maintain the task attribute table and generate corresponding outputs, perform corresponding operations on the task attribute table according to the central processor message, set the boolean value for triggering dynamic resource allocation, and extract the task queue;

[0024] The processing unit allocation quantity generation module calculates and allocates the quantity of processing units according to the task priority and the number of nodes in the data flow graph, and restricts it not to exceed the theoretical maximum value.

[0025] The processing unit shape generation module uses a coordinate-based heuristic method to allocate regular-shaped processing units to each task, so as to enrich the routing resources and improve the task throughput.

[0026] The configuration preloading module is used to send the configuration sent by the central processing unit to each processing unit via the bus after supplementing the correct processing unit number, and controls the configuration to be written into the buffer of the processing unit configuration memory.

[0027] The finite state machine controls the working order of the task management module, the processing unit allocation quantity generation module, the processing unit shape generation module and the configuration preloading module, triggers task management, resource allocation, configuration preloading and task execution in sequence, and realizes dynamic resource allocation.

[0028] Furthermore, the dynamic resource allocator is connected to the CGRA circuit architecture processing unit through the configuration transmission bus between the processing units, and each processing unit includes a configuration acceptance circuit inside the processing unit.

[0029] Furthermore, the hardware circuit is deployed on the FPGA, without the need for BRAM resources, only requiring DSP, FF and LUT resources, and the resource occupancy decreases as the CGRA scale increases.

[0030] The beneficial effects of the present invention are as follows:

[0031] 1. In the case of multi-task parallel execution, dynamic creation and destruction of multiple applications, the present invention can achieve throughput improvement compared with the representative dynamic method and static method.

[0032] 2. The present invention reduces the hardware delay to about 200 - 500 clock cycles, while weighing the processing unit utilization rate to an average of 82.7%, which is better than the 5000 clock cycle hardware delay and nearly 100% processing unit utilization rate of the representative dynamic method, and the 43.5% processing unit utilization rate of the representative static method.

[0033] 3. The present invention deploys the CGRA multi-task dynamic resource allocation method and the hardware circuit on the FPGA corresponding to different specifications of the CGRA, without the need for BRAM resources, and the required DSP, FF, and LUT resources account for a relatively low proportion, which is suitable for CGRA applications of different scales. Description of the Drawings

[0034] Figure 1It is a schematic structural diagram of an implementation manner of the hardware circuit for realizing the CGRA multi-task dynamic resource allocation method of the present invention;

[0035] Figure 2 It is a method flowchart of the task management module of the present invention;

[0036] Figure 3 It is a method flowchart of the processing unit allocation quantity generation module of the present invention;

[0037] Figure 4 It is a method flowchart of the processing unit shape quantity generation module of the present invention;

[0038] Figure 5 It is a multi-task dynamic resource allocation working flowchart under the control of the finite state machine of the present invention. Specific implementation manner

[0039] Specific implementation manner 1: Combined with Figures 1-5 Describe this implementation manner. The CGRA multi-task dynamic resource allocation method described in this implementation manner is realized through the following steps:

[0040] Step 1: Use the task attribute table to record the attributes of each task. The task attribute table includes task number, number of data flow graph nodes, status, task priority, and the number of processing units assigned to each task. The task attribute table is a RAM or register heap, which is used to be accessed by all other modules to update and query the task status in real time.

[0041] Step 2: Through the task management module, update the task attribute table according to the central processor message, generate a task queue and a Boolean value for triggering dynamic resource allocation, maintain the task attribute table and generate corresponding outputs.

[0042] Step 3: When dynamic resource allocation is triggered, through the processing unit allocation quantity generation module, calculate and allocate the number of processing units according to the task priority and the number of data flow graph nodes. For each task in the task queue, first query the priority of the current task from the task attribute table. The processing unit allocation quantity generation module calculates the allocated number of processing units through weighted average, and uses the proportion obtained from the number of nodes of the task data flow graph after priority weighting or the total number of nodes of all task data flow graphs as the weight.

[0043] Step 4: Through the processing unit shape generation module, determine the shape of the allocated processing units according to the allocated number of processing units. The processing unit shape generation module adopts a coordinate-based heuristic method to allocate regular-shaped processing units for each task to enrich the routing resources and improve the task throughput.

[0044] Step 5: Preload the configuration information into the configuration memory of the processing unit by configuring the preloading module and the bus to match the allocated shape of the processing unit; the design of the configuration preloading module and the bus includes a dual-buffer configuration memory, and the processing unit configuration memory is a dual-buffer, which is controlled by a 1-bit counter to write alternately to hide the latency of reloading the configuration, and the configuration information is transmitted to the processing unit through the specified bus.

[0045] Step 6: Use a finite state machine to control the working order of the modules in Steps 2 - 5 to achieve dynamic resource allocation, and achieve seamless dynamic resource allocation by overlapping configuration preloading with CGRA execution.

[0046] As Figure 1 shown, a hardware circuit for implementing the CGRA multi-task dynamic resource allocation method, on the basis of the original CGRA circuit architecture, adds a dynamic resource allocator, a configuration transmission bus between processing units, and a configuration acceptance circuit inside the processing unit. The dynamic resource allocator is connected to the processing units of the CGRA circuit architecture through the configuration transmission bus between processing units, and each processing unit includes the configuration acceptance circuit inside the processing unit. The dynamic resource allocator is integrated into the CGRA circuit architecture and includes a task attribute table, a task management module, a processing unit allocation quantity generation module, a processing unit shape generation module, a finite state machine, and a configuration preloading module; the configuration transmission bus between processing units is used to transmit configuration information; the configuration acceptance circuit inside the processing unit is used to receive and store configuration information.

[0047] The task attribute table is a RAM or register bank for recording the attributes of each task. The fields include task number (unique task identifier), number of data flow graph nodes, status (whether the task has been executed), task priority affecting the number of processing units allocated, and the processing unit numbers allocated to each task. The task attribute table will be accessed by all other modules.

[0048] As Figure 2As shown, the task management module is used to maintain the task attribute table and generate corresponding outputs, perform corresponding operations on the task attribute table according to the central processing unit (CPU) messages, set the boolean value for triggering dynamic resource allocation, and extract the task queue. The inputs to the task management module include the CPU messages and the task attribute table. The outputs are the updated task queue and a boolean value (indicating whether dynamic resource allocation should be triggered). The function of the task management module is to maintain the task attribute table and generate corresponding outputs. For each message in the CPU messages, if the operation is task construction, all attributes of this task will be recorded in a new row of the task attribute table. If the operation is destruction or task execution completion, the status field of the specific row in the task attribute table identified by the task number will be set to 0. Whenever there is a new task construction, forced destruction, or execution completion, the boolean value will be set to true, meaning that dynamic resource allocation should be triggered. After all these operations are completed, the task numbers with a status field of 1 will be extracted into the task queue.

[0049] As Figure 3 shown, the processing unit allocation quantity generation module calculates and allocates the number of processing units according to the task priority and the number of data flow graph nodes, and restricts it not to exceed the theoretical maximum value. In addition to the outputs of the task management module (including the task queue and a boolean value), the processing unit allocation quantity generation module also takes the task attribute table and the total number of CGRA processing units as inputs, and finally outputs the allocated processing unit numbers for each task in the task queue. If the boolean value is true, the module will generate a new resource allocation decision. For each task in the task queue, first query the priority of the current task from the task attribute table. Then, calculate the allocated number of processing units through weighted average, which uses the ratio obtained by multiplying the number of nodes in the task data flow graph weighted by the priority by the total number of nodes in all task data flow graphs as the weight. In addition, before writing to the "allocated processing unit number" field of the task attribute table, the number of processing units needs to be further normalized using the theoretical maximum value, because additional processing units exceeding the theoretical maximum value will be wasted by the CGRA compiler.

[0050] As Figure 4As shown, the processing unit shape generation module adopts a coordinate-based heuristic method to assign regular-shaped processing units to each task, so as to enrich the routing resources and improve the task throughput. This shape generation is used to further determine the shape of the allocated processing units, so that all the shapes of different tasks correctly form the shape of the CGRA processing unit array. Since regular shapes provide richer routing resources than irregular shapes, the CGRA compiler has the potential to achieve higher task throughput. Therefore, an ideal shape generation algorithm should provide regular shapes for the allocated processing units of each task to maximize the throughput of all tasks. However, the above goal belongs to a combinatorial problem, and its computational complexity increases with the increase in the size of the CGRA and the number of tasks, which does not meet the real-time scheduling requirements and has significant hardware overhead.

[0051] To balance shape regularity and time overhead, the processing unit shape generation module proposes a coordinate-based heuristic method. Taking a 4x4-sized CGRA as an example, both the X and Y coordinates start from 0 and have a maximum value of 3. For each task in the task queue, first obtain the number of allocated processing units for it, and then query two sets of vertical and horizontal lengths corresponding to the number. For example, if the number of processing units = 6, the vertical and horizontal lengths are (horizontal = 2, vertical = 3) and (horizontal = 1, vertical = 6), corresponding to regular and irregular shapes respectively. Then, this method will query the two sets of vertical and horizontal lengths respectively and check whether the current coordinates (X, Y) are still legal after adding the vertical and horizontal lengths. If legal, all the processing unit numbers within the range will be written back to the task attribute table, and (X, Y) will be updated, which indicates that the shape of the processing units allocated for the current task has been generated. Then, move to the next task. Otherwise, the number of processing units will be reduced to relax the constraints and retry.

[0052] The configuration preloading module is used to send the configuration supplemented with the correct processing unit numbers sent by the central processor to each processing unit via the bus and control the configuration to be written into the buffer of the processing unit configuration memory. The task execution on the CGRA is configuration-driven, and under various processing unit shapes, the positions of the configuration in the configuration memory of the processing unit are completely different. Therefore, after dynamic resource allocation, the configuration of each processing unit must be reloaded to match the shape of the allocated processing unit. As Figure 1As shown, in order to hide the latency of reloading, the configuration memory of each processing unit is designed as a double buffer. The configuration sent by the central processing unit will first be supplemented with the correct processing unit number by the preloading module, and then sent to the first processing unit through the specified bus. Each processing unit will always latch the received configuration and then send it to the next processing unit. If the processing unit number field of the latched configuration is equal to the local processing unit number, they will be written into one of the buffers in the configuration memory, which is controlled by a 1-bit counter and written alternately. This design is hardware-friendly and can improve the timing.

[0053] As Figure 5 shown, the finite state machine controls the working order of the task management module, the processing unit allocation quantity generation module, the processing unit shape generation module, and the configuration preloading module, triggering task management, resource allocation, configuration preloading, and task execution in sequence to achieve dynamic resource allocation. After each dynamic resource allocation through task creation or destruction, the finite state machine will iteratively enable the task management module (S1), the processing unit allocation quantity generation module (S2), and the processing unit shape generation module (S3) to generate new resource allocation decisions, which are recorded in the task attribute table. Then it notifies the central processing unit to start the configuration preloading module and transmit the configuration to one of the buffers. Finally, it triggers the CGRA execution after the preloading is completed. The finite state machine realizes seamless dynamic resource allocation by overlapping the configuration preloading with the CGRA execution.

[0054] The present invention proposes a CGRA multi-task dynamic resource allocation method and hardware circuit, finds a new CGRA multi-task dynamic resource allocation method, and designs corresponding hardware modules, so that when tasks are created or destroyed, dynamic resource allocation can be completed within hundreds of clock cycles, thereby improving the CGRA resource utilization rate and further increasing the multi-task throughput rate.

[0055] Embodiment

[0056] 1. RTL Design and Simulation

[0057] First, according to Figures 1 to 4 the hardware architecture and the flowchart of the dynamic resource allocation method, complete the RTL design and functional simulation of the hardware modules. Secondly, use the open-source CGRA tool CGRA-Flow to complete the addition of the circuit and bus interface for receiving the configuration in each processing unit through PyMTL3 and complete the simulation. Finally, generate a synthesizable RTL design through CGRA-Flow.

[0058] 2. Module Instantiation and Compilation

[0059] Use AMD Xilinx Vivado to encapsulate the hardware module and CGRA into IP cores respectively, and open BlockDesign. Instantiate the Processing System IP as the central processor, AXIDMA IP, MIG IP, etc., and complete the correct connection of each module. Subsequently, set the Floorplan, and carry out Synthesis and Implementation, and finally output the hardware platform containing the bitstream.

[0060] 3. Testing

[0061] Use the AMD Xilinx Vivado SDK to open the hardware platform, write C++ code to start the DMA to add multitasks to the CGRA, and set interrupts to detect the completion time of multitask execution and capture the execution results.

[0062] 4. Results

[0063] (1) Throughput improvement

[0064] The present invention, together with the 4x4 CGRA, is deployed on the FPGA of a self-developed acquisition board of a storage recorder, and is tested in the scenario of measuring the "dynamic characteristics of the motor". In the case of parallel execution, dynamic creation and destruction of 17 tasks in 3 applications, the present invention can achieve throughput improvements of 1.34 times and 2.09 times respectively compared with the representative dynamic method and static method.

[0065] (2) Hardware latency and processing unit utilization

[0066] The processing unit utilization of the representative dynamic method is close to 100%, but the hardware latency is about 5000 clock cycles, which is not conducive to the frequent execution of dynamic resource allocation. The present invention reduces the hardware latency to about 200 - 500 clock cycles, but balances the processing unit utilization to an average of 82.7%. While the hardware utilization of the representative static method is only 43.5%.

[0067] (3) Hardware resource cost

[0068] The present invention does not require BRAM resources when deployed on the FPGA, and requires DSP, FF, and LUT resources. It accounts for 16.67%, 9.68%, and 11.67% of the DSP, FF, and LUT resources required by the 4x4 CGRA. It accounts for 4.17%, 2.56%, and 2.86% of the DSP, FF, and LUT resources required by the 8x8 CGRA. It accounts for 1.85%, 1.17%, and 0.76% of the DSP, FF, and LUT resources required by the 12x12 CGRA. It accounts for 1.04%, 1.18%, and 0.79% of the DSP, FF, and LUT resources required by the 16x16 CGRA.

[0069] The above are only the preferred embodiments of the present invention and do not impose any formal restrictions on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments within the spirit and principles of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A CGRA multi-task dynamic resource allocation method, characterized in that: The above-mentioned CGRA multi-task dynamic resource allocation method is implemented through the following steps: Step 1: Use the task attribute table to record the attributes of each task; Step 2: Through the task management module, update the task attribute table according to the central processor message, and generate a task queue and a Boolean value for triggering dynamic resource allocation; Step 3: When dynamic resource allocation is triggered, through the processing unit allocation quantity generation module, calculate and allocate the number of processing units according to the task priority and the number of data flow graph nodes; Step 4: Through the processing unit shape generation module, determine the shape of the allocated processing units according to the number of allocated processing units; Step 5: Through the configuration preloading module and the bus, preload the configuration information into the configuration memory of the processing unit to match the allocated processing unit shape; Step 6: Use a finite state machine to control the working order of the modules in Steps 2 - 5 to achieve dynamic resource allocation.

2. The CGRA multi-task dynamic resource allocation method according to claim 1, characterized in that: The task attribute table described in Step 1 includes task number, number of data flow graph nodes, status, task priority, and the number of processing units allocated to each task.

3. A CGRA multi-task dynamic resource allocation method according to claim 2, characterized in that: The task attribute table described in Step 1 is a RAM or register file.

4. A CGRA multi-task dynamic resource allocation method according to claim 1, characterized in that: The processing unit allocation quantity generation module described in Step 3 calculates the allocated number of processing units through weighted average, and uses the proportion obtained from the number of task data flow graph nodes after priority weighting or the total number of nodes of all task data flow graphs as the weight.

5. A CGRA multi-task dynamic resource allocation method according to claim 1, characterized in that: The processing unit shape generation module described in Step 4 uses a coordinate-based heuristic method to allocate regular-shaped processing units to each task.

6. The CGRA multi-task dynamic resource allocation method according to claim 1, characterized in that: The configuration preloading module and bus design described in Step 5 includes a dual-buffer configuration memory to hide the latency of reloading the configuration, and transmits the configuration information to the processing unit through a specified bus.

7. A hardware circuit for implementing the CGRA multi-task dynamic resource allocation method according to any one of claims 1 to 6, characterized in that: It includes: A CGRA circuit architecture, a dynamic resource allocator, a configuration transmission bus between processing units, and a configuration acceptance circuit inside the processing unit. The dynamic resource allocator is integrated into the CGRA circuit architecture and includes a task attribute table, a task management module, a processing unit allocation quantity generation module, a processing unit shape generation module, a finite state machine, and a configuration preloading module; the configuration transmission bus between processing units is used to transmit configuration information; the configuration acceptance circuit inside the processing unit is used to receive and store configuration information.

8. The hardware circuit for implementing the CGRA multi-task dynamic resource allocation method according to claim 7, characterized in that: In the above-mentioned dynamic resource allocator, The task attribute table is used to record the attributes of each task, including task number, number of data flow graph nodes, task status, priority, and the number of allocated processing units, and is accessed by other modules; The task management module is used to maintain the task attribute table and generate corresponding outputs, perform corresponding operations on the task attribute table according to the central processor message, set the Boolean value for triggering dynamic resource allocation, and extract the task queue; The processing unit allocation quantity generation module calculates and allocates the number of processing units according to the task priority and the number of data flow graph nodes, and restricts it not to exceed the theoretical maximum value; The processing unit shape generation module uses a coordinate-based heuristic method to allocate regular-shaped processing units to each task to enrich routing resources and improve task throughput; Configure a preloading module, which is used to send the configuration supplemented with the correct processing unit number sent by the central processing unit to each processing unit via the bus, and control the configuration to be written into the buffer of the processing unit configuration memory; A finite state machine controls the working order of the task management module, the processing unit allocation quantity generation module, the processing unit shape generation module, and the configuration preloading module, triggering task management, resource allocation, configuration preloading, and task execution in sequence to achieve dynamic resource allocation.

9. The hardware circuit for implementing the CGRA multi-task dynamic resource allocation method according to claim 8, wherein: The dynamic resource allocator is connected to the CGRA circuit architecture processing unit through the configuration transmission bus between the processing units, and each processing unit includes a configuration receiving circuit inside the processing unit.

10. The hardware circuit for implementing the CGRA multi-task dynamic resource allocation method according to claim 9, characterized in that: The hardware circuit is deployed on the FPGA, without the need for BRAM resources, only requiring DSP, FF, and LUT resources, and the resource occupancy decreases as the CGRA scale increases.