Parallel processor dynamic resource allocation system and method based on reconfigurable hardware

By integrating reconfigurable hardware within the parallel processor chip, dynamic adjustment and optimized allocation of hardware resources are achieved, solving the inefficiency problem caused by the fixed allocation of resources in traditional parallel processors, improving resource utilization and computing performance, and reducing energy consumption.

CN120994369APending Publication Date: 2025-11-21YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511048560.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The fixed allocation of hardware resources in traditional parallel processors leads to low resource utilization efficiency and cannot meet the needs of complex and ever-changing computing tasks. Existing dynamic resource allocation methods at the software level cannot achieve deep reconfiguration and efficient utilization of hardware resources.

Method used

By integrating reconfigurable hardware within the parallel processor chip, combined with a workload monitoring module and a resource allocation control module, dynamic adjustment and optimized allocation of hardware resources can be achieved, including workload monitoring, resource allocation control, and dynamic reconfiguration of the reconfigurable hardware module, to adapt to different task requirements.

Benefits of technology

It improves resource utilization and computing performance, reduces energy consumption, and meets the requirements of complex computing tasks, especially significantly improving overall performance and resource utilization efficiency when handling mixed tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994369A_ABST
    Figure CN120994369A_ABST
Patent Text Reader

Abstract

The invention discloses a parallel processor dynamic resource allocation system and method based on reconfigurable hardware, and relates to the technical field of computer chips, and the system comprises a workload monitoring module which is responsible for monitoring the workload type and the resource demand of a task processed by a parallel processor in real time, and determining the task type and the resource demand by identifying the task type and the resource demand; a monitoring result is fed back to the resource allocation control module; the resource allocation control module is responsible for generating a corresponding resource allocation control signal based on feedback information of the workload monitoring module and dynamically configuring the reconfigurable hardware module; the reconfigurable hardware module is responsible for realizing efficient adaptation to diversified tasks through a plurality of reconfigurable units formed by programmable logic devices according to different working load dynamic reconfiguration functions and connection modes; and the data caching and transmission module is responsible for cross-module data circulation and global data storage. Different task requirements can be accurately adapted, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer chips, in particular to a parallel processor dynamic resource allocation system and method based on reconfigurable hardware. BACKGROUND

[0002] With the rapid development of artificial intelligence, virtual reality, high-definition video rendering and other technologies, the application scenarios of parallel processors are increasingly diverse, and the workloads they face are becoming more complex and diverse. In the traditional architecture design of parallel processors, hardware resources are usually fixedly allocated, such as the configuration of computing units, cache units, memory bandwidth and other resources, which cannot be changed after the chip is manufactured.

[0003] Taking two common application scenarios of graphics rendering and deep learning training as examples, in the graphics rendering process, the parallel processor needs a large amount of resources for texture processing, rasterization and other operations; while in the deep learning training task, matrix operation and data storage access become the focus of resource consumption. When the parallel processor switches between different types of tasks or processes multiple tasks at the same time, due to the inability to dynamically adjust hardware resources, it will lead to low resource utilization efficiency. After the graphics rendering task is completed, a large number of texture processing units originally used for graphics processing may be idle, while the matrix operation unit cannot fully exert its performance due to insufficient resources, resulting in decreased overall computing efficiency and increased energy consumption.

[0004] In addition, some existing dynamic resource allocation methods are mostly based on software-level scheduling strategies, which adjust the resource usage of parallel processors through operating systems or application programs. Although this approach can alleviate the problem of unreasonable resource allocation to some extent, it is limited by software scheduling and the fixed nature of hardware architecture, and cannot achieve deep reconfiguration and efficient utilization of hardware resources, making it difficult to meet the performance requirements of current complex and diverse computing tasks for parallel processors. Therefore, there is an urgent need for a parallel processor architecture that can dynamically allocate resources at the hardware level to improve the performance and resource utilization efficiency of parallel processors under different workloads. SUMMARY

[0005] The present application provides a parallel processor dynamic resource allocation system and method based on reconfigurable hardware to address the needs and deficiencies of current technology. By integrating reconfigurable hardware within the parallel processor chip, the system achieves dynamic adjustment and optimized allocation of hardware resources, improving the performance and resource utilization efficiency of parallel processors under different workloads, and reducing energy consumption.

[0006] In the first aspect, the present application provides a parallel processor dynamic resource allocation system based on reconfigurable hardware, which solves the above technical problems by adopting the following technical solutions:

[0007] A reconfigurable hardware-based parallel processor dynamic resource allocation system, the structure of which comprises:

[0008] A workload monitoring module, responsible for real-time monitoring of the workload type and resource demand of the tasks processed by the parallel processor, and feeding back the monitoring results to the resource allocation control module to provide decision basis for resource allocation;

[0009] A resource allocation control module, responsible for generating corresponding resource allocation control signals based on the feedback information of the workload monitoring module to dynamically configure the reconfigurable hardware module;

[0010] A reconfigurable hardware module, responsible for dynamically reconfiguring the function and connection mode of a plurality of reconfigurable units composed of programmable logic devices according to different workloads to achieve efficient adaptation to diversified tasks;

[0011] A data cache and transmission module, responsible for cross-module data flow and global data storage.

[0012] Optionally, the workload monitoring module involved is the sensing core of the parallel processor resource scheduling, and its implementation and operation mechanism are as follows:

[0013] The workload monitoring module is constructed by a special monitoring chip or a monitoring circuit integrated in the parallel processor. The monitoring chip or monitoring circuit is directly connected to the instruction execution unit and data storage unit of the parallel processor, and real-time acquisition of key operation data in the task execution process is performed, including instruction stream sequence, data access address distribution and calculation cycle time consumption;

[0014] After the acquisition is completed, the processing unit inside the workload monitoring module inputs the instruction stream sequence into the load type identification algorithm, inputs the data access address distribution into the data access frequency analysis algorithm, and inputs the calculation cycle time consumption into the computing power demand estimation algorithm, and deep analysis of the data is performed with the aid of digital signal processing technology. The workload type and resource demand results obtained by analysis are fed back to the resource allocation control module in real time to provide accurate and real-time decision basis for dynamically adjusting the resource allocation strategy, ensuring efficient matching of parallel processor resources and task demand.

[0015] Preferably, the load type identification algorithm involves a pattern matching algorithm based on instruction sequence characteristics or a lightweight deep learning classification model with a parameter quantity <100K.

[0016] Optionally, the specific implementation and operation mechanism of the resource allocation control module involved are as follows:

[0017] The resource allocation control module takes a microcontroller or a special control logic circuit as a hardware carrier, first receives task workload information transmitted by the workload monitoring module, then analyzes and decides the received information according to a preset resource allocation strategy and algorithm, generates a control signal containing a hardware resource allocation ratio, a reconfigurable hardware module working mode and data cache capacity allocation content, and synchronously transmits the generated control signal to the reconfigurable hardware module and the data cache and transmission module through a special control bus, and finally completes real-time dynamic configuration and flexible adjustment of the hardware resources, ensuring efficient matching of system resources and task load.

[0018] Optionally, the plurality of reconfigurable units specifically include:

[0019] The computing unit is configured to support dynamic configuration of operation precision and parallelism, adjust hardware logic according to task requirements to perform various computing tasks, and the like.

[0020] The local cache unit is configured to temporarily store intermediate data in a computing process, and dynamically adjust the cache capacity according to requirements to adapt to data throughput requirements of different tasks.

[0021] The internal transmission unit is configured to realize data interaction between the computing unit and the local cache unit, support dynamic adaptation of bandwidth, and adjust transmission logic and rate according to data transmission requirements.

[0022] Preferably, the computing unit is constructed based on a reconfigurable logic architecture of an FPGA, utilizes logic units of the FPGA to realize diversified computing functions, and is flexibly configured as a general arithmetic logic unit, a floating point operation unit and a special matrix operation unit, realizes dynamic reconfiguration of hardware functions through programmable characteristics of the FPGA, and meets requirements of different computing tasks.

[0023] The local cache unit relies on on-chip storage resources of the FPGA to build a multi-level cache structure, includes instruction cache and data cache, and optimizes cache hierarchy and capacity allocation through reconfigurable configuration to improve data read-write efficiency.

[0024] The internal transmission unit realizes efficient transmission and exchange of data between units by means of high-speed I / O interfaces and internal buses of the FPGA, dynamically adjusts interface bandwidth and bus paths according to data transmission requirements, ensures efficient and low-delay data interaction between the computing unit and the local cache unit, and between each reconfigurable unit and external modules, and adapts to transmission load in different scenarios.

[0025] Optionally, the data cache and transmission module is built through high-speed cache chips and high-speed data transmission interface chips at a hardware level to realize efficient storage, rapid interaction and global consistency management of data.

[0026] Optionally, the data caching and transmission modules involved specifically include:

[0027] The global cache unit is a multi-level cache structure built on static random access memory and dynamic random access memory. Through the coordinated scheduling of the multi-level cache structure, the speed and capacity requirements of data storage are balanced. Among them, static random access memory serves as the first-level cache, which is used to store real-time task data, core instructions and intermediate calculation results that are accessed more frequently than a preset threshold; dynamic random access memory serves as the second-level cache, which is used to cache all data, non-real-time instructions and historical interaction records generated during task execution.

[0028] The cross-module transmission unit relies on a dedicated data transmission interface chip based on the high-speed serial interface standard to build a high-speed data channel with low latency and high bandwidth. Through a standardized interface protocol, it realizes bidirectional interaction between real-time load data of the workload monitoring module, scheduling instructions of the resource allocation control module, and execution status data of the reconfigurable hardware module, ensuring the timeliness and reliability of cross-module data transmission and supporting the collaborative operation of various modules.

[0029] Secondly, the present invention provides a method for dynamic resource allocation of parallel processors based on reconfigurable hardware, and the technical solution adopted to solve the above-mentioned technical problems is as follows:

[0030] A method for dynamic resource allocation of parallel processors based on reconfigurable hardware, which is based on the dynamic resource allocation system for parallel processors described in the first aspect, specifically includes the following steps:

[0031] S1. When the parallel processor starts processing a task, the workload monitoring module is immediately activated, continuously collecting information during task execution. After real-time analysis and processing, it identifies the workload type and specific resource requirements of the current task and feeds it back to the resource allocation control module.

[0032] S2. After receiving the information from the workload monitoring module, the resource allocation control module makes resource allocation decisions based on the pre-set resource allocation strategy and algorithm. Then, combined with the current resource status of the reconfigurable hardware module, it generates corresponding resource allocation control signals to determine the specific reconfigurable unit that needs to be reconfigured and its reconfiguration method.

[0033] S3. After receiving the control signal from the resource allocation control module, the reconfigurable hardware module performs a hardware reconfiguration operation on the specific reconfigurable unit.

[0034] S4. After the reconfigurable hardware unit completes its reconfiguration, the parallel processor begins to execute the task. During the task execution, the workload monitoring module continuously monitors the load changes. If adjustments are needed, the S1-S3 process is repeated until the task is completed, thereby achieving dynamic resource adaptation and efficient utilization.

[0035] The present invention provides a parallel processor dynamic resource allocation system and method based on reconfigurable hardware, which has the following advantages compared with the prior art:

[0036] 1. This invention precisely adapts to different task requirements through hardware reconfiguration and dynamic allocation, improves resource utilization and computing performance while reducing energy consumption, meets the requirements of complex computing tasks, and can be widely used in fields with high requirements for parallel processor performance, such as graphics processing, deep learning computing, and scientific computing.

[0037] 2. This invention integrates a reconfigurable hardware module inside the parallel processor chip, and combines it with a workload monitoring and resource allocation control module to achieve dynamic adjustment and optimized allocation of hardware resources. It can flexibly adjust the allocation of internal resources of the parallel processor according to different workload requirements, avoid resource idleness and waste, and significantly improve resource utilization efficiency. Especially when processing mixed tasks, the reconfigurable hardware module can quickly allocate resources to the most needed tasks, so that the overall performance of the parallel processor can be fully utilized.

[0038] 3. This invention provides the most suitable computing resources and architecture for different types of tasks through hardware reconfiguration, thereby accelerating task execution speed and improving the computing performance of parallel processors under various workloads. In particular, in deep learning training tasks, the specially configured matrix operation unit can significantly improve training speed and shorten training time.

[0039] 4. By making reasonable resource allocation, this invention can reduce unnecessary resource consumption and lower the overall energy consumption of parallel processors. Especially when the task has low requirements for computing resources, the reconfigurable hardware module can reduce the power consumption of some units and achieve energy-saving operation. Attached Figure Description

[0040] Appendix Figure 1 This is a module connection block diagram of Embodiment 1 of the present invention;

[0041] Appendix Figure 2 This is a flowchart of the method according to Embodiment 2 of the present invention. Detailed Implementation

[0042] To make the technical solution, the technical problem solved, and the technical effect of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments.

[0043] Example 1:

[0044] Combined with appendix Figure 1 This embodiment proposes a parallel processor dynamic resource allocation system based on reconfigurable hardware, the structure of which includes:

[0045] The workload monitoring module is responsible for real-time monitoring of the workload type and resource requirements of the tasks processed by the parallel processor. By identifying the task type and resource requirements, the monitoring results are fed back to the resource allocation control module to provide a basis for resource allocation decisions.

[0046] The resource allocation control module is responsible for generating corresponding resource allocation control signals based on feedback information from the workload monitoring module, and dynamically configuring the reconfigurable hardware modules.

[0047] The reconfigurable hardware module is responsible for dynamically reconfiguring functions and connection methods according to different workloads through multiple reconfigurable units composed of programmable logic devices, so as to achieve efficient adaptation to diverse tasks.

[0048] The data caching and transmission module is responsible for cross-module data flow and global data storage.

[0049] In this embodiment, the workload monitoring module serves as the core of parallel processor resource scheduling, and its implementation and operation mechanism are as follows:

[0050] The workload monitoring module is constructed using a dedicated monitoring chip or a monitoring circuit integrated within the parallel processor. This chip or circuit is directly connected to the instruction execution unit and data storage unit of the parallel processor, collecting key operational data in real time during task execution, including instruction stream sequences, data access address distribution, and computation cycle time. After collection, the processing unit within the workload monitoring module inputs the instruction stream sequence into a load type identification algorithm, the data access address distribution into a data access frequency analysis algorithm, and the computation cycle time into a computing power requirement estimation algorithm. It then uses digital signal processing (DSP) technology to perform in-depth data analysis, feeding back the resulting workload type and resource requirements to the resource allocation control module in real time. This provides accurate and real-time decision-making basis for dynamically adjusting resource allocation strategies, ensuring efficient matching between parallel processor resources and task requirements.

[0051] Specifically, the load type identification algorithm uses a pattern matching algorithm based on instruction sequence features (to distinguish load types by comparing feature codes such as graphics rendering instructions and matrix operation instructions in the instruction stream) or a lightweight deep learning classification model with fewer than 100K parameters (such as a small CNN or decision tree to classify instruction streams and data access patterns).

[0052] In this embodiment, the specific implementation and operation mechanism of the resource allocation control module are as follows:

[0053] The resource allocation control module uses a microcontroller (MCU) or dedicated control logic circuit as its hardware carrier. First, it receives task workload information (such as task type, computational load, and real-time priority) transmitted by the workload monitoring module. Then, based on preset resource allocation strategies (such as load balancing strategies and priority scheduling strategies) and algorithms (such as dynamic programming algorithms and greedy algorithms), it analyzes and makes decisions on the received information, generating control signals that include hardware resource allocation ratios, reconfigurable hardware module operating modes, and data cache capacity allocation. These control signals are synchronously transmitted via a dedicated control bus to the reconfigurable hardware module (to activate, combine, or put the computing unit to sleep) and the data cache and transmission module (to adjust cache partition size and data transmission bandwidth). Ultimately, this completes the real-time dynamic configuration and flexible adjustment of hardware resources, ensuring efficient matching between system resources and task load.

[0054] In this embodiment, the multiple reconfigurable units involved include a computing unit, a local cache unit, and an internal transmission unit.

[0055] The computing unit supports dynamic configuration of computational precision and parallelism, adjusting hardware logic according to task requirements to execute various computational tasks. Specifically, the computing unit is built on a reconfigurable logic architecture based on FPGA, utilizing the FPGA's logic units to implement diverse computing functions. It can be flexibly configured as a general-purpose arithmetic logic unit (ALU, supporting basic addition, subtraction, multiplication, and division), a floating-point unit (FPU, handling high-precision decimal operations), and a dedicated matrix operation unit (adapted to large-scale matrix calculations in scenarios such as deep learning). Through the programmable characteristics of the FPGA, the hardware functions can be dynamically reconfigured to meet the needs of different computing tasks.

[0056] The local cache unit temporarily stores intermediate data during the computation process, and its capacity is dynamically adjusted as needed to adapt to the data throughput requirements of different tasks. Specifically, the local cache unit relies on the on-chip storage resources of the FPGA to build a multi-level cache structure, including an instruction cache (used to temporarily store operation instructions to be executed, speeding up instruction reading) and a data cache (temporarily storing frequently accessed data during the computation process, reducing data access latency). Through reconfigurable configuration, the cache hierarchy and capacity allocation are optimized to improve data read and write efficiency.

[0057] The internal transmission unit enables data interaction between the computing unit and the local cache unit, supporting dynamic bandwidth adaptation and adjusting the transmission logic and rate according to data transmission requirements. Specifically, the internal transmission unit utilizes the FPGA's high-speed I / O interface and internal bus to achieve efficient data transmission and exchange between units, and dynamically adjusts the interface bandwidth and bus path according to data transmission requirements. This ensures efficient and low-latency data interaction between the computing unit and the local cache unit, as well as between each reconfigurable unit and external modules, adapting to transmission loads in different scenarios.

[0058] In this embodiment, the data caching and transmission module is built using a high-speed cache chip and a high-speed data transmission interface chip at the hardware level to achieve efficient data storage, fast interaction, and global consistency management.

[0059] The data caching and transmission modules involved specifically include:

[0060] The global cache unit is based on a multi-level cache structure built with static random access memory (SRAM) and dynamic random access memory (DRAM). Through the coordinated scheduling of the multi-level cache structure, the speed and capacity requirements of data storage are balanced. Among them, static random access memory (SRAM) serves as the first-level cache, used to store real-time task data, core instructions, and intermediate calculation results that are accessed more frequently than a preset threshold; dynamic random access memory (DRAM) serves as the second-level cache, used to cache all data generated during task execution, non-real-time instructions, and historical interaction records.

[0061] The cross-module transmission unit relies on dedicated data transmission interface chips based on high-speed serial interface standards (such as PCIe, SerDes, etc.) to build a high-speed data channel with low latency and high bandwidth. Through standardized interface protocols, it realizes bidirectional interaction between real-time load data of the workload monitoring module, scheduling instructions of the resource allocation control module, and execution status data of the reconfigurable hardware module, ensuring the timeliness and reliability of cross-module data transmission and supporting the collaborative operation of various modules.

[0062] Example 2:

[0063] Combined with appendix Figure 2 This embodiment proposes a method for dynamic resource allocation of parallel processors based on reconfigurable hardware. It is based on the dynamic resource allocation system for parallel processors described in Embodiment 1, and specifically includes the following steps:

[0064] S1. When the parallel processor begins processing a task, the workload monitoring module immediately starts, continuously collecting information during task execution: For graphics rendering tasks, the workload monitoring module focuses on information such as the frequency of texture sampling instructions and the amount of vertex processing data; for deep learning training tasks, the workload monitoring module focuses on monitoring information such as the type and number of matrix operation instructions and the batch size of data access. After real-time analysis and processing of the collected information, the workload type and specific resource requirements of the current task are identified and fed back to the resource allocation control module.

[0065] S2. After receiving information from the workload monitoring module, the resource allocation control module makes resource allocation decisions based on pre-set resource allocation strategies and algorithms. If it determines that the current task is a deep learning training task with high demand for matrix operation resources, the resource allocation control module calculates resource parameters such as the required number of matrix operation units, cache size, and data transmission bandwidth. Then, combined with the current resource status of the reconfigurable hardware module, it generates corresponding resource allocation control signals to determine the specific reconfigurable unit (computation unit / local cache unit / internal transmission unit) that needs to be reconfigured and its reconfiguration method.

[0066] S3. After receiving the control signal from the resource allocation control module, the reconfigurable hardware module performs hardware reconfiguration operation on the specific reconfigurable unit (computing unit / local cache unit / internal transmission unit).

[0067] For FPGA-based computing units, the logic functions and connectivity of the FPGA are changed by loading a new configuration file. When a computing unit is reconfigured into a matrix operation unit, the FPGA reconfigures its internal logic units and builds dedicated matrix multiplication and addition circuits. At the same time, the local cache unit adjusts the size and mapping method of the cache according to the resource allocation control signal to adapt to the storage requirements of matrix operation data. The internal transmission unit adjusts the priority and bandwidth allocation of the data transmission channel to ensure that matrix operation data can be transmitted to each computing unit quickly and accurately.

[0068] S4. After the reconfigurable hardware unit completes its reconfiguration, the parallel processor begins to execute the task. During the task execution, the workload monitoring module continuously monitors the load changes. If adjustments are needed, the S1-S3 process is repeated until the task is completed, thereby achieving dynamic resource adaptation and efficient utilization.

[0069] In summary, the parallel processor dynamic resource allocation system and method based on reconfigurable hardware of this invention can accurately adapt to different task requirements through hardware reconfiguration and dynamic allocation, improve resource utilization and computing performance, reduce energy consumption, meet the requirements of complex computing tasks, and solve the problem that existing dynamic resource allocation methods cannot achieve deep reconfiguration and efficient utilization of hardware resources due to the limitations of software scheduling and the fixed nature of hardware architecture.

[0070] The above specific examples illustrate the principles and implementation methods of the present invention in detail. These embodiments are merely for the purpose of helping to understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made to the present invention by those skilled in the art without departing from the principles of the present invention should fall within the patent protection scope of the present invention.

Claims

1. A parallel processor dynamic resource allocation system based on reconfigurable hardware, characterized in that, Its structure includes: The workload monitoring module is responsible for real-time monitoring of the workload type and resource requirements of the tasks processed by the parallel processor. By identifying the task type and resource requirements, the monitoring results are fed back to the resource allocation control module to provide a basis for resource allocation decisions. The resource allocation control module is responsible for generating corresponding resource allocation control signals based on feedback information from the workload monitoring module, and dynamically configuring the reconfigurable hardware modules. The reconfigurable hardware module is responsible for dynamically reconfiguring functions and connection methods according to different workloads through multiple reconfigurable units composed of programmable logic devices, so as to achieve efficient adaptation to diverse tasks. The data caching and transmission module is responsible for cross-module data flow and global data storage.

2. The parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 1, characterized in that, The workload monitoring module, as the core of parallel processor resource scheduling, is implemented and operates as follows: The workload monitoring module is constructed through a dedicated monitoring chip or a monitoring circuit integrated inside the parallel processor. The monitoring chip or monitoring circuit is directly connected to the instruction execution unit and data storage unit of the parallel processor to collect key operational data in real time during task execution, including instruction stream sequence, data access address distribution and computation cycle time. After data collection is completed, the processing unit inside the workload monitoring module inputs the instruction stream sequence into the load type identification algorithm, the data access address distribution into the data access frequency analysis algorithm, and the computation cycle time into the computing power requirement estimation algorithm. It then uses digital signal processing technology to perform in-depth analysis of the data and feeds back the analyzed workload type and resource requirement results to the resource allocation control module in real time. This provides the module with accurate and real-time decision-making basis for dynamically adjusting resource allocation strategies, ensuring efficient matching between parallel processor resources and task requirements.

3. The parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 2, characterized in that, The load type identification algorithm employs a pattern matching algorithm based on instruction sequence features or a lightweight deep learning classification model with fewer than 100K parameters.

4. The parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 1, characterized in that, The specific implementation and operation mechanism of the resource allocation control module are as follows: The resource allocation control module uses a microcontroller or dedicated control logic circuit as its hardware carrier. First, it receives task workload information transmitted by the workload monitoring module. Then, based on a preset resource allocation strategy and algorithm, it analyzes and makes decisions on the received information, generating control signals that include hardware resource allocation ratios, reconfigurable hardware module operating modes, and data cache capacity allocation. The generated control signals are synchronously transmitted to the reconfigurable hardware module and the data cache and transmission module through a dedicated control bus, ultimately completing the real-time dynamic configuration and flexible adjustment of hardware resources to ensure efficient matching between system resources and task load.

5. A parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 1, characterized in that, The plurality of reconfigurable units specifically include: The computing unit is used to support dynamic configuration of computational precision and parallelism, and adjusts the hardware logic according to task requirements to execute various computing tasks. Local cache units are used to temporarily store intermediate data during the computation process, and the cache capacity is dynamically adjusted as needed to adapt to the data throughput requirements of different tasks. The internal transmission unit is used to realize data interaction between the computing unit and the local cache unit, supports dynamic bandwidth adaptation, and adjusts the transmission logic and rate according to data transmission requirements.

6. A parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 5, characterized in that, The computing unit is built on the reconfigurable logic architecture of FPGA, which uses the logic units of FPGA to realize diverse computing functions and can be flexibly configured as a general arithmetic logic unit, a floating-point operation unit and a dedicated matrix operation unit. The programmable characteristics of FPGA enable dynamic reconfiguration of hardware functions to meet the needs of different computing tasks. The local cache unit relies on the on-chip storage resources of the FPGA to build a multi-level cache structure, including instruction cache and data cache. The cache hierarchy and capacity allocation are optimized through reconfigurable configuration to improve data read and write efficiency. The internal transmission unit utilizes the FPGA's high-speed I / O interface and internal bus to achieve efficient data transmission and exchange between units. It dynamically adjusts the interface bandwidth and bus path according to data transmission requirements to ensure efficient and low-latency data interaction between the computing unit and the local cache unit, as well as between each reconfigurable unit and external modules, adapting to transmission loads in different scenarios.

7. A parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 1, characterized in that, The data caching and transmission module is built using a high-speed cache chip and a high-speed data transmission interface chip at the hardware level to achieve efficient data storage, fast interaction, and global consistency management.

8. A parallel processor dynamic resource allocation system based on reconfigurable hardware according to claim 7, characterized in that, The data caching and transmission module specifically includes: The global cache unit is a multi-level cache structure built on static random access memory and dynamic random access memory. Through the coordinated scheduling of the multi-level cache structure, the speed and capacity requirements of data storage are balanced. Among them, static random access memory serves as the first-level cache, which is used to store real-time task data, core instructions and intermediate calculation results that are accessed more frequently than a preset threshold; dynamic random access memory serves as the second-level cache, which is used to cache all data, non-real-time instructions and historical interaction records generated during task execution. The cross-module transmission unit relies on a dedicated data transmission interface chip based on the high-speed serial interface standard to build a high-speed data channel with low latency and high bandwidth. Through a standardized interface protocol, it realizes bidirectional interaction between real-time load data of the workload monitoring module, scheduling instructions of the resource allocation control module, and execution status data of the reconfigurable hardware module, ensuring the timeliness and reliability of cross-module data transmission and supporting the collaborative operation of various modules.

9. A method for dynamic resource allocation of parallel processors based on reconfigurable hardware, characterized in that, Based on the parallel processor dynamic resource allocation system as described in any one of claims 1-8, it specifically includes the following steps: S1. When the parallel processor starts processing a task, the workload monitoring module is immediately activated, continuously collecting information during task execution. After real-time analysis and processing, it identifies the workload type and specific resource requirements of the current task and feeds it back to the resource allocation control module. S2. After receiving the information from the workload monitoring module, the resource allocation control module makes resource allocation decisions based on the pre-set resource allocation strategy and algorithm. Then, combined with the current resource status of the reconfigurable hardware module, it generates corresponding resource allocation control signals to determine the specific reconfigurable unit that needs to be reconfigured and its reconfiguration method. S3. After receiving the control signal from the resource allocation control module, the reconfigurable hardware module performs a hardware reconfiguration operation on the specific reconfigurable unit. S4. After the reconfigurable hardware unit completes its reconfiguration, the parallel processor begins to execute the task. During the task execution, the workload monitoring module continuously monitors the load changes. If adjustments are needed, the S1-S3 process is repeated until the task is completed, thereby achieving dynamic resource adaptation and efficient utilization.

Citation Information

Cited By

  • Quantization precision adaptive switching hardware acceleration architecture

    CN121351741A

  • Industrial edge computing terminal system and reconstruction method

    CN121985056A

  • Memory allocation method and device, computer equipment and storage medium

    CN122044827A

  • A multi-modal device and a storage resource calling method, configuration method and device thereof

    CN122507528A