Dynamic allocation method and system for workload of CPU-GPU (Central Processing Unit-Graphics Processing Unit) heterogeneous system

By using FPGA chips and PSO algorithms in CPU-GPU heterogeneous systems, dynamically compute and adjust the CPU/GPU mapping ratio, the problem of static task allocation method in the existing technology is solved, and the system resource utilization rate and task processing efficiency are improved.

CN120066791AActive Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510189285.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The task allocation method in existing CPU-GPU heterogeneous systems is relatively simple and static, and the dynamic characteristics of the task and the real-time status of the system resources cannot be fully considered, resulting in low system resource utilization and extended task processing time.

Method used

By introducing FPGA chips into the CPU-GPU heterogeneous system, the PSO algorithm combines the task completion rate and delay model, the optimal CPU/GPU mapping ratio is dynamically calculated, and the task allocation strategy is adjusted in real time.

Benefits of technology

Real-time adjustments are realized based on the dynamic characteristics of the task and the real-time status of the system resources, improving the utilization rate of system resources and task processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066791A_ABST
    Figure CN120066791A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of memory management, in particular to a CPU-GPU heterogeneous system workload dynamic allocation method and system. According to the CPU-GPU heterogeneous system workload dynamic allocation method, a CPU performs preliminary analysis and preprocessing on task data, extracts key feature information of a task and judges a task type; the FPGA chip realizes a PSO algorithm, performs calculation in combination with a task completion rate and a delay model, and finds an optimal CPU / GPU mapping proportion; the scheduling and distribution unit accurately distributes tasks between the CPU and the GPU according to the returned optimization result; the task completion condition and delay data in the execution process are fed back to the FPGA chip, and the current PSO algorithm calculation is adjusted in real time. According to the CPU-GPU heterogeneous system workload dynamic allocation method and system, the workload can be adjusted in real time according to the dynamic characteristics of the task and the real-time state of the system resources, the utilization rate of the system resources is increased, and the task processing efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of memory management, and particularly to a method and system for dynamically allocating workloads in a CPU-GPU heterogeneous system. Background Art

[0002] With the rapid development of technology, the complexity and scale of various computing tasks have increased explosively. In scientific research, such as climate simulation, molecular dynamics simulation, etc., the amount of data to be processed and the amount of computation are extremely large. If these tasks are processed only by a single CPU or GPU, the efficiency is often low and cannot meet the requirements of practical applications.

[0003] To cope with the increasing computing demands, CPU-GPU heterogeneous systems have emerged and been widely used. The CPU has powerful logical control and complex instruction processing capabilities and is good at handling serial tasks and system management; while the GPU, with its highly parallel computing architecture, shows great advantages in processing large-scale data parallel computing tasks. This heterogeneous system combines the strengths of both and has become an important infrastructure in the modern computing field. However, how to fully utilize the performance advantages of this heterogeneous system and achieve efficient allocation of tasks between the CPU and the GPU has become a key problem to be solved urgently.

[0004] Currently, in CPU-GPU heterogeneous systems, the task allocation methods are mostly relatively simple and static. Many applications adopt a fixed-ratio task allocation strategy, or simply make a rough division according to the type of tasks, such as allocating all computationally intensive tasks to the GPU and other tasks to the CPU. This method does not fully consider the dynamic characteristics of tasks and the real-time state of system resources and cannot be adjusted in real time according to these changes, resulting in low utilization of system resources and extended task processing time.

[0005] Based on the above problems, the present invention proposes a method and system for dynamically allocating workloads in a CPU-GPU heterogeneous system. Summary of the Invention

[0006] The present invention provides a simple and efficient method and system for dynamically allocating workloads in a CPU-GPU heterogeneous system to make up for the defects of the prior art.

[0007] The present invention is implemented by the following technical solutions:

[0008] A method for dynamically allocating workloads in a CPU-GPU heterogeneous system, characterized by comprising the following steps:

[0009] Step S1, task data preprocessing

[0010] After receiving the task data, the CPU performs preliminary parsing and preprocessing on the task data through the preprocessing unit, extracts the key feature information of the task, including the task data volume, task priority, and the load conditions of the CPU and GPU in the current system; meanwhile, it determines the task type;

[0011] Step S2, Data Transfer and PSO Algorithm Calculation

[0012] After preprocessing, the CPU transfers the data related to the task to the FPGA chip. After receiving the data, the FPGA chip uses its internal hardware logic to implement the PSO algorithm, calculates in combination with the task completion rate and the latency model, and finds the optimal CPU / GPU mapping ratio;

[0013] In the said step S2, the FPGA chip evaluates the advantages and disadvantages of different CPU / GPU mapping ratio schemes according to the task completion rate and the latency model;

[0014] Among them, the task completion rate is calculated by statistically calculating the ratio of the number of tasks completed within a unit time to the total number of tasks; the latency is obtained by recording the time difference from the task submission to completion;

[0015] The PSO algorithm continuously performs iterative search according to the evaluation results until the optimal CPU / GPU mapping ratio is found.

[0016] In the said step S2, the implementation process of the PSO algorithm is as follows:

[0017] Step S2.1, Initialize the particle swarm

[0018] First, determine the number of particles. Each particle represents a CPU-GPU load allocation ratio value, and the value range is between 0 and 1. 0 means all tasks are executed by the CPU, and 1 means all tasks are executed by the GPU;

[0019] Randomly initialize the velocity for each particle. The velocity determines the moving direction and step size of the particle in the search space; set the initial position of each particle as its individual best position (pbest); at the same time, assign random values to the two parameters of the task completion rate and latency, and calculate the fitness value at this time; find the position and fitness value with the best fitness among all particles, and set them as the global best position (gbest) and the global best fitness;

[0020] Step S2.2, Construct the fitness function and calculate the fitness value Fitness;

[0021] Comprehensively considering the task completion rate and latency, construct the fitness function:

[0022] Fitness = ω 1 ×R + ω 2 / log(L + 1)

[0023] Among them, R is the task completion rate, and the calculation formula is as follows:

[0024] R = (N completed / N total ) / T

[0025] Among them, N completed is the number of tasks completed after the task execution time T under the current CPU-GPU load allocation ratio, and N total is the total number of tasks; the statistics of the number of completed tasks here are obtained through the CPU / GPU task execution records and the monitoring unit.

[0026] L is the latency time. The CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit use timestamps to record the task submission time and the time when the result of the completed task is obtained, and calculate the difference between the two to obtain the latency;

[0027] The calculation formula is as follows:

[0028] L = T result -T submit

[0029] Among them, T submit is the task submission time, and T result is the time when the result is obtained;

[0030] ω 1 and ω 2 are weight coefficients, and ω 1 + ω 2 = 1. The weight coefficients are custom-adjusted according to the actual requirements and characteristics of the task; the numerical values actually participating in the fitness function calculation are mapped according to the task type obtained by the CPU preprocessing unit;

[0031] Step S2.3, Update the individual best and global best

[0032] For each particle, compare its current fitness value with the individual best fitness. If the current fitness is better, update the individual best position (pbest) to the current position, and the individual best fitness to the current fitness value;

[0033] Then, compare the individual best fitness of all particles, find the best particle among them. If the individual best fitness of the found particle is better than the current global best fitness, update the global best position (gbest) and the global best fitness;

[0034] Step S2.4, Update the particle velocity and position

[0035] Update the particle velocity according to the basic formula of the PSO algorithm:

[0036]

[0037] Among them, is the velocity of particle i at the (k + 1)-th iteration, ω is the inertia weight with a value of 0.6, is the velocity of particle i at the k-th iteration; c 1 and c 2 are learning factors, both with a value of 1.5; r 1 and r 2 are random numbers with values greater than or equal to 0 and less than or equal to 1, pbest i is the individual optimal position of particle i, is the position of particle i at the k-th iteration, gbest is the global optimal position;

[0038] Update the particle position according to the updated velocity:

[0039]

[0040] Values exceeding 1 are set to 1, and values less than 0 are set to 0;

[0041] Step S2.5, Judge the termination condition

[0042] Judge whether the fitness value converges, that is, in consecutive multiple iterations, the change in the global optimal fitness value is less than a user-defined threshold (a very small threshold, such as 0.001). If the convergence condition is met, terminate the algorithm and output the optimization result.

[0043] Step S3, Task allocation and execution

[0044] After the FPGA chip calculates the optimized CPU / GPU mapping ratio, it returns the result to the scheduling and allocation unit in the CPU; the scheduling and allocation unit makes precise allocations of tasks between the CPU and the GPU based on the returned optimization result;

[0045] For the tasks allocated to the CPU, the CPU processes them according to its own execution logic and scheduling strategy;

[0046] For the tasks allocated to the GPU, the CPU sends the task data and relevant instructions to the GPU; after receiving the tasks, the GPU reasonably allocates the tasks to the idle SM (stream multiprocessors) and efficiently executes the tasks using its parallel computing ability;

[0047] In step S3, during the task execution, the CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit respectively record the execution status of tasks in the CPU and GPU in real time, including the amount of completed tasks and the execution progress of the current task, and simultaneously monitor the latency situation during the task execution, including data transfer latency and computing latency.

[0048] Step S4, Feedback and Adjustment

[0049] The CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit feed back the task completion status and latency data during the execution process to the FPGA chip. After receiving the feedback data, the FPGA chip makes real-time adjustments to the current PSO algorithm calculation;

[0050] If it is found that the current CPU / GPU mapping ratio results in a low task completion rate or high latency, the FPGA chip will adjust the search direction of the PSO algorithm, recalculate a more optimal CPU / GPU mapping ratio, and then return the result to the CPU again for adjustment of task allocation. This process repeats in a loop to achieve the dynamic optimization of the system.

[0051] A CPU-GPU heterogeneous system workload dynamic allocation system, including a CPU-GPU heterogeneous system and an FPGA chip;

[0052] The CPU system is provided with a preprocessing unit, a CPU task execution record and monitoring unit, and a scheduling and allocation unit;

[0053] The preprocessing unit is responsible for, after the CPU receives task data, performing preliminary parsing and preprocessing on the task data, extracting key feature information of the task, including the amount of task data, task priority, and the load conditions of the CPU and GPU in the current system, and determining the task type;

[0054] The CPU task execution record and monitoring unit is responsible for, during the task execution, recording the execution status of tasks in the CPU in real time, including the amount of completed tasks and the execution progress of the current task, and simultaneously monitoring the latency situation during the task execution, including data transfer latency and computing latency;

[0055] The scheduling and allocation unit is responsible for achieving precise allocation of tasks between the CPU and GPU according to the optimized CPU / GPU mapping ratio calculated by the FPGA chip;

[0056] The GPU system is provided with a GPU task execution record and monitoring unit;

[0057] The GPU task execution record and monitoring unit is responsible for recording in real time the execution status of tasks in the GPU during task execution, including the amount of completed tasks and the execution progress of the current task, and at the same time monitoring the latency situation during task execution, including data transfer latency and computing latency;

[0058] The FPGA chip is provided with a PSO algorithm hardware acceleration module, and the PSO algorithm hardware acceleration module includes a data storage unit, a computing unit, and an iteration number controller;

[0059] The data storage unit is responsible for storing particle information, task information, and intermediate result information;

[0060] Each particle represents a CPU / GPU mapping ratio in the PSO algorithm. The example information includes the current position value, speed value, and individual best position (pbest) and its corresponding fitness value; among them, the current position of the particle is the corresponding CPU / GPU mapping ratio value, and the value range is between 0 and 1; the speed value determines the moving direction and step size of the particle in the search space;

[0061] The data storage unit uses a register file to implement the storage of particle information. The task information includes the total task volume, the amount of completed tasks, and the type information of the task;

[0062] The data storage unit uses a random access memory RAM to implement the storage of task information.

[0063] The intermediate result information is the intermediate result generated during the calculation process of the PSO algorithm.

[0064] The data storage unit uses a first-in-first-out queue FO to store intermediate result information.

[0065] The computing unit is responsible for updating the calculation of the particle speed and position according to the PSO algorithm formula, and at the same time, based on the task completion rate and latency data, completing the calculation of the fitness value, providing core operation support for the iteration of the PSO algorithm, and finding the optimal CPU / GPU mapping ratio.

[0066] A CPU-GPU heterogeneous system workload dynamic allocation device, characterized by including:

[0067] One or more processors, one or more memories, and one or more programs, where one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.

[0068] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the method described above is implemented.

[0069] The beneficial effect of the present invention is that: the CPU-GPU heterogeneous system workload dynamic allocation method and system can adjust the workload in real time according to the dynamic characteristics of tasks and the real-time state of system resources, which not only improves the utilization rate of system resources, but also greatly improves the task processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0071] Attached Figure 1 is a schematic diagram of the system architecture of the CPU-GPU heterogeneous system workload dynamic allocation system of the present invention.

[0072] Attached Figure 2 is a schematic diagram of the PSO hardware acceleration method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] The CPU-GPU heterogeneous system workload dynamic allocation method includes the following steps:

[0075] Step S1, task data preprocessing

[0076] After receiving the task data, the CPU preliminarily parses and preprocesses the task data through a preprocessing unit, extracts the key feature information of the task, including the task data volume, task priority, and the load conditions of the CPU and GPU in the current system; at the same time, it judges the task type (whether it is a compute-intensive type, data-intensive type, or other type);

[0077] Step S2, data transfer and PSO algorithm calculation

[0078] After preprocessing, the CPU transfers task-related data to the FPGA chip. After receiving the data, the FPGA chip uses its internal hardware logic to implement the PSO algorithm, calculates in combination with the task completion rate and the latency model, and finds the optimal CPU / GPU mapping ratio;

[0079] In step S2, the FPGA chip evaluates the advantages and disadvantages of different CPU / GPU mapping ratio schemes according to the task completion rate and the latency model;

[0080] Among them, the task completion rate is calculated by counting the ratio of the number of tasks completed within a unit time to the total number of tasks; the latency is obtained by recording the time difference from task submission to completion;

[0081] The PSO algorithm continuously performs iterative search according to the evaluation results until the optimal CPU / GPU mapping ratio is found.

[0082] In step S2, the implementation process of the PSO algorithm is as follows:

[0083] Step S2.1, Initialize the particle swarm

[0084] First, determine the number of particles. Each particle represents a CPU-GPU load distribution ratio value, and the value range is between 0 and 1. 0 means all tasks are executed by the CPU, and 1 means all tasks are executed by the GPU;

[0085] Randomly initialize the velocity for each particle. The velocity determines the moving direction and step size of the particle in the search space; set the initial position of each particle as its personal best position (pbest); at the same time, assign random values to the two parameters of the task completion rate and the latency, and calculate the fitness value at this time; find the position and fitness value with the best fitness among all particles, and set them as the global best position (gbest) and the global best fitness;

[0086] Step S2.2, Construct the fitness function and calculate the fitness value Fitness

[0087] Comprehensively consider the task completion rate and the latency, and construct the fitness function:

[0088] Fitness = ω 1 ×R + ω 2 / log(L + 1)

[0089] Among them, R is the task completion rate, and the calculation formula is as follows:

[0090] R = (N completed / N total ) / T

[0091] Among them, N completedLet \(N\) be the number of tasks completed after task execution time \(T\) under the current CPU-GPU load distribution ratio. total Let \(N_{total}\) be the total number of tasks;

[0092] Let \(L\) be the latency time. The CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit use timestamps to record the task submission time and the time when the result is obtained after the task is completed, and calculate the difference between the two to get the latency.

[0093] The calculation formula is as follows:

[0094] \(L = T_{result}\) result - \(T_{submission}\) submit

[0095] where \(T_{submission}\) submit is the task submission time, and \(T_{result}\) result is the time when the result is obtained;

[0096] Let \(\omega_1\) 1 and \(\omega_2\) 2 be the weight coefficients, and \(\omega_1\) 1 + \(\omega_2\) 2 = 1. The weight coefficients are custom-adjusted according to the actual requirements and characteristics of the tasks; map the task types obtained by the CPU preprocessing unit to the values actually participating in the fitness function calculation;

[0097] For tasks with extremely high real-time requirements, \((\omega_1,\omega_2)=(0.3,0.7)\) to highlight the emphasis on latency; 1 , \(\omega_2\) 2 )=(0.3,0.7) to highlight the emphasis on latency;

[0098] For batch processing tasks, \((\omega_1,\omega_2)=(0.8,0.2)\) to emphasize the importance of the task completion rate. 1 , \(\omega_2\) 2 )=(0.8,0.2) to emphasize the importance of the task completion rate.

[0099] Step S2.3, Update the individual best and global best

[0100] For each particle, compare its current fitness value with the individual best fitness. If the current fitness is better, update the individual best position (pbest) to the current position, and the individual best fitness to the current fitness value;

[0101] Then, compare the individual best fitness of all particles, find the best particle among them. If the individual best fitness of the found particle is better than the current global best fitness, update the global best position (gbest) and the global best fitness;

[0102] Step S2.4, Update the particle velocity and position

[0103] Update the particle velocity according to the basic formula of the PSO algorithm:

[0104]

[0105] Among them, is the velocity of particle i at the (k + 1)-th iteration, ω is the inertia weight with a value of 0.6, is the velocity of particle i at the k-th iteration; c 1 and c 2 are learning factors, both with a value of 1.5; r 1 and r 2 are random numbers with values greater than or equal to 0 and less than or equal to 1, pbest i is the individual optimal position of particle i, is the position of particle i at the k-th iteration, and gbest is the global optimal position;

[0106] Update the particle position according to the updated velocity:

[0107]

[0108] Values exceeding 1 are set to 1, and values less than 0 are set to 0;

[0109] Step S2.5, judge the termination condition

[0110] Judge whether the fitness value converges, that is, in continuous multiple iterations, the change in the global optimal fitness value is less than a user-defined threshold (a very small threshold, such as 0.001). If the convergence condition is met, terminate the algorithm and output the optimization result.

[0111] Step S3, task allocation and execution

[0112] After the FPGA chip calculates the optimized CPU / GPU mapping ratio, it returns the result to the scheduling and allocation unit in the CPU; the scheduling and allocation unit makes an accurate allocation of tasks between the CPU and the GPU based on the returned optimization result;

[0113] For the tasks allocated to the CPU, the CPU processes them according to its own execution logic and scheduling strategy;

[0114] For the tasks allocated to the GPU, the CPU sends the task data and related instructions to the GPU; after receiving the tasks, the GPU reasonably allocates the tasks to the idle SM (stream multiprocessor) and efficiently executes the tasks using its parallel computing ability;

[0115] In step S3, during the task execution process, the CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit respectively record the execution status of tasks in the CPU and GPU in real time, including the amount of completed tasks and the execution progress of the current task. At the same time, the delay situation during the task execution process is monitored, including data transmission delay and calculation delay.

[0116] Step S4, Feedback and Adjustment

[0117] The CPU task execution record and monitoring unit and the GPU task execution record and monitoring unit feed back the task completion status and delay data during the execution process to the FPGA chip. After receiving the feedback data, the FPGA chip makes real-time adjustments to the current PSO algorithm calculation;

[0118] If it is found that the current CPU / GPU mapping ratio results in a low task completion rate or high delay, the FPGA chip will adjust the search direction of the PSO algorithm, recalculate a more optimal CPU / GPU mapping ratio, and then return the result to the CPU again for adjustment of task allocation. This process repeats in a loop to achieve the dynamic optimization of the system.

[0119] This CPU-GPU heterogeneous system workload dynamic allocation system includes a CPU-GPU heterogeneous system and an FPGA chip;

[0120] The CPU serves as the control core of the system, responsible for processing system-level management tasks, logical judgment, and some serial computing tasks. It has powerful complex instruction execution capabilities and good logical control capabilities; a preprocessing unit, a CPU task execution record and monitoring unit, and a scheduling and allocation unit are provided in the CPU system;

[0121] The preprocessing unit is responsible for, after the CPU receives task data, performing preliminary parsing and preprocessing on the task data, extracting key feature information of the task, including the amount of task data, task priority, and the load conditions of the CPU and GPU in the current system, and judging the task type (whether it is compute-intensive, data-intensive, or other types);

[0122] The CPU task execution record and monitoring unit is responsible for, during the task execution process, recording the execution status of tasks in the CPU in real time, including the amount of completed tasks and the execution progress of the current task. At the same time, the delay situation during the task execution process is monitored, including data transmission delay and calculation delay;

[0123] The scheduling and allocation unit is responsible for achieving precise allocation of tasks between the CPU and GPU according to the optimized CPU / GPU mapping ratio calculated by the FPGA chip;

[0124] The GPU focuses on large-scale data parallel computing tasks. With its large number of computing cores and highly parallel architecture, it can efficiently handle compute-intensive tasks. In the GPU system, there is a GPU task execution record and monitoring unit.

[0125] The GPU task execution record and monitoring unit is responsible for, during the task execution process, recording in real time the execution status of tasks in the GPU, including the amount of completed tasks and the execution progress of the current task, and at the same time monitoring the latency situation during the task execution process, including data transfer latency and computing latency.

[0126] The FPGA chip plays a key role in acceleration and flexible resource allocation. It can achieve the accelerated execution of specific algorithms through hardware programming, especially playing an important role in optimizing the task allocation between the CPU and the GPU. In the FPGA chip, there is a PSO algorithm hardware acceleration module, and the PSO algorithm hardware acceleration module includes a data storage unit, a computing unit, and an iteration number controller.

[0127] The data storage unit is responsible for storing particle information, task information, and intermediate result information.

[0128] Each particle represents a CPU / GPU mapping ratio in the PSO algorithm. The example information includes the current position value, speed value, and individual best position (pbest) and its corresponding fitness value. Among them, the current position of the particle is the corresponding CPU / GPU mapping ratio value, and the value range is between 0 and 1. The speed value determines the moving direction and step size of the particle in the search space.

[0129] The above particle information is crucial for the update of particles and the convergence of the algorithm.

[0130] The data storage unit uses a register file to implement the storage of particle information. The register file has the characteristics of high-speed reading and writing and is suitable for storing data that is frequently accessed.

[0131] The task information includes the total task amount, the amount of completed tasks, and the type information of the task.

[0132] The task completion rate is calculated using the total task amount and the amount of completed tasks. Tasks are classified into compute-intensive tasks or data-intensive tasks according to their types.

[0133] The data storage unit uses a random access memory (RAM) to implement the storage of task information.

[0134] The intermediate result information is the intermediate result generated during the calculation process of the PSO algorithm.

[0135] The data storage unit uses a FIFO (First In First Out) queue to store intermediate result information. FIFO is suitable for scenarios where data needs to be processed in order and can ensure the order of data.

[0136] The computing unit is responsible for updating the calculation of particle velocity and position according to the PSO algorithm formula. At the same time, based on the task completion rate and latency data, it calculates the fitness value to provide core operation support for the iteration of the PSO algorithm and finds the optimal CPU / GPU mapping ratio.

[0137] During the calculation process, the PSO algorithm hardware acceleration module coordinates the data transfer order between the data storage unit and the computing unit. When calculating the particle velocity update, it is necessary to first read data such as the current position, velocity, individual best position, and global best position of the particle from the data storage unit, and then input these data into the computing unit of the computing unit in sequence according to the calculation order. After the calculation is completed, the control unit controls the newly calculated velocity and position values to be written back to the corresponding positions in the data storage unit. At the same time, when calculating the fitness value, it is also necessary to read task-related data and intermediate results in a similar order, and store the fitness value in the data storage unit for subsequent comparison and update operations.

[0138] The CPU-GPU heterogeneous system workload dynamic allocation device includes:

[0139] One or more processors, one or more memories, and one or more programs, where one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the methods in any of the above methods.

[0140] The readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method as described above.

[0141] The above embodiments are only one of the specific implementation manners of the present invention, and the common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A CPU-GPU heterogeneous system workload dynamic allocation method, characterized by: Step S1: Task data preprocessing After receiving the task data, the CPU performs preliminary analysis and preprocessing on the task data through the preprocessing unit to extract key feature information of the task, including the task data volume, task priority, and the load of the CPU and GPU in the current system; at the same time, the task type is determined; Step S2: Data transmission and PSO algorithm calculation After preprocessing, the CPU passes the task-related data to the FPGA chip. After receiving the data, the FPGA chip uses its internal hardware logic to implement the PSO algorithm, combining the task completion rate and delay model for calculation to find the optimal CPU / GPU mapping ratio; Step S3: Task allocation and execution After the FPGA chip calculates the optimized CPU / GPU mapping ratio, it returns the result to the scheduling and allocation unit in the CPU; the scheduling and allocation unit accurately allocates tasks between the CPU and GPU based on the returned optimization result; For tasks assigned to the CPU, the CPU processes them according to its own execution logic and scheduling strategy; For tasks assigned to the GPU, the CPU sends the task data and related instructions to the GPU. After receiving the task, the GPU reasonably assigns the task to the idle streaming multiprocessor SM and uses its parallel computing capability to efficiently execute the task. Step S4: Feedback and adjustment The CPU task execution recording and monitoring unit and the GPU task execution recording and monitoring unit feed back the task completion status and delay data during the execution process to the FPGA chip. After receiving the feedback data, the FPGA chip makes real-time adjustments to the current PSO algorithm calculation; If it is found that the current CPU / GPU mapping ratio results in a low task completion rate or too high a delay, the FPGA chip will adjust the search direction of the PSO algorithm, recalculate a better CPU / GPU mapping ratio, and then return the result to the CPU to adjust the task allocation. This cycle repeats to achieve dynamic optimization of the system.

2. The CPU-GPU heterogeneous system workload dynamic allocation method according to claim 1, characterized in that: In step S2, the FPGA chip evaluates the pros and cons of different CPU / GPU mapping ratio schemes according to the task completion rate and the delay model; The task completion rate is calculated by counting the ratio of the number of completed tasks to the total number of tasks in a unit of time; the delay is obtained by recording the time difference from task submission to completion; The PSO algorithm continuously performs iterative search based on the evaluation results until the optimal CPU / GPU mapping ratio is found.

3. The CPU-GPU heterogeneous system workload dynamic allocation method according to claim 2, characterized in that: In step S2, the PSO algorithm implementation process is as follows: Step S2.1: Initialize particle swarm First, determine the number of particles. Each particle represents a CPU-GPU load distribution ratio value, ranging from 0 to 1. 0 means that all tasks are executed by the CPU, and 1 means that all tasks are executed by the GPU. Initialize the speed of each particle randomly, and the speed determines the moving direction and step size of the particle in the search space; set the initial position of each particle to its individual optimal position pbest; assign random values ​​to the task completion rate and delay parameters, and calculate the fitness value at this time; find the position and fitness value with the best fitness among all particles, and set them as the global optimal position gbest and the global optimal fitness; Step S2.2, construct a fitness function and calculate the fitness value Fitness; Considering the task completion rate and delay comprehensively, the fitness function is constructed: Fitness=ω1×R+ω2 / log(L+1) Among them, R is the task completion rate, and the calculation formula is as follows: R=(N completed / N total ) / T Among them, N completed N is the number of completed tasks after task execution time T under the current CPU-GPU load distribution ratio, total is the total number of tasks; L is the delay time. The CPU task execution recording and monitoring unit and the GPU task execution recording and monitoring unit use timestamps to record the task submission time and the time to complete the task and get the result. The difference between the two is calculated to get the delay. The calculation formula is as follows: L=T result -T submit Among them, T submit is the task submission time, T result Time to get results; ω1 and ω2 are weight coefficients, and ω1+ω2=1. The weight coefficients are customized according to the actual needs and characteristics of the task; the value actually involved in the fitness function calculation is mapped according to the task type obtained by the CPU preprocessing unit; step S2.3, update the individual optimum and the global optimum For each particle, compare its current fitness value with the individual optimal fitness. If the current fitness is better, update the individual optimal position pbest to the current position, and the individual optimal fitness to the current fitness value. Then, compare the individual optimal fitness of all particles to find the best particle. If the individual optimal fitness of the found particle is better than the current global optimal fitness, update the global optimal position gbest and the global optimal fitness. Step S2.4: Update particle speed and position Update the particle velocity according to the basic formula of the PSO algorithm: in, is the velocity of particle i at the (k+1)th iteration, ω is the inertia weight with a value of 0.6, is the velocity of particle i in the kth iteration; c1 and c2 are learning factors, both of which are 1.5; r1 and r2 are random numbers with values ​​greater than or equal to 0 and a sum less than or equal to 1, pbest i is the individual optimal position of particle i, is the position of particle i at the kth iteration, gbest is the global optimal position; Update the particle position according to the updated velocity: Values ​​greater than 1 are set to 1, and values ​​less than 0 are set to 0; Step S2.5: Determine the termination condition Determine whether the fitness value converges, that is, in multiple consecutive iterations, the change in the global optimal fitness value is less than the custom threshold. If the convergence condition is met, the algorithm is terminated and the optimization result is output.

4. The CPU-GPU heterogeneous system workload dynamic allocation method according to claim 1, characterized in that: In step S3, during the task execution process, the CPU task execution recording and monitoring unit and the GPU task execution recording and monitoring unit respectively record the execution status of the tasks in the CPU and GPU in real time, including the amount of completed tasks and the execution progress of the current task, and monitor the delays in the task execution process, including data transmission delays and calculation delays.

5. A CPU-GPU heterogeneous system workload dynamic allocation system, characterized by: Including CPU-GPU heterogeneous systems and FPGA chips; The CPU system is equipped with a pre-processing unit, a CPU task execution recording and monitoring unit, and a scheduling and allocation unit; The preprocessing unit is responsible for performing preliminary analysis and preprocessing on the task data after the CPU receives the task data, extracting key feature information of the task, including the task data volume, task priority, and the load of the CPU and GPU in the current system, and determining the task type; The CPU task execution recording and monitoring unit is responsible for recording the execution status of tasks in the CPU in real time during the task execution process, including the amount of completed tasks and the execution progress of the current task, and monitoring the delays in the task execution process, including data transmission delays and calculation delays; The scheduling and allocation unit realizes accurate allocation of tasks between the CPU and the GPU according to the optimized CPU / GPU mapping ratio calculated by the FPGA chip; The GPU system is provided with a GPU task execution recording and monitoring unit; The GPU task execution recording and monitoring unit is responsible for recording the execution status of tasks in the GPU in real time during the task execution process, including the amount of completed tasks and the execution progress of the current task, and monitoring the delay during the task execution process, including data transmission delay and calculation delay; The FPGA chip is provided with a PSO algorithm hardware acceleration module, which includes a data storage unit, a calculation unit and an iteration number controller; The data storage unit is responsible for storing particle information, task information and intermediate result information; Each particle represents a CPU / GPU mapping ratio in the PSO algorithm. The example information includes the current position value, speed value, individual optimal position pbest and its corresponding fitness value. The current position of the particle is the corresponding CPU / GPU mapping ratio value, which ranges from 0 to 1. The speed value determines the moving direction and step length of the particle in the search space. The task information includes the total task amount, the completed task amount and the task type information; The intermediate result information is the intermediate result generated during the calculation process of the PSO algorithm; The calculation unit is responsible for updating the particle speed and position according to the PSO algorithm formula, and calculating the fitness value based on the task completion rate and delay data, providing core computing support for the iteration of the PSO algorithm and finding the optimal CPU / GPU mapping ratio.

6. The CPU-GPU heterogeneous system workload dynamic allocation system according to claim 5, characterized in that: The data storage unit uses a register file to store particle information.

7. The CPU-GPU heterogeneous system workload dynamic allocation system according to claim 5, characterized in that: The data storage unit uses a random access memory (RAM) to store task information.

8. The CPU-GPU heterogeneous system workload dynamic allocation system according to claim 5, characterized in that: The data storage unit uses a first-in-first-out queue FO to store intermediate result information.

9. A CPU-GPU heterogeneous system workload dynamic allocation device, characterized by: include: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 4.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • A load balancing method and apparatus based on CPU-GPU

    CN109213601A

  • Spark-based adaptive task scheduling method aiming at heterogeneous environment

    CN109376012A

  • Automatic tuning load balancing method for heterogeneous computing system

    CN118034909A

  • Calculation task scheduling method and device in multivariate heterogeneous environment and medium

    CN119415240A

  • Dynamically provisioning and scaling graphic processing units for data analytic workloads in a hardware cloud

    US20170293994A1

Cited By

  • Hardware acceleration equipment management system and method applied to network security product

    CN120909969A