DNN reasoning dynamic partitioning and task scheduling method based on GA-DQN collaboration

Through the GA-DQN collaborative optimization method, dynamic partitioning and server selection, the computing resource and scheduling efficiency problems of DNN inference tasks in multi-user and multi-server environments are solved, and efficient task completion and accuracy improvement are achieved, which is suitable for edge computing scenarios.

CN120276816APending Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510283566.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the complex edge environment of multi-user and multi-server, the efficient execution of DNN inference tasks faces the problems of limited computing resources and low task scheduling efficiency. Especially in the multi-user concurrency scenario, traditional scheduling algorithms are difficult to meet the requirements of task time limit and resource utilization.

Method used

The deep Q network (GA-DQN) collaborative optimization method based on genetic algorithm enhancement is adopted. Through dynamic partitioning, early exit mechanism and server selection, the optimal partition point, early exit decision and scheduling decisions are formulated, the Markov decision-making process is established, and the GA-DQN collaborative optimization strategy is used to solve complex optimization problems in high-dimensional decision space.

Benefits of technology

It significantly improves the task completion rate and inference accuracy, especially suitable for edge computing scenarios with high requirements for real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276816A_ABST
    Figure CN120276816A_ABST
Patent Text Reader

Abstract

A DNN reasoning dynamic partitioning and task scheduling method based on GA-DQN collaboration comprises the steps that a trained multi-outlet DNN model is deployed for each device in a multi-user and multi-server scene to conduct task reasoning, and in a time slot, tasks reach user devices, meet Poisson distribution and are stored in task queues of the user devices; establishing an optimization problem P aiming at maximizing the total completion rate and reasoning accuracy of task reasoning in a time slot by taking a strategy feasibility interval as a constraint and taking a partition point of the task, advanced exit selection, server selection and a scheduling decision as optimization variables; and establishing a Markov decision process, and solving a problem P by using a GA-DQN collaborative optimization strategy. The method is suitable for dynamic DNN reasoning partitioning and scheduling of multiple users and multiple servers, and the task completion rate and the reasoning accuracy within a period of time can be effectively increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge intelligence, and is a method for dynamically partitioning and task scheduling of Deep Neural Network (DNN) inference based on an early exit mechanism in a multi-user and multi-server scenario. This method formulates near-optimal partition point selection, early exit decision-making, server selection, and scheduling decisions through a collaborative optimization method based on a Genetic Algorithm-enhanced Deep Q-Network (GA-DQN), thereby maximizing the task completion rate and inference accuracy. Background Art

[0002] With the deep integration of edge computing and artificial intelligence technologies, edge intelligence has gradually become a key technology for realizing low-latency and high-efficiency real-time data processing. However, in the complex edge environment of multi-users and multi-servers, the efficient execution of DNN inference tasks faces severe challenges. On the one hand, the computing resources of edge devices are limited and it is difficult to independently complete large-scale DNN inference tasks; on the other hand, in the multi-user concurrent scenario, the task scheduling efficiency and the dynamic nature of resource allocation have become the main bottlenecks restricting system performance. To address these challenges, researchers have proposed various optimization strategies, such as offloading some DNN inference tasks to edge servers, designing task partitioning models, and introducing early exit mechanisms. These methods have alleviated the computing pressure on edge devices to a certain extent and improved the inference efficiency. However, most existing studies focus on single optimization strategies and lack in-depth exploration of multi-objective collaborative optimization. Especially in the multi-server and multi-user scenario, how to achieve the collaborative optimization of task partitioning, early exit mechanism, server selection, and task scheduling remains an urgent problem to be solved.

[0003] In recent years, task partitioning models and early exit mechanisms have become important research directions for optimizing DNN inference performance. The task partitioning model realizes the flexible allocation of computing resources by intelligently splitting DNN inference tasks between local devices and edge servers; the early exit mechanism reduces unnecessary computational volume by setting multiple intermediate exits in the DNN architecture and dynamically adjusting the inference depth according to task complexity. However, with the surge in DNN computing requirements, especially in the multi-user and multi-task concurrent scenario, edge servers need to process a large number of offloaded tasks simultaneously, and traditional scheduling algorithms are difficult to meet the requirements of task deadlines and resource utilization. Therefore, designing an efficient optimization strategy by combining task partitioning, early exit mechanism, and server task scheduling is an important research topic in the field of edge intelligence. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention proposes an optimization method for dynamically partitioning and scheduling DNN inference tasks in a multi-user and multi-server scenario to maximize the task completion rate and inference accuracy. First, the present invention considers a system scenario consisting of multiple user devices and multiple edge servers, adopts a partitioning model and an early-exit architecture model, and combines server allocation and task scheduling to model DNN task inference. This model is represented as an optimization problem of maximizing the completion rate and inference accuracy of all tasks. Then, the present invention proposes a multi-objective collaborative strategy of GA-DQN. Among them, DQN is responsible for learning and predicting the optimal task partitioning, early-exit selection, and server selection decisions in the current state; GA optimizes the task queue within the server and quickly generates an optimal scheduling scheme through genetic operations based on a priority mechanism. The two achieve efficient collaboration through a feedback mechanism, thereby maximizing the task completion rate while ensuring the inference accuracy.

[0005] To implement the above process, the present invention provides the following technical solutions:

[0006] A DNN inference dynamic partitioning and task scheduling method based on GA-DQN collaboration deploys a trained multi-exit DNN model for each device in a multi-user and multi-server scenario for task inference. During a time slot, tasks arriving at user devices follow a Poisson distribution and are stored in the task queues of user devices. Constrained by the policy feasibility interval and with the partitioning points, early-exit selection, server selection, and scheduling decisions of tasks as optimization variables, an optimization problem P is established with the goal of maximizing the total completion rate and inference accuracy of task inference within a time slot. A Markov decision process is established, and the GA-DQN collaborative optimization strategy is used to solve problem P.

[0007] Furthermore, the method includes the following steps:

[0008] Step 1. The system consists of N user devices UE and S edge servers connected to the base station with a capacity of M server composed of, represents the set of UEs, S represents the set of servers. On each user device and edge server, a trained DNN model is deployed for performing DNN task inference; use cd i represents the communication distance limit of user device i. Task inference is divided into two parts. The first half of the partitioning point is calculated locally by the user device, and the second half is offloaded to the server within the user's communication distance for calculation;

[0009] Step 2: The task selects a partition at a fixed layer preset in the DNN network. At each partition point except the 0th layer, a classifier composed of a fully connected layer and a softmax layer is constructed as an exit for early exit inference. A threshold is set at the exit. If the maximum confidence value output by the softmax layer of the task at the exit is greater than its threshold, the task can select this exit to early exit inference and return the prediction result. Let L denote the total number of layers of the DNN model;

[0010] Step 3: During time slot t, the tasks arriving at the user equipment follow a Poisson distribution. The number of tasks arriving at user equipment i is denoted as M i (t). The jth task of user equipment i is denoted as task ij ={d ij , τ ij , dl ij}, where d ij is the size of the task input data, τ ij is the time when the task arrives at the device, and dl ij is the task deadline; The partition selection strategy for each task is represented by , and the decision of whether to exit is represented by α ij ∈{0, 1}, and the inference accuracy of the exit at the partition point is represented by ;

[0011] Step 4: After each user equipment i receives M i (t) tasks, it decides whether the task exits the inference early. If α ij = 1, the user equipment directly gives the task inference result and exits, without needing to offload it to the server for calculation after partitioning; Otherwise, the part of the task after the partition point is offloaded to the server within its communication range for further task scheduling and processing. Let S i denote the set of servers that can be offloaded within the communication range of device i. The server selection decision for each task is represented by ;

[0012] Step 5: According to the server selection decisions of all tasks, each server obtains the set of tasks scheduled for execution by itself. The task scheduling decision on the server is represented by the scheduling order ;

[0013] Step 6: Calculate the completion status of the task during the calculation process: The completion delay of the task is represented by . If the completion time of the task execution does not exceed its deadline, it is considered completed; otherwise, it is considered uncompleted. Then the task completion status can be expressed as where 1(x) is the indicator function, and the value of x is 1 when it satisfies the condition, otherwise the value is 0;

[0014] Step 7: Taking the partition point selection decision of the task, early exit decision, server selection decision, and server scheduling decision as optimization variables, and maximizing the task completion rate and inference accuracy rate within time slot t as the optimization objective, an optimization problem is established. Let ω1 and ω2 represent the weights of the completion rate and accuracy rate, and its mathematical model is as follows:

[0015]

[0016] Constraint ① is the range constraint of the partition point, constraints ② and ③ are the range and condition constraints of the early exit decision, constraints ④ and ⑤ are the interval and condition constraints of server selection, and constraint ⑥ is the server task capacity constraint;

[0017] Step 8: Model the problem P as a Markov decision process. Let s i (t) = {d i (t), τ i (t), dl i (t), c i (t)} represent the state of agent i at the t-th time slot, where d i (t) is the set of input data sizes of the tasks of device i, τ i (t) is the set of arrival times of the tasks, dl i (t) is the set of deadline times of the tasks, and c i (t) is the computing power of device i; Let represent the action of agent i at the t-th time slot, and the reward of agent i at the t-th time slot is expressed as

[0018] Step 9: Establish a solution process based on a priority-based genetic algorithm (GA). Let M s represent the number of task schedules of server s. In GA, each individual is represented by a chromosome for the task scheduling scheme on the server, and the chromosome is encoded as a vector where t k represents the number of the k-th task, and the order on the chromosome corresponds to the scheduling order of the tasks;

[0019] Step 10: Initialize the genetic algorithm population. The priority of the task on the server is expressed as Let Randomly generate the task numbers higher than the threshold at the front positions of the chromosome;

[0020] Step 11: Use F i tness(g) = aR(g) - bP(g) to represent the fitness function based on the reward and punishment mechanism, where represents the score for successful scheduling of the task, It represents the penalty score when the task is not completed. σ is a parameter controlling the penalty strength, and a and b are weight parameters;

[0021] Step 12: Calculate the fitness value of each individual according to the fitness function, perform selection, crossover, mutation, and update the population. Execute this step until the number of iterations is reached, and the optimal task scheduling order can be obtained;

[0022] Step 13: Solve the optimal policy through the method of Deep Q-Network enhanced by Genetic Algorithm (GA-DQN) to maximize the completion rate and inference accuracy. Agent i obtains the current state s i (t) from the environment and selects the action a i (t) according to the ε-greedy policy;

[0023] Step 14: According to the action decisions of all agents The server obtains the set of tasks to be scheduled, executes Step 9, Step 10, Step 11, and Step 12, and obtains the optimal scheduling order of the server. Calculate the delay of the task on the server according to the scheduling order, and thus calculate the task completion situation TC ij ;

[0024] Step 15: Calculate the reward r i (t), and the environment feedbacks the reward r i (t) and the next state s i (t + 1) to Agent i. Agent i stores the transition (s i (t), a i (t), r i (t), s i (t + 1)) into the experience pool and samples from the experience pool to update the parameters of the network Q;

[0025] Step 16: The GA algorithm gives the optimal scheduling decision for the server task queue, and the trained Q network outputs the global optimal decisions of partition point selection, whether to exit early, and server selection.

[0026] The beneficial effects of the present invention are mainly manifested in: For the DNN inference task in the multi-user multi-server scenario, it comprehensively considers the joint optimization problems of dynamic partitioning, early exit mechanism, server allocation, and task scheduling, and proposes a collaborative optimization strategy based on the Deep Q-Network enhanced by Genetic Algorithm, effectively solving the complex optimization problem in the high-dimensional decision space. The present invention can significantly improve the task completion rate on the premise of ensuring the inference accuracy, and is especially suitable for the edge computing scenario with high requirements for real-time performance and accuracy. Description of the Drawings

[0027] Figure 1 It is the user-server collaboration model for the DNN inference task.

[0028] Figure 2 It is a structural model for multi - exit DNN partitioning.

[0029] Figure 3 It is a flowchart for optimizing server task scheduling based on the genetic algorithm.

[0030] Figure 4 It is a flowchart of the deep Q - network algorithm based on the genetic algorithm. Detailed implementation manner

[0031] The present invention will be further described below with reference to the accompanying drawings.

[0032] Refer to Figure 1 、 Figure 2 、 Figure 3 and Figure 4 A DNN inference dynamic partitioning and task scheduling method based on GA - DQN collaboration includes the following steps:

[0033] Step 1: In the scenario model as Figure 1 , the system consists of N user devices UE and S edge servers connected to the base station with a capacity of M server . represents the set of UEs, S represents the set of servers. On each user device and edge server, a trained DNN model is deployed to perform DNN task inference; use cd i to represent the communication distance limit of user device i. Task inference is divided into two parts. The first half of the division point is calculated locally by the user device, and the second half is offloaded to the server within the user's communication distance for calculation;

[0034] Step 2: The task selects a partition at the fixed layer preset in the DNN network. At each partition point except the 0th layer, a classifier composed of a fully - connected layer and a softmax layer is constructed as an exit for early - termination inference. As shown in the structural model of Figure 2 , a threshold is set at the exit. If the maximum confidence value output by the softmax layer of the task at the exit is greater than its threshold, the task can select this exit to terminate the inference early and return the prediction result. Use L to represent the total number of layers of the DNN model;

[0035] Step 3: In time slot t, the tasks arriving at the user device follow a Poisson distribution. The number of tasks arriving at user device i is represented as M i (t). The j - th task of user device i is represented as task ij = {d ij , τ ij , dl ij} where d ij is the size of the task input data, τ ijis the time when the task arrives at the device, \(d_l\) ij is the task deadline; the partitioning selection strategy for each task is determined by Whether to exit the decision is indicated by \(\alpha\) ij \(\in \{0, 1\}\), and the inference accuracy rate at the exit at the partitioning point is represented by ;

[0036] Step 4. After each user device \(i\) receives \(M\) i (t) tasks, it decides whether to exit the inference of the task in advance. If \(\alpha\) ij = 1, the user device directly gives the task inference result and exits, without the need to unload it to the server for calculation after partitioning; otherwise, the part of the task after the partitioning point is unloaded to the server within its communication range for further task scheduling processing, and \(S\) i represents the set of servers that can be unloaded within the communication range of device \(i\). The server selection decision for each task is determined by ;

[0037] Step 5. According to the server selection decisions of all tasks, each server obtains the set of tasks scheduled and executed by itself. The task scheduling decision on the server is represented by the scheduling order ;

[0038] Step 6. Calculate the completion status of the tasks during the calculation process: the completion delay of the task is represented by . If the completion time of the task execution does not exceed its deadline, it is considered completed; otherwise, it is considered uncompleted. Then the task completion status can be expressed as where \(1(x)\) is the indicator function, \(x\) satisfies the condition and the value is 1, otherwise the value is 0;

[0039] Step 7. Taking the partitioning point selection decision, the early exit decision, the server selection decision, and the server scheduling decision of the task as optimization variables, and maximizing the task completion rate and inference accuracy rate within time slot \(t\) as the optimization objective, an optimization problem is established. Using \(\omega_1\) and \(\omega_2\) to represent the weights of the completion rate and accuracy rate, its mathematical model is:

[0040]

[0041]

[0042] Constraint ① is the range constraint of the partitioning point, Constraint ②③ are the range and condition constraints of the early exit decision, Constraint ④⑤ are the interval and condition constraints of the server selection, and Constraint ⑥ is the server task capacity constraint;

[0043] Step 8. Model the problem \(P\) as a Markov decision process, using \(s\) i (t)=\{d i (t),\tau i(t), dl i (t), c i (t) represents the state of the t-th time slot agent i, where d i (t) is the set of input data sizes of the tasks of device i, τ i (t) is the set of arrival times of the tasks, dl i (t) is the set of deadlines of the tasks, c i (t) is the computing power of device i; use to represent the action of the t-th time slot agent i, and the reward of the t-th time slot agent i is expressed as

[0044] Step 9, establish a priority-based genetic algorithm (GA) solution process, and the solution process is as Figure 3 shown. Use M s to represent the number of task schedules of server s. Each individual in GA is represented by a chromosome for the task scheduling scheme on the server, and the chromosome is encoded as a vector where, t k represents the number of the k-th task, and the order on the chromosome corresponds to the scheduling order of the task;

[0045] Step 10, initialize the genetic algorithm population. The priority of the task on the server is expressed as Put the task numbers higher than the threshold are randomly generated at the front positions of the chromosome;

[0046] Step 11, use Fitness(g) = aR(g) - bP(g) to represent the fitness function based on the reward and punishment mechanism, where represents the score when the task is successfully scheduled, represents the penalty score when the task is not completed. σ is a parameter to control the penalty intensity, and a, b are weight parameters;

[0047] Step 12, calculate the fitness value of each individual according to the fitness function, perform selection, crossover, mutation and update the population. Execute this step until the number of iterations is reached, and the optimal task scheduling order can be obtained;

[0048] Step 13, solve the optimal strategy through the genetic algorithm enhanced deep Q network (GA-DQN) method to maximize the completion rate and inference accuracy, and the process is as Figure 4 shown. Agent i obtains the current state s i (t) from the environment, and selects the action a i (t) according to the ε-greedy strategy;

[0049] Step 14, according to the action decisions of all agents The server obtains the set of tasks to be scheduled, executes Step 9, Step 10, Step 11, and Step 12 to obtain the optimal scheduling order of the server. Calculate the latency of the tasks on the server according to the scheduling order, and thus calculate the task completion situation TC ij ;

[0050] Step 15. Calculate the reward r i (t), and the environment feeds back the reward r i (t) and the next state s i (t + 1) to the agent i. The agent i stores the transition (s i (t), a i (t), r i (t), s i (t + 1)) into the experience pool and samples from the experience pool to update the parameters of the network Q;

[0051] Step 16. The GA algorithm gives the optimal scheduling decision of the server task queue, and the trained Q network outputs the global optimal decisions of partition point selection, whether to exit early, and server selection.

[0052] The solution of this embodiment aims at the DNN inference tasks in the multi - user and multi - server scenario, comprehensively considers the joint optimization problems of dynamic partitioning, early - exit mechanism, server allocation, and task scheduling, and proposes a collaborative optimization strategy based on a genetic - algorithm - enhanced deep Q - network, effectively solving the complex optimization problems in the high - dimensional decision space. The present invention can significantly improve the task completion rate on the premise of ensuring the inference accuracy, and is especially suitable for edge - computing scenarios with high requirements for real - time performance and accuracy.

[0053] For Figure 1 the scenario diagram of multiple user devices collaborating with multiple edge servers to process tasks as shown

[0054] First, there are N user devices and S edge servers in the system. On each user device and edge server, a trained DNN model is deployed to perform DNN task inference. During a time slot, the arrival of tasks at the user devices follows a Poisson distribution. Considering the communication distance of the user devices, the task inference is divided into two parts. The first half of the division point is calculated locally by the user devices, and the second half is offloaded to the servers within the communication distance of the users for calculation.

[0055] Secondly, after each user device receives a task, it decides whether to exit the inference early. If the decision of the task is to exit early, the user device directly gives the task inference result and exits, without offloading it to the server for post - partition calculation; otherwise, it selects a server within its communication range and offloads the part after the task partition point to the server for further task scheduling and processing.

[0056] Again, with the strategic feasibility interval as the constraint, and the decisions on the partition point selection, early exit, server selection, and scheduling of tasks as the optimization variables, an optimization problem P is established with the goal of maximizing the total completion rate and inference accuracy of task inference within a time slot.

[0057] Finally, the problem is formulated as a Markov decision process, and GA-DQN is used to solve the optimal policy. The optimal decisions on task partitioning, early exit point selection, and server allocation are achieved through DQN. At the same time, the server task queue is scheduled and optimized using priority-based GA, and the two achieve efficient cooperation through a feedback mechanism.

[0058] The present invention addresses the scenario of multi-user devices and multi-servers, comprehensively considering the joint optimization problems of dynamic partitioning of DNN inference tasks, early exit mechanisms, server allocation, and task scheduling, and can dynamically formulate efficient optimization strategies to maximize the task completion rate and inference accuracy.

[0059] The content described in the embodiments of this specification is only a list of the implementation forms of the inventive concept and is for illustrative purposes only. The protection scope of the present invention should not be regarded as limited to the specific forms stated in these embodiments, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept.

Claims

1. A DNN inference dynamic partitioning and task scheduling method based on GA-DQN collaboration, characterized in that Deploy the trained multi-exit DNN model for task inference for each device in a multi-user and multi-server scenario. During a time slot, tasks arriving at the user device follow a Poisson distribution and are stored in the task queue of the user device; With the policy feasibility interval as the constraint, and the partition point, early exit selection, server selection, and scheduling decision of the task as the optimization variables, establish an optimization problem P with the goal of maximizing the total completion rate and inference accuracy of task inference within a time slot; establish a Markov decision process, and use a deep Q network based on a genetic algorithm to solve problem P through collaborative optimization strategies.

2. The DNN inference dynamic partitioning and task scheduling method based on GA-DQN collaboration according to claim 1, wherein The method includes the following steps: Step 1. The system consists of N user equipments UE and S edge servers connected to the base station with a capacity of . represents the set of UEs, represents the set of servers. On each user equipment and edge server, a trained DNN model is deployed to perform DNN task inference; Let cd i represent the communication distance limit of user equipment i. The task inference is divided into two parts. The first half of the division point is calculated locally by the user equipment, and the second half is offloaded to the server within the user's communication distance for calculation; Step 2: The task selects a partition at a fixed layer preset in the DNN network. At each partition point except the 0th layer, a classifier composed of a fully connected layer and a softmax layer is constructed as an exit for early exit inference. A threshold is set at the exit. If the maximum confidence value output by the softmax layer of the task at the exit is greater than its threshold, the task can select this exit to early exit the inference and return the prediction result. Let L represent the total number of layers of the DNN model; Step 3. During time slot t, the tasks arriving at the user equipment follow a Poisson distribution. The number of tasks arriving at user equipment i is denoted as M i (t). The j-th task of user equipment i is denoted as task ij ={d ij , τ ij , dl ij}, where d ij is the size of the task input data, τ ij is the time when the task arrives at the device, and dl ij is the task deadline; the partition selection strategy for each task is represented by . Whether to exit the decision is represented by α ij ∈{0, 1}, and the inference accuracy at the exit at the partition point is represented by . Step 4. After each user device i receives M i (t) tasks, it decides whether to prematurely exit the inference of the task. If α ij = 1, the user device directly gives the task inference result and exits, without offloading it to the server for the calculation after partitioning; otherwise, it offloads the part after the task partitioning point to the server within its communication range to perform further task scheduling processing, using to represent the set of servers that can be offloaded within the communication range of device i. The server selection decision for each task is represented by ; Step 5. According to the server selection decisions of all tasks, each server obtains its respective set of tasks to be scheduled and executed, and the task scheduling decision on the server is represented by the scheduling order indicated; Step 6. Calculate the completion status of tasks during the calculation process: The completion delay of a task is represented by . If the completion time of task execution does not exceed its deadline, it is considered completed; otherwise, it is considered uncompleted. Then, the task completion status can be expressed as where 1(x) is an indicator function, with a value of 1 when x satisfies the condition and 0 otherwise; Step 7: With the partition point selection decision, early exit decision, server selection decision, and server scheduling decision of the task as the optimization variables, establish an optimization problem with maximizing the task completion rate and inference accuracy within time slot t as the optimization goal. Use ω1 and ω2 to represent the weights of the completion rate and accuracy. Its mathematical model is: Constraint ① is the range constraint of the partition point, constraints ② and ③ are the range and condition constraints of the early exit decision, constraints ④ and ⑤ are the interval and condition constraints of server selection, and constraint ⑥ is the server task capacity constraint; Step 8, model the problem P as a Markov decision process, and use s i (t) = {d i (t), τ i (t), dl i (t), c i (t)} to represent the state of agent i at the t-th time slot, where d i (t) is the set of input data sizes of the tasks of device i, τ i (t) is the set of arrival times of the tasks, dl i (t) is the set of deadline times of the tasks, and c i (t) is the computing power of device i; use to represent the action of agent i at the t-th time slot, and the reward of agent i at the t-th time slot is expressed as Step 9: Establish a priority-based genetic algorithm (GA) solution process. Use M s to represent the number of task schedules of server s. Each individual in GA is represented by a chromosome for the task scheduling scheme on the server, and the chromosome is encoded as a vector where t k represents the number of the k-th task, and the order on the chromosome corresponds to the scheduling order of the tasks; Step 10, initialize the genetic algorithm population, and the priority of tasks on the server is represented as The task numbers higher than the threshold are randomly generated at the front positions of the chromosome; Step 11. Use Fitness(g) = aR(g) - bP(g) to represent the fitness function based on the reward and punishment mechanism, where, represents the score when the task is successfully scheduled, represents the penalty score when the task is not completed, σ is a parameter controlling the penalty intensity, and a, b are weight parameters; Step 12: Calculate the fitness value of each individual according to the fitness function, perform selection, crossover, mutation, and update the population. Execute this step until the number of iterations is reached to obtain the optimal task scheduling order; Step 13: Solve for the optimal policy by the method of Deep Q-Network Enhanced by Genetic Algorithm (GA-DQN) to maximize the completion rate and inference accuracy. Agent i obtains the current state s i (t) from the environment and selects an action a i (t) according to the ε-greedy policy; Step 14. According to the action decisions of all agents The server obtains the task set to be scheduled, executes Step 9, Step 10, Step 11, and Step 12 to obtain the optimal scheduling order of the server, calculates the delay of the task on the server according to the scheduling order, and thereby calculates the task completion situation TC ij ; Step 15: Calculate the reward r i (t), the environment feeds back the reward r i (t) and the next state s i (t + 1) to the agent i. The agent i will store the transition (s i (t), a i (t), r i (t), s i (t + 1)) into the experience pool and sample from the experience pool to update the parameters of the network Q Step 16: The GA algorithm gives the optimal scheduling decision for the server task queue, and the trained Q network outputs the global optimal decisions of partition point selection, whether to exit early, and server selection.

Citation Information

Cited By

  • Quantization-aware multi-model online inference scheduling method for heterogeneous GPU cluster

    CN122547552A

  • Quantization-aware multi-model online inference scheduling method for heterogeneous GPU cluster

    CN122547552B