Heterogeneous dynamic scheduling strategy based on reinforcement learning

Through a heterogeneous dynamic scheduling strategy based on reinforcement learning, and using graph convolutional networks and the A2C algorithm to optimize the scheduling strategy, the inefficiency problem in dynamic and heterogeneous task scheduling environments is solved, and efficient task completion time optimization and resource utilization improvement are achieved.

CN120762830AActive Publication Date: 2025-10-10KUNMING 705 TECH DEV CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510652713.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-10-10
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Existing scheduling algorithms lack broad adaptability when faced with dynamic, changeable and heterogeneous task scheduling environments. They find it difficult to effectively deal with task duration, communication delays and performance differences of heterogeneous platforms, resulting in low scheduling efficiency.

Method used

A heterogeneous dynamic scheduling strategy based on reinforcement learning is adopted. Task information is processed through graph convolutional networks, and the A2C algorithm is combined to train RL agents. Strategy and value network optimization are used to achieve dynamic scheduling and coordinated scheduling of heterogeneous resources.

Benefits of technology

It improves scheduling efficiency, reduces task completion time, adapts to dynamic changes, improves resource utilization, has generalization capabilities, and is suitable for a variety of heterogeneous scheduling environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762830A_ABST
    Figure CN120762830A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous dynamic scheduling strategy based on reinforcement learning, and the strategy comprises the following steps: modeling a scheduling problem as a Markov decision process, defining a state, an action, a transfer function and a return function, extracting task graph structure features through a graph convolution network, and constructing state representation; outputting action probability distribution through a strategy network, performing strategy optimization in combination with a dominant function and entropy regularization, estimating a state value by using a value network, and training by minimizing a Bellman error; in system interaction, the intelligent agent selects task scheduling actions according to states and continuously optimizes strategies, and finally the target of minimizing task completion time is achieved. According to the method, allocation and scheduling decisions can be made according to the system state during operation, and the scheduling efficiency and adaptability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a heterogeneous dynamic scheduling strategy based on reinforcement learning. Background Art

[0002] Task scheduling is a typical combinatorial optimization problem that is NP-hard. Its complexity is particularly exacerbated in dynamic environments. For example, in practical applications, some task graphs (DAGs) may gradually expand with the addition of new resources or computing nodes, making the scheduling process highly uncertain and real-time. This type of problem is widely encountered in complex computing environments such as parallel computing, cloud computing, and edge computing.

[0003] Existing scheduling algorithms primarily include static heuristics (such as HEFT) and some rule-based adaptive methods. However, these algorithms are often limited to specific task graph structures or problem sizes, making them ineffective in addressing uncertainties such as task duration, communication delays, and performance differences across heterogeneous platforms. Furthermore, they lack a universal strategy with broad adaptability.

[0004] As computing systems scale and task complexity continue to increase, traditional scheduling methods face challenges such as weak adaptability and poor scalability. In recent years, reinforcement learning, an intelligent method with self-learning and policy optimization capabilities, has gradually demonstrated its potential for application in dynamic scheduling. Although previous studies have shown that reinforcement learning has advantages in handling small-scale or static task graphs, its policy generalization capabilities and training stability in dynamic, changing, and heterogeneous scheduling environments still need to be further improved. Therefore, there is an urgent need for an intelligent scheduling method that integrates graph structure modeling and reinforcement learning mechanisms, can adapt to dynamic task changes, and improve scheduling efficiency. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an intelligent scheduling method that integrates graph structure modeling and reinforcement learning mechanism, can adapt to dynamic changes in tasks and improve scheduling efficiency.

[0006] The technical solution of the present invention is:

[0007] A heterogeneous dynamic scheduling strategy based on reinforcement learning includes the following steps:

[0008] Step 1. Problem definition and modeling: Formalize the scheduling problem as task scheduling given a directed acyclic graph (DAG) and heterogeneous computing units, with the goal of minimizing task completion time. Model the dynamic scheduling problem as a Markov decision process, defining the state S, action A, transition function P, and reward function R.

[0009] The state S contains information about ready tasks and their descendants within a certain depth, as well as the state of computing resources. The task information is processed using a graph convolutional network (GCN). Each task is represented by original features, and the state is enriched by stacking GCN layers to represent the DAG.

[0010] The action A is defined as: when a computing resource is available, select a task from the ready tasks to run on the resource, or choose to keep it idle;

[0011] The transfer function P is a function in which the computer system transfers from one state to another on its own, and reinforcement learning does not explicitly model this function.

[0012] The reward function R is defined as follows: before the DAG scheduling is completed, the immediate reward of each step is 0; after the scheduling is completed, the reward is the normalized difference between the completion time of the RL algorithm and the completion time of the HEFT algorithm;

[0013] Step 2: Select the A2C algorithm and train the RL agent, which includes the following steps:

[0014] Step 2.1. Network structure design: The A2C algorithm includes two neural networks:

[0015] Policy network: Input state s, output action probability distribution π(a|s;θ), which represents the probability of choosing action a in state s, where θ is a parameter of the policy network. By optimizing θ, the policy network can adjust the action probability distribution to maximize the cumulative reward;

[0016] Value Network: Input state s, output state value V(s;φ), which represents the expected cumulative return of the state, where φ is a parameter of the value network. By optimizing φ, the value network can more accurately estimate the value of the state;

[0017] Step 2.2: Neural network optimization, including:

[0018] Policy network optimization: The objective function combines policy gradient, advantage function and entropy regularization, and is expressed as:

[0019] J(θ)=E[A(s,a)·logπ(a∣s;θ)+βH(π(·∣s;θ))],

[0020] A(s,a)=Q(s,a)-V(s;φ),

[0021] Q(s,a)=r t +γV(s t+1 ; φ),

[0022] H(π(·|s;θ))=-Σ a π(a∣s; θ)logπ(a∣s; θ),

[0023] Where J(θ) is the objective function of the policy network, which represents the expected return of policy π under parameter θ. It is used to optimize the policy network. The goal is to maximize J(θ) through gradient ascent to improve the policy. E[·] is the expectation operator, which represents the expectation of the distribution of state s and action a. A(s,a) is the advantage function, which measures the pros and cons of choosing action a in state s compared to the average action. Q(s,a) is the action-state value, which represents the immediate reward r obtained after taking action a in state s. t Add the discounted future value; V(s;φ) is the state value, which represents the expected cumulative return of state s and is output by the value network; r t is the immediate reward obtained at time step t, γ is the discount factor used to balance the immediate and future rewards; V(s t+1 ;φ) is the next state s t+1 The state value of , s is the current state, a is the current action, H(π(·|s;θ)) is the policy entropy, which measures the uncertainty of the policy under state s. The larger the entropy, the more random the policy; the smaller the entropy, the more certain the policy; π(a|s;θ) is the action probability distribution, β is the entropy regularization hyperparameter, which controls the effect of the entropy term on the optimization and prevents the policy from converging to the local optimum too early.

[0024] Value network optimization: The value network is optimized by minimizing the Bellman error, which is expressed as: L(φ)=E[(r t +γV(s t+1 ;φ)-V(s t ;φ)) 2 ],

[0025] Where L(φ) is the loss function of the value network, which is used to optimize the value network parameter φ. The goal is to make the value network estimate the state value more accurately by minimizing L(φ); E[·] is the expectation operator, r t is the immediate reward obtained at time step t, γ is the discount factor;

[0026] Step 2.3: Training process, including:

[0027] The agent interacts with the environment and generates a series of state-action-reward sequences (s t ,a t ,r t ,s t+1 );

[0028] After collecting multiple steps of data, update θ and φ using stochastic gradient descent;

[0029] Repeat the interaction and optimization until the strategy converges;

[0030] Step 3: Build the RL agent architecture, which includes the following steps:

[0031] Step 3.1, data input: input the current DAG task graph and resource state into the agent;

[0032] Step 3.2, DAG information processing: use stacked graph convolutional layers to extract DAG features, the expression is:

[0033] h v (l+1) = ReLU(W (l) · Aggregate({h u (l) | u e N(v) U {v}})),

[0034] where h v (l+1) is the feature vector of node v in the l+1 layer, which represents the updated feature of node v after l+1 layer GCN processing, which is used to capture the information of node and its neighbors; W (l) is the weight matrix of the l layer, which is used to linearly transform the aggregated features; h u (l) is the feature vector of node u in the l layer, and N(v) is the neighbor node set of node v;

[0035] Each layer of GCN aggregates the information of nodes and their neighbors to generate more rich feature representation;

[0036] The last layer of GCN outputs the global representation of DAG h G ;

[0037] Step 3.3, resource state embedding, including:

[0038] Embedding the resource state into a vector h R ;

[0039] Combine the global representation of DAG h G and the vector h R to get the complete state representation h = [h G , h R ];

[0040] Step 3.4, calculate action probability: input h into the fully connected layer of the policy network, the expression is:

[0041] π(a | s; θ) = Softmax(W[h G , h R ] + b),

[0042] where W is the weight of the fully connected layer, and b is the bias;

[0043] Step 3.5, action sampling and execution: Sample an action from π(a|s;θ), select a ready task and assign it to the current computing resource for execution; if the sampling result is 0 action, no task scheduling is performed;

[0044] Step 3.6, status update and loop: Update task status and resource utilization, determine whether there are ready tasks, and repeat scheduling until all tasks are completed.

[0045] Furthermore, in step 1, the graph convolutional network (GCN) is used to process task information. Each task is represented by original features, and the state is enriched by stacking GCN layers to express the DAG:

[0046]

[0047] Among them, H (l+1) is the node feature matrix of the l+1th layer, H (l) is the node feature matrix of the lth layer, is the adjacency matrix with self-loops added, A is the original adjacency matrix, I is the identity matrix, for The corresponding degree matrix, W (l) is the learnable weight matrix of the lth layer, which is used to map the input features to the feature space of the next layer.

[0048] Furthermore, the expression of the reward function R in step 1 is:

[0049]

[0050] Among them, T RL is the completion time of the RL algorithm, T HEFT is the completion time of HEFT algorithm, when T RL <T HEFT When , R>0, it means that the RL algorithm is better than the HEFT algorithm.

[0051] Beneficial effects of the present invention:

[0052] 1. Improve scheduling efficiency and optimize task completion time: By introducing the A2C algorithm from reinforcement learning, the intelligent agent can adaptively select the optimal scheduling action under different states, effectively reducing the overall completion time of task scheduling. It outperforms traditional heuristic algorithms such as HEFT in most task graphs.

[0053] 2. Adaptability to dynamically changing scheduling environments: By modeling the scheduling problem as a Markov decision process and combining dynamic graph structure input with resource status, the present invention can dynamically respond to the gradual expansion of the task graph structure and the real-time changes in resource status, and has strong environmental adaptability.

[0054] 3. Support for heterogeneous resource scheduling and multi-task parallelism: By combining resource state embedding with a policy network, the present invention enables intelligent agents to reasonably distinguish the characteristics of different computing resources (such as CPUs and GPUs), achieve collaborative scheduling of heterogeneous resources, and improve overall system resource utilization;

[0055] 4. The reinforcement learning strategy has self-learning and generalization capabilities: The RL agent proposed in this paper gradually optimizes the scheduling strategy during multiple rounds of interactive training. It does not rely on fixed rules and can show strong generalization capabilities when facing different task sizes, different DAG structures, and unknown task graphs.

[0056] 5. Introducing Graph Convolutional Networks to Enhance Task Graph Structure Understanding: This paper uses GCN to perform multi-layer embedding learning on the DAG task graph, capturing the topological dependencies between nodes. This makes state representation richer, policy decisions more reasonable, and significantly improves the accuracy of scheduling decisions.

[0057] 6. Collaborative optimization of strategy and value networks to improve stability and convergence speed: This paper adopts the dual-network structure of the A2C algorithm, introduces a value benchmark in strategy optimization, and improves the stability of gradient estimation. It also uses entropy regularization to prevent the strategy from falling into local optimality too early, thereby improving overall training efficiency and the diversity of scheduling strategies.

[0058] 7. Strong portability and applicability to various scheduling scenarios: The method of the present invention has no task graph-specific dependencies and strategy learning capabilities, and is applicable to various heterogeneous scheduling environments such as cloud computing, edge computing, and high-performance parallel systems, with good engineering practicality and deployability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a basic flow chart of a heterogeneous dynamic scheduling strategy based on reinforcement learning of the present invention.

[0060] Figure 2 It is a schematic diagram of the principle of a heterogeneous dynamic scheduling strategy based on reinforcement learning of the present invention. DETAILED DESCRIPTION

[0061] like Figure 1-2 As shown in Figure 1, a heterogeneous dynamic scheduling strategy based on reinforcement learning includes the following steps:

[0062] Step 1. Problem definition and modeling: Formalize the scheduling problem as task scheduling given a directed acyclic graph (DAG) and heterogeneous computing units, with the goal of minimizing task completion time. Model the dynamic scheduling problem as a Markov decision process, defining the state S, action A, transition function P, and reward function R.

[0063] The state S contains information about ready tasks and their descendants within a certain depth, as well as the state of computing resources. The task information is processed using a graph convolutional network (GCN). Each task is represented by original features, and the state is enriched by stacking GCN layers to represent the DAG.

[0064] The action A is defined as: when a computing resource is available, select a task from the ready tasks to run on the resource, or choose to keep it idle;

[0065] The transfer function P is a function in which the computer system transfers from one state to another on its own, and reinforcement learning does not explicitly model this function.

[0066] The reward function R is defined as follows: before the DAG scheduling is completed, the immediate reward of each step is 0; after the scheduling is completed, the reward is the normalized difference between the completion time of the RL algorithm and the completion time of the HEFT algorithm;

[0067] Step 2: Select the A2C algorithm and train the RL agent, which includes the following steps:

[0068] Step 2.1. Network structure design: The A2C algorithm includes two neural networks:

[0069] Policy network: Input state s, output action probability distribution π(a|s;θ), which represents the probability of choosing action a in state s, where θ is a parameter of the policy network. By optimizing θ, the policy network can adjust the action probability distribution to maximize the cumulative reward;

[0070] Value Network: Input state s, output state value V(s;φ), which represents the expected cumulative return of the state, where φ is a parameter of the value network. By optimizing φ, the value network can more accurately estimate the value of the state;

[0071] Step 2.2: Neural network optimization, including:

[0072] Policy network optimization: The objective function combines policy gradient, advantage function and entropy regularization, and is expressed as:

[0073] J(θ)=E[A(s,a)·logπ(a∣s;θ)+βH(π(·∣s;θ))],

[0074] A(s,a)=Q(s,a)-V(s;φ),

[0075] Q(s,a)=r t +γV(s t+1 ; φ),

[0076] H(π(·|s;θ))=-Σ a π(a∣s; θ)logπ(a∣s; θ),

[0077] where J(θ) is the objective function of the policy network, representing the expected return of the policy π at parameters θ, used to optimize the policy network, the goal is to maximize J(θ) through gradient ascent, thus improving the policy; E[·] is the expectation operator, representing the expectation over the distribution of states s and actions a; A(s, a) is the advantage function, measuring the superiority of choosing action a in state s compared to the average action; Q(s, a) is the action-state value, representing the immediate reward r t obtained after taking action a in state s, discounted by the future value; V(s; φ) is the state value, representing the expected cumulative return of state s, output by the value network; r t is the immediate reward obtained at time step t, γ is the discount factor, used to balance immediate and future rewards; V(s t+1 ; φ) is the state value of the next state s t+1 , s is the current state, a is the current action, H(π(·|s; θ)) is the policy entropy, measuring the uncertainty of the policy in state s, the greater the entropy, the more random the policy; the smaller the entropy, the more deterministic the policy; π(a|s; θ) is the action probability distribution, β is the entropy regularization hyperparameter, controlling the influence of the entropy term on optimization, preventing the policy from converging to a local optimum too early;

[0078] Value network optimization: optimize the value network by minimizing the Bellman error, the expression is: L(φ) = E[(r t + γV(s t+1 ; φ) - V(s t ; φ)) 2 ],

[0079] where L(φ) is the loss function of the value network, used to optimize the value network parameters φ, the goal is to minimize L(φ) to make the value network more accurately estimate the state value; E[·] is the expectation operator, r t is the immediate reward obtained at time step t, γ is the discount factor;

[0080] Step 2.3 training process, including:

[0081] The agent interacts with the environment, generating a series of state-action-reward sequences (s t ,a t ,r t ,s t+1 );

[0082] After collecting multi-step data, update θ and φ using stochastic gradient descent;

[0083] Repeat the interaction and optimization until the policy converges;

[0084] Step 3, build the RL agent architecture, including the following steps:

[0085] Step 3.1, Data input: Send the current DAG task graph and resource status to the agent;

[0086] Step 3.2, DAG information processing: Use stacked graph convolution layers to extract DAG features, the expression is:

[0087] h v (l+1) =ReLU(W (l) ·Aggregate({h u (l) |u∈N(v)∪{v}})),

[0088] Among them, h v (l+1) is the feature vector of node v in the l+1 layer, which represents the updated features of node v after being processed by the l+1 layer GCN, and is used to capture the information of the node and its neighbors; W (l) is the weight matrix of the lth layer, which is used to perform linear transformation on the aggregated features; h u (l) is the feature vector of node u in layer l, N(v) is the set of neighbor nodes of node v;

[0089] Each layer of GCN aggregates node and neighbor information to generate richer feature representations;

[0090] The last layer of GCN outputs the global representation h of DAG G ;

[0091] Step 3.3: Resource status embedding, including:

[0092] Embed the computing resource state into vector h R ;

[0093] Combined with the global representation h of DAG G and vector h R , and the complete state representation h=[h G ,h R ];

[0094] Step 3.4, calculate the action probability: input h into the fully connected layer of the policy network, the expression is:

[0095] π(a|s;θ)=Softmax(W[h G ,h R ]+b),

[0096] Among them, W is the weight of the fully connected layer, and b is the bias;

[0097] Step 3.5, action sampling and execution: Sample an action from π(a|s;θ), select a ready task and assign it to the current computing resource for execution; if the sampling result is 0 action, no task scheduling is performed;

[0098] Step 3.6, status update and loop: Update task status and resource utilization, determine whether there are ready tasks, and repeat scheduling until all tasks are completed.

[0099] Furthermore, in step 1, the graph convolutional network (GCN) is used to process task information. Each task is represented by original features, and the state is enriched by stacking GCN layers to express the DAG:

[0100]

[0101] Among them, H (l+1) is the node feature matrix of the l+1th layer, H (l) is the node feature matrix of the lth layer, is the adjacency matrix with self-loops added, A is the original adjacency matrix, I is the identity matrix, for The corresponding degree matrix, W (l) is the learnable weight matrix of the lth layer, which is used to map the input features to the feature space of the next layer.

[0102] Furthermore, the expression of the reward function R in step 1 is:

[0103]

[0104] Among them, T RL is the completion time of the RL algorithm, T HEFT is the completion time of HEFT algorithm, when T RL <T HEFT When , R>0, it means that the RL algorithm is better than the HEFT algorithm.

[0105] The basic principles, main features and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A heterogeneous dynamic scheduling strategy based on reinforcement learning, characterized in that: The following steps are involved: Step 1. Problem definition and modeling: Formalize the scheduling problem as task scheduling given a directed acyclic graph (DAG) and heterogeneous computing units, with the goal of minimizing task completion time. Model the dynamic scheduling problem as a Markov decision process, defining the state S, action A, transition function P, and reward function R. The state S contains information about ready tasks and their descendants within a certain depth, as well as the state of computing resources. The task information is processed using a graph convolutional network (GCN). Each task is represented by original features, and the state is enriched by stacking GCN layers to represent the DAG. The action A is defined as: when a computing resource is available, select a task from the ready tasks to run on the resource, or choose to keep it idle; The transfer function P is transferred from one state to another by the computer system itself, and reinforcement learning does not explicitly model it; The reward function R is defined as follows: before the DAG scheduling is completed, the immediate reward of each step is 0; after the scheduling is completed, the reward is the normalized difference between the completion time of the RL algorithm and the completion time of the HEFT algorithm; Step 2: Select the A2C algorithm and train the RL agent, which includes the following steps: Step 2.

1. Network structure design: The A2C algorithm includes two neural networks: Policy network: Input state s, output action probability distribution π(a|s; θ), which represents the probability of selecting action a in state s, where θ is the parameter of the policy network. By optimizing θ, the policy network can adjust the action probability distribution to maximize the cumulative return; Value network: Input state s, output state value V(s; φ), which represents the expected cumulative return of the state, where φ is the parameter of the value network. By optimizing φ, the value network can more accurately estimate the state value; Step 2.2: Neural network optimization, including: Policy network optimization: The objective function combines policy gradient, advantage function and entropy regularization, and is expressed as: J(θ)=E[A(s,a)·logπ(a∣s;θ)+βH(π(·∣s;θ))], A(s,a)=Q(s,a)-V(s;φ), Q(s,a)=r t +γV(s t+1 ;φ), H(π(·∣s;θ))=-∑ a π(a∣s;θ)logπ(a∣s;θ), Where J(θ) is the objective function of the policy network, which represents the expected return of policy π under parameter θ. It is used to optimize the policy network. The goal is to maximize J(θ) through gradient ascent to improve the policy. E[·] is the expectation operator, which represents the expectation of the distribution of state s and action a. A(s,a) is the advantage function, which measures the pros and cons of choosing action a in state s compared to the average action. Q(s,a) is the action-state value, which represents the immediate reward r obtained after taking action a in state s. t Add the discounted future value; V(s;φ) is the state value, which represents the expected cumulative return of state s and is output by the value network; r t is the immediate reward obtained at time step t, γ is the discount factor used to balance the immediate and future rewards; V(s t+1 ;φ) is the next state s t+1 The state value of , s is the current state, a is the current action, H(π(·|s;θ)) is the policy entropy, which measures the uncertainty of the policy under state s. The larger the entropy, the more random the policy; the smaller the entropy, the more certain the policy; π(a|s;θ) is the action probability distribution, β is the entropy regularization hyperparameter, which controls the impact of the entropy term on optimization and prevents the policy from converging to the local optimum too early; Value network optimization: The value network is optimized by minimizing the Bellman error, and the expression is: L(φ)=E[(r t +γV(s t+1 ;φ)-V(s t ;φ)) 2 ], Where L(φ) is the loss function of the value network, which is used to optimize the value network parameter φ. The goal is to make the value network estimate the state value more accurately by minimizing L(φ); E[·] is the expectation operator, r t is the immediate reward obtained at time step t, γ is the discount factor; Step 2.3: Training process, including: The agent interacts with the environment and generates a series of state-action-reward sequences (s t ,a t ,r t ,s t+1 ); After collecting multiple steps of data, update θ and φ using stochastic gradient descent; Repeat the interaction and optimization until the strategy converges; Step 3: Build the RL agent architecture, which includes the following steps: Step 3.1, Data Input: Send the current DAG task graph and resource status to the agent; Step 3.2, DAG information processing: Use stacked graph convolution layers to extract DAG features, the expression is: h v (l+1) =ReLU(W (l) ·Aggregate({h u (l) ∣u∈N(v)∪{v}}))), Among them, h v (l+1) is the feature vector of node v in the l+1 layer, which represents the updated features of node v after being processed by the l+1 layer GCN, and is used to capture the information of the node and its neighbors; W (l) is the weight matrix of the lth layer, which is used to perform linear transformation on the aggregated features; h u (l) is the feature vector of node u in layer l, N(v) is the set of neighbor nodes of node v; Each layer of GCN aggregates node and neighbor information to generate richer feature representations; The last layer of GCN outputs the global representation h of DAG G ; Step 3.3: Resource status embedding, including: Embed the computing resource state into vector h R ; Combined with the global representation h of DAG G and vector h R , and the complete state representation h=[h G ,h R ]; Step 3.4, calculate the action probability: input h into the fully connected layer of the policy network, the expression is: π(a∣s;θ)=Softmax(W[h G ,h R ]+b), Where W is the weight of the fully connected layer and b is the bias; Step 3.5, action sampling and execution: Sample an action from π(a|s;θ), select a ready task and assign it to the current computing resource for execution; if the sampling result is 0 action, no task scheduling is performed; Step 3.6, status update and loop: Update task status and resource utilization, determine whether there are ready tasks, and repeat scheduling until all tasks are completed.

2. A heterogeneous dynamic scheduling strategy based on reinforcement learning according to claim 1, characterized in that: In step 1, the graph convolutional network (GCN) is used to process task information. Each task is represented by original features, and the state is enriched by stacking GCN layers to express the DAG: Among them, H (l+1) is the node feature matrix of the l+1th layer, H (l) is the node feature matrix of the lth layer, is the adjacency matrix with self-loops added, A is the original adjacency matrix, I is the identity matrix, for The corresponding degree matrix, W (l) is the learnable weight matrix of the lth layer, which is used to map the input features to the feature space of the next layer.

3. The heterogeneous dynamic scheduling strategy based on reinforcement learning according to claim 1, characterized in that: The expression of the reward function R in step 1 is: Among them, T RL is the completion time of the RL algorithm, T HEFT is the completion time of HEFT algorithm, when T RL <T HEFT When , R>0, it means that the RL algorithm is better than the HEFT algorithm.

Citation Information

Patent Citations

  • Multi-agent cooperative computing resource scheduling method, device and system

    CN116909742A

  • Virtualized digital twin traffic resource scheduling model based on cloud edge collaboration

    CN119417022A

  • Intelligent construction scheduling method based on Internet of Things

    CN119476881A

  • Early research and judgment method and device for online social network unreal public opinion information

    CN119646597A

  • Flexible job shop scheduling method based on act-critic multi-agent deep reinforcement learning

    CN119740803A