Rescheduling method of edge service task based on HG-Trans
By modeling the task scheduling problem as a Markov decision-making process and using the HG-Trans network combined with reinforcement learning strategies, the problem that existing scheduling methods are difficult to quickly adjust the task execution path after a failure occurs, and the task scheduling performance improvement and resource optimization allocation in the fault scenario are achieved.
Patent Information
- Application Number
- CN202510150247.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing scheduling methods are difficult to quickly adjust the task execution path after a failure occurs, and lack an efficient failure recovery mechanism, resulting in task loss, data inconsistency and service quality decline.
The dynamic task scheduling problem is modeled as a Markov decision-making process, and the HG-Trans network is combined with reinforcement learning strategies to dynamically adapt to node failures in edge computing scenarios, adjust the scheduling strategy in real time, and automatically identify the faulty nodes through the status representation of the heterogeneous graph structure.
It realizes the improvement of task scheduling performance in failure scenarios, can quickly adjust the task execution path, reduce scheduling delay, and achieve global optimal allocation of resources, which improves the execution success rate of task rescheduling.
Smart Images

Figure CN119987972A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of task management, and in particular relates to a method for rescheduling edge service tasks based on HG-Trans. Background Art
[0002] With the rapid development of Internet of Things (IoT) and Mobile Edge Computing (MEC) technologies, more and more computing tasks are transferred from centralized cloud processing to edge devices close to the data source to shorten response time, reduce data transmission delay and improve computing efficiency. However, the edge computing environment faces many challenges due to its inherent distributed and heterogeneous characteristics, especially in terms of device failure and uncertainty. Compared with the traditional centralized computing architecture, edge nodes are often resource-constrained devices, such as low-power servers, IoT sensors or smart terminals, which are susceptible to hardware damage, power consumption limitations, unstable network connections and other issues. Whether the edge device suddenly drops offline or the communication delay between nodes caused by network congestion, it will have an impact on the tasks being executed. If the fault is not handled in time, it may lead to task loss, data inconsistency, service quality degradation, and even affect the stability of the entire system. Therefore, in the fault scenario, how to effectively recover the interrupted task and minimize the recovery cost becomes a key issue to ensure the reliability and performance of the edge computing system.
[0003] Directed acyclic graph (DAG) is a common task representation, which is widely used in fields such as workflow management, data processing and computing task allocation. DAG task scheduling aims to assign tasks to appropriate computing nodes according to task dependencies and resource constraints. At present, the DAG task scheduling problem in distributed heterogeneous computing environments has been widely studied. H. Topcuoglu proposed the HEFT algorithm to assign each subtask to the processor with the least execution time. Q. Shen et al. proposed DTOSC to perform DAG task offloading and service caching in vehicle edge computing. Y. Liu developed MAMTS to assign priorities to different task computing topologies. The above methods rely on static scheduling, which predetermines the scheduling scheme before the task starts to execute, without considering the changes at runtime, and thus it is difficult to deal with dynamic failures. The dynamic scheduling method adjusts according to the actual situation at runtime and has better adaptability. J. Yan proposed an actor-critic DRL to learn the optimal DAG subtask allocation of access points. J. Wang proposed a DAG task offloading method based on meta-reinforcement learning. X. Wei developed a joint optimization drone trajectory planning and DAG task scheduling algorithm based on DRL. In order to effectively extract the dependencies between subtasks, H.Lee used graph convolutional networks (GCN) for DAG task scheduling. J.Chen proposed a DAG task offloading algorithm called ACE, which uses GCN to capture the topological information of DAG subtasks. None of the above studies considered the impact of fault events on task interruption in the dynamic scheduling process. In response to the fault problem, Lee and Gil proposed a checkpoint and replication based on clustering heuristic (CRCH) to schedule jobs and tolerate faults in cloud frameworks. Wu proposed an improved workflow scheduling and fault-tolerant process integrated scheduling mechanism. Malik completed the fault scheduling and detection mechanism through the dynamic standby replication (LSR) strategy. The above studies provide many valuable methods for task scheduling and fault recovery problems, but do not consider how to extract information state features in the scheduling process under fault conditions, effectively integrate task information and scheduling information, and quickly handle interrupted tasks. In this context, it is necessary to develop a rescheduling method that can perceive task dependencies and environmental states when a fault occurs and migrate failed tasks to other edge devices for execution at the lowest cost. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the HG-Trans-based edge service task rescheduling method provided by the present invention solves the problem that the existing scheduling method relies on local information or simple rules, lacks an efficient fault recovery mechanism, and cannot quickly adjust the task execution path after a fault occurs.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a rescheduling method of edge service tasks based on HG-Trans, comprising the following sub-steps:
[0006] S1. Model the dynamic task scheduling problem as a Markov decision process and convert the scheduling state into a heterogeneous graph structure;
[0007] S2, input the heterogeneous graph structure into the HG-Trans network to obtain the estimated state value and the probability distribution of all actions;
[0008] S3. Input the estimated state value and the probability distribution of all actions into the decision network and output the task scheduling strategy.
[0009] Further: In S1, the method of modeling the dynamic task scheduling problem as a Markov decision process is specifically as follows:
[0010] (1) Setting the state: The comprehensive reflection of all task arrangements and computing resource states at any time constitutes the state at the corresponding time;
[0011] (2) Set actions: define the actions at any time as TS pairs, where T is the task node and S is the server node;
[0012] (3) Set the transition function: The transition function represents the probability of action a transitioning from the current state s to the next state s′. The transition function P a The specific expression of (s,s′) is:
[0013] P a (s,s′)=P(s t+1 =s′|s t =s,a t =a)
[0014] In the formula, s t is the state at time t, s t+1 is the state at time t+1, a t is the action at time t;
[0015] (4) Set rewards: For any time, the reward function R (Makespan) is expressed as:
[0016]
[0017] Where R(·) is the reward function, Makespan is the maximum time span for task completion, Makespan(HEFT) is the maximum time for HEFT to complete, and HEFT is the heterogeneous earliest completion time algorithm, which is the most commonly used heuristic algorithm in task scheduling.
[0018] Further: In said S1, the heterogeneous graph structure H t =(T,S,ε t ), where ε t is the arc set of TS pairs, and each task node corresponds to a server node.
[0019] Further: S2 comprises the following sub-steps:
[0020] S21, input the heterogeneous graph structure into the multi-layer TransformerConv, update the node features, and generate node embedding;
[0021] S22. Embed the nodes into a multi-layer perceptron to generate estimated state values and probability distributions of actions.
[0022] Further: In S21, the method of updating the features of each node by a single-layer TransformerConv comprises the following steps:
[0023] S211, input the heterogeneous graph structure into the multi-layer TransformerConv, perform weighted aggregation on the nodes and their neighboring nodes based on the Transformer self-attention mechanism, and calculate the relationship between node i and its first-order domain The attention coefficient A between nodes j in ij ;
[0024]
[0025] In the formula, Q i is the query vector, K j is the key vector, d k is the dimension of the key vector, T is the matrix transpose,
[0026]
[0027] S212, use the softmax function to normalize the attention coefficient in the neighborhood, and calculate the relationship between node i and its first-order domain The attention weight α between nodes j in ij ;
[0028]
[0029] S213, weighted aggregation of neighbor node information is performed through attention weights to update node features. The updated node features h i The specific expression is:
[0030]
[0031] In the formula, σ is a nonlinear activation function, V jis the feature representation of node j.
[0032] Further: In S22, the expression for generating the estimated state value V is specifically:
[0033]
[0034] Where N is the total number of nodes, W v is the projection matrix, W v ∈R d×1 , R is a matrix.
[0035] Further: In S22, the method for generating the probability distribution of the action includes the following steps:
[0036] S221, embed the nodes of the available tasks and perform aggregation operations to generate a batch matrix H T ;
[0037] H T =[h1,h2,...,h M ] T
[0038] In the formula, h M Node embedding for the Mth task;
[0039] S222, map the batch matrix into a one-dimensional vector space, generate a score for each task, and normalize the score through the Softmax function of the fully connected layer to obtain the probability distribution of the action, where the probability π of the lth action l The specific expression is:
[0040]
[0041] In the formula, o l It is the lth component in the Softmax function output vector o.
[0042] Further: S3 is specifically:
[0043] The policy gradient theorem optimizes the policy network. The objective function of the policy network is established based on the estimated state value and the probability distribution of all actions. The objective function is minimized to search for the optimal scheduling action.
[0044] Among them, the expression of the objective function is specifically:
[0045]
[0046] In the formula, is the entropy function of the probability distribution, β is the hyperparameter that controls the effect of entropy regularization, and A(s,a) is the advantage function that measures whether action a is better than the average value in state s.t ,a t ,r t ,s t+1 ) and the current estimated value of the estimated state value V, π(a|s) is the policy function, is the gradient of the policy network parameter θ.
[0047] The beneficial effects of the present invention are as follows: the present invention provides a rescheduling method for edge service tasks based on HG-Trans, and utilizes an edge fault scenario task rescheduling strategy based on a heterogeneous graph neural network to comprehensively improve the task scheduling performance of the edge computing system in a fault scenario. This method solves the problem that the existing scheduling methods rely on local information or simple rules, lack an efficient fault recovery mechanism, and cannot quickly adjust the task execution path after a fault occurs, thereby reducing scheduling delays while achieving global optimal allocation of resources. Compared with the prior art, this method has the following effects:
[0048] (1) The task rescheduling problem is modeled as a Markov decision process and combined with a reinforcement learning strategy, which can dynamically adapt to node failures in edge computing scenarios and adjust the scheduling strategy in real time.
[0049] (2) A fault-tolerant mechanism is introduced into state modeling. The faulty nodes are automatically identified through the state representation of the heterogeneous graph structure, and the interrupted tasks are adaptively adjusted, thereby improving the execution success rate of task rescheduling.
[0050] (3) The Transformer-based heterogeneous graph embedding method can efficiently capture the complex relationship between tasks and computing resources, optimize scheduling strategies, reduce scheduling delays, and achieve global optimal allocation of resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flow chart of the HG-Trans-based edge service task rescheduling method of the present invention.
[0052] Figure 2 A framework for task rescheduling strategies for edge failure scenarios based on heterogeneous graph neural networks.
[0053] Figure 3 This is the HG-Trans network structure diagram. DETAILED DESCRIPTION
[0054] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0055] like Figure 1 As shown, in one embodiment of the present invention, a method for rescheduling edge service tasks based on HG-Trans includes the following steps:
[0056] S1. Model the dynamic task scheduling problem as a Markov decision process and convert the scheduling state into a heterogeneous graph structure;
[0057] S2, input the heterogeneous graph structure into the HG-Trans network to obtain the estimated state value and the probability distribution of all actions;
[0058] S3. Input the estimated state value and the probability distribution of all actions into the decision network and output the task scheduling strategy.
[0059] In this embodiment, the method of the present invention uses a task rescheduling strategy for edge failure scenarios based on heterogeneous graph neural networks, and its framework is as follows: Figure 2 As shown, the present invention builds an edge simulation environment based on the edge network model and the resource fault model, and establishes a task computing model according to the edge network architecture. Task rescheduling is a continuous decision-making process. By iteratively taking scheduling actions, tasks are assigned to executable computing resources in each state until all tasks are scheduled. In each iteration, the scheduling state is first converted into a heterogeneous graph structure, and then the heterogeneous graph structure with a two-stage embedding process is input into the HG-Trans network, which is applied to the processing of heterogeneous graph convolution to extract feature embeddings of operations and nodes. The decision network uses these feature embeddings to generate action probability distributions and samples scheduling actions from them.
[0060] In S1, the method of modeling the dynamic task scheduling problem as a Markov decision process is specifically as follows:
[0061] (1) Setting the state: The comprehensive reflection of all task arrangements and computing resource states at any time constitutes the state at the corresponding time;
[0062] The comprehensive reflection of all task arrangements and computing resource states at time t constitutes state s t , will s t The information represented in is approximately information about running tasks, ready tasks, and their descendants. The state of server resources is represented by a vector that contains the type of each computing resource and the estimated time it will be available.
[0063] (2) Set actions: define the actions at any time as TS pairs, where T is the task node and S is the server node;
[0064] In the action definition process, task selection and resource allocation are considered as a composite decision. t ∈At Defined as a TS pair. When the server resource is available, the action is to select the task to run on this computing resource. When the server resource is unavailable due to a failure, the RL agent quickly responds to the task being executed on the failed server and puts it back into the pending queue, and clears all computing resource allocation information related to the task on the failed resource.
[0065] (3) Set the transition function: The transition function represents the probability of action a transitioning from the current state s to the next state s′. The transition function P a The specific expression of (s,s′) is:
[0066] P a (s,s′)=P(s t+1 =s′|s t =s,a t =a)
[0067] In the formula, s t is the state at time t, s t+1 is the state at time t+1, a t For the action at time t, the present invention distinguishes two different states through the topological structure and characteristics of the heterogeneous graph.
[0068] (4) Set rewards: For any time, the reward function r (Makespan) is expressed as:
[0069]
[0070] Where R(·) is the reward function, Makespan is the maximum time span for task completion, Makespan(HEFT) is the maximum time for HEFT to complete, and HEFT is the heterogeneous earliest completion time algorithm, which is the most commonly used heuristic algorithm in task scheduling.
[0071] In this embodiment, the reward provides feedback information to the RL agent on how well it performs in optimizing the objective. Due to the stability and applicability of the HEFT algorithm in scheduling problems, the reward function is defined by normalizing the current scheduling Makespan with the Makespan of the HEFT baseline algorithm. The RL agent performs operations based on the visited state and the current strategy, interacts with the problem to be solved, and gradually adjusts the strategy to optimize the objective function, that is, minimize the Makespan.
[0072] The present invention models the dynamic task scheduling problem as a Markov decision process (MDP). At time t, the agent observes the current system state s t and make a decision tThat is, at the current time T(t), the unexecuted tasks are assigned to the available computing resources. Then, the environment is transferred to t+1, and the above process is iterated continuously until all tasks are scheduled.
[0073] In S1, the heterogeneous graph structure H t =(T,S,ε t ), where ε t is the arc set of TS pairs, and each task node corresponds to a server node.
[0074] The S2 comprises the following sub-steps:
[0075] S21, input the heterogeneous graph structure into the multi-layer TransformerConv, update the node features, and generate node embedding;
[0076] S22. Embed the nodes into a multi-layer perceptron to generate estimated state values and probability distributions of actions.
[0077] In this embodiment, the HG-Trans network is as follows Figure 3 As shown, the network of HG-Trans is used to process graph topology and node features.
[0078] In S21, the method of updating the features of each node by a single-layer TransformerConv includes the following steps:
[0079] S211, input the heterogeneous graph structure into the multi-layer TransformerConv, perform weighted aggregation on the nodes and their neighboring nodes based on the Transformer self-attention mechanism, and calculate the relationship between node i and its first-order domain The attention coefficient A between nodes j in ij ;
[0080]
[0081] In the formula, Q i is the query vector, K j is the key vector, d k is the dimension of the key vector, T is the matrix transpose,
[0082]
[0083] S212, use the softmax function to normalize the attention coefficient in the neighborhood, and calculate the relationship between node i and its first-order domain The attention weight α between nodes j in ij ;
[0084]
[0085] S213, weighted aggregation of neighbor node information is performed through attention weights to update node features. The updated node features h i The specific expression is:
[0086]
[0087] In the formula, σ is a nonlinear activation function, V j is the feature representation of node j.
[0088] In S22, the expression for generating the estimated state value V is specifically:
[0089]
[0090] Where N is the total number of nodes, W v is the projection matrix, W v ∈R d×1 , R is a matrix.
[0091] In this embodiment, mean pooling is used to aggregate all node embeddings, and the state value V is estimated through one-dimensional projection. is the global feature vector after mean pooling. The present invention improves the neural network structure of the "actor-critic" algorithm (A2C) to learn the RL agent, and improves the value function V by minimizing the mean square error of the Bellman function.
[0092] Among them, the "actor-critic" algorithm (A2C) is a reinforcement learning algorithm based on policy function and value function. The algorithm contains two networks. One is the policy network, which outputs an action a and takes an action based on the current state s. The second is the value network, which is responsible for evaluating the quality of the current action a. When optimizing decisions, the two neural networks in A2C are improved. One is to improve the estimation of the value network by minimizing the mean square error of the Bellman function. The other is to transform the problem of maximizing the cumulative reward into a combined optimization of the policy network gradient and the advantage function through the policy gradient theorem, and add the entropy of the policy to the objective function minimized by the policy network.
[0093] In S22, the method for generating the probability distribution of the action includes the following steps:
[0094] S221, embed the nodes of the available tasks and perform aggregation operations to generate a batch matrix H T ;
[0095] H T =[h1,h2,...,h M ] T
[0096] In the formula, hM Node embedding for the Mth task;
[0097] S222, map the batch matrix into a one-dimensional vector space, generate a score for each task, and normalize the score through the Softmax function of the fully connected layer to obtain the probability distribution of the action, where the probability π of the lth action l The specific expression is:
[0098]
[0099] In the formula, o l It is the lth component in the Softmax function output vector o.
[0100] The S3 is specifically:
[0101] The policy gradient theorem optimizes the policy network, establishes the objective function of the policy network based on the estimated state value and the probability distribution of all actions, and minimizes the objective function to search for the optimal scheduling action; the policy network is based on the policy function π θ (a t |s t ), θ is the weight parameter of the policy network. Given the state s at time t, it outputs an action probability distribution, which indicates the probability of the agent taking a certain action in each state.
[0102] Among them, the expression of the objective function is specifically:
[0103]
[0104] In the formula, is the entropy function of the probability distribution, β is the hyperparameter that controls the effect of entropy regularization, and A(s,a) is the advantage function that measures whether action a is better than the average value in state s. t ,a t ,r t ,s t+1 ) and the current estimated value of the estimated state value V, π(a|s) is the policy function, is the gradient of the policy network parameter θ.
[0105] In this embodiment, the present invention optimizes the decision-making of task scheduling by minimizing the objective function, selecting actions according to the current strategy, and interacting with the environment until a terminal state is reached. Secondly, at each time step t, the agent's experience is collected, and the network weights are corrected using stochastic gradient descent to continuously improve the strategy. Through relevant experiments, the optimal scheduling strategy can be effectively found to achieve efficient rescheduling of tasks and resource reallocation.
[0106] The beneficial effects of the present invention are as follows: the present invention provides a rescheduling method for edge service tasks based on HG-Trans, and utilizes an edge fault scenario task rescheduling strategy based on a heterogeneous graph neural network to comprehensively improve the task scheduling performance of the edge computing system in a fault scenario. This method solves the problem that the existing scheduling methods rely on local information or simple rules, lack an efficient fault recovery mechanism, and cannot quickly adjust the task execution path after a fault occurs, thereby reducing scheduling delays while achieving global optimal allocation of resources. Compared with the prior art, this method has the following effects:
[0107] (1) The task rescheduling problem is modeled as a Markov decision process and combined with a reinforcement learning strategy, which can dynamically adapt to node failures in edge computing scenarios and adjust the scheduling strategy in real time.
[0108] (2) A fault-tolerant mechanism is introduced into state modeling. The faulty nodes are automatically identified through the state representation of the heterogeneous graph structure, and the interrupted tasks are adaptively adjusted, thereby improving the execution success rate of task rescheduling.
[0109] (3) The Transformer-based heterogeneous graph embedding method can efficiently capture the complex relationship between tasks and computing resources, optimize scheduling strategies, reduce scheduling delays, and achieve global optimal allocation of resources.
[0110] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.
Claims
1. The rescheduling method of edge service tasks based on HG-Trans is characterized by: The following steps are included: S1. Model the dynamic task scheduling problem as a Markov decision process and convert the scheduling state into a heterogeneous graph structure; S2, input the heterogeneous graph structure into the HG-Trans network to obtain the estimated state value and the probability distribution of all actions; S3. Input the estimated state value and the probability distribution of all actions into the decision network and output the task scheduling strategy.
2. The method for rescheduling edge service tasks based on HG-Trans according to claim 1, characterized in that: In S1, the method of modeling the dynamic task scheduling problem as a Markov decision process is specifically as follows: (1) Setting the state: The comprehensive reflection of all task arrangements and computing resource states at any time constitutes the state at the corresponding time; (2) Set actions: define the actions at any time as TS pairs, where T is the task node and S is the server node; (3) Set the transition function: The transition function represents the probability of action a transitioning from the current state s to the next state s′. The transition function P a The specific expression of (s,s′) is: P a (s,s′)=P(s t+1 =s′|s t =s,a t =a) In the formula, s t is the state at time t, s t+1 is the state at time t+1, a t is the action at time t; (4) Set rewards: For any time, the reward function R (Makespan) is expressed as: Where R(·) is the reward function, Makespan is the maximum time span for task completion, Makespan(HEET) is the maximum time for HEFT to complete, and HEFT is the heterogeneous earliest completion time algorithm, which is the most commonly used heuristic algorithm in task scheduling.
3. The method for rescheduling edge service tasks based on HG-Trans according to claim 2, characterized in that: In S1, the heterogeneous graph structure H t =(T,S,ε t ), where ε t is the arc set of TS pairs, and each task node corresponds to a server node.
4. The method for rescheduling edge service tasks based on HG-Trans according to claim 3, characterized in that: The S2 comprises the following sub-steps: S21, input the heterogeneous graph structure into the multi-layer TransformerConv, update the node features, and generate node embedding; S22. Embed the nodes into a multi-layer perceptron to generate estimated state values and probability distributions of actions.
5. The method for rescheduling edge service tasks based on HG-Trans according to claim 4, characterized in that: In S21, the method of updating the features of each node by a single-layer TransformerConv includes the following steps: S211, input the heterogeneous graph structure into the multi-layer TransformerConv, perform weighted aggregation on the nodes and their neighboring nodes based on the Transformer self-attention mechanism, and calculate the relationship between node i and its first-order domain The attention coefficient A between nodes j in ij ; In the formula, Q i is the query vector, K j is the key vector, d k is the dimension of the key vector, T is the matrix transpose, S212, use the softmax function to normalize the attention coefficient in the neighborhood, and calculate the relationship between node i and its first-order domain The attention weight α between nodes j in ij ; S213: weighted aggregation of neighbor node information is performed through attention weights to update node features. The updated node features h i The specific expression is: In the formula, σ is a nonlinear activation function, V j is the feature representation of node j.
6. The method for rescheduling edge service tasks based on HG-Trans according to claim 5, characterized in that: In S22, the expression for generating the estimated state value V is specifically: Where N is the total number of nodes, W v is the projection matrix, W v ∈R d×1 , R is a matrix.
7. The method for rescheduling edge service tasks based on HG-Trans according to claim 6, characterized in that: In S22, the method for generating the probability distribution of the action includes the following steps: S221, embed the nodes of the available tasks and perform aggregation operations to generate a batch matrix H T ; H T =[h1,h2,...,h M ] T In the formula, h M Node embedding for the Mth task; S222, map the batch matrix into a one-dimensional vector space, generate a score for each task, and normalize the score through the Softmax function of the fully connected layer to obtain the probability distribution of the action, where the probability π of the lth action l The specific expression is: In the formula, o l It is the lth component in the Softmax function output vector o.
8. The method for rescheduling edge service tasks based on HG-Trans according to claim 7, characterized in that: The S3 is specifically: The policy gradient theorem optimizes the policy network. The objective function of the policy network is established based on the estimated state value and the probability distribution of all actions. The objective function is minimized to search for the optimal scheduling action. Among them, the expression of the objective function is specifically: In the formula, is the entropy function of the probability distribution, β is the hyperparameter that controls the effect of entropy regularization, and A(s,a) is the advantage function that measures whether action a is better than the average value in state s. t ,a t ,r t ,s t+1 ) and the current estimated value of the estimated state value V, π(a|s) is the policy function, is the gradient of the policy network parameter θ.
Citation Information
Patent Citations
MEC-oriented dependent task unloading method
CN117806730A
Self-adaptive task scheduling execution unit management method and system
CN119376903A
Cited By
Unmanned aerial vehicle group scheduling method and system based on multi-layer recurrent neural network
CN120746207A
Intelligent production scheduling decision method and system based on multi-oil product sequential transportation
CN122694137A