Dynamic scheduling optimization method and system for DAG application based on deadline constraint
By optimizing the selection strategy of task nodes and computing nodes through DAGTransformer encoder and self-supervised pre-trained model, the scheduling problem caused by the limited computing resources of terminal devices is solved, achieving efficient and energy-saving task scheduling and meeting the deadline constraints of industrial applications.
Patent Information
- Application Number
- CN202511275543.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-09
AI Technical Summary
Limited computing resources on terminal devices lead to unreasonable task scheduling in mobile edge computing architectures, which cannot meet the deadline constraints of industrial applications, resulting in latency and resource waste.
A dynamic scheduling optimization method based on deadline constraints for DAG applications is adopted. By combining a DAGTransformer encoder and a self-supervised pre-trained model with Markov decision process and reinforcement learning, a selection policy network for task nodes and computing nodes is constructed to optimize the scheduling scheme.
It achieves efficient and energy-saving task scheduling in multi-heterogeneous DAG application scenarios, shortens cold start time, improves model convergence quality and scheduling performance, and avoids invalid decisions on unreachable nodes and overloaded nodes.
Smart Images

Figure CN121092292A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of DAG application scheduling technology, and particularly relates to a dynamic scheduling optimization method and system for DAG applications based on deadline constraints. Background Technology
[0002] End devices typically have limited hardware capacity, supporting only simple local computing. For example, camera-equipped end devices move through discrete manufacturing plants to detect anomalies such as product stacking, material spillage, and worker tripping. These devices need to process large amounts of data to complete these tasks, often with stringent requirements for response time and energy consumption. Due to the limited computing resources of end devices, data-intensive tasks are currently mainly offloaded by cloud computing. However, traditional mobile cloud computing, while increasing the burden on the core network, cannot meet stringent latency requirements. To address this issue, mobile edge computing compensates for the shortcomings of cloud computing in Industrial Internet of Things (IIoT) applications.
[0003] While mobile edge computing alleviates latency issues, the increasing number of industrial terminal devices and mobile applications, coupled with the complex structures of most mobile applications, varying resource requirements, and heterogeneous application challenges such as application deadlines, still result in significant latency for terminal devices. Furthermore, resource contention among multiple mobile terminals can lead to unbalanced workload distribution across communication channels and processors, increasing potential latency. Inadequate scheduling can cause applications to fail to complete within deadlines, impacting industrial production efficiency. Traditional heuristic scheduling algorithms rely heavily on prior information about the system environment and cannot provide millisecond-level decision responses for real-time dynamic scheduling of mobile industrial applications, further increasing task latency, wasting resources, and reducing overall system efficiency. Summary of the Invention
[0004] The purpose of this application is to provide a dynamic scheduling optimization method and system for DAG applications based on deadline constraints, which is applicable to multi-heterogeneous DAG scheduling under deadline constraints and can complete the mobile task flow in a timely, efficient and energy-saving manner.
[0005] To achieve the above objectives, one aspect of this application provides a dynamic scheduling optimization method for multi-heterogeneous DAG applications based on deadline constraints, comprising: S1: Modeling a multi-time-slot dynamic system for mobile edge computing; The modeling includes: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling. S2: Construct a Markov decision process based on the modeling and set a reward function with respect to the deadline constraint; S3: Construct a DAG structured encoder and a self-supervised pre-trained model of the encoder, and train the encoder in the cloud; S4: Construct a task node action selection strategy network and a computing node action selection strategy network based on the encoder, and select the corresponding task node and computing node at a preset time step; S5: Input the state features obtained after encoding by the encoder into the Markov decision process, and output the scheduling scheme through the Markov decision process; S6: Collect execution data at the edge layer, train the task node action selection policy network and the computing node action selection policy network using the near-end policy optimization algorithm, and update the scheduling scheme according to the reward function.
[0006] Preferably, S1 further includes: S11: Construct a DAG application dataset by synthesizing a DAG generator. The DAG application dataset includes multiple DAG application graphs for representing heterogeneous applications. The parameters controlling the characteristics of the DAG application graphs include fat, density, and CCR. Among them, fat is used to control the width and height of the DAG, density is used to determine the number of edges between adjacent layers, and CCR is used to represent the ratio of communication cost to computation cost; S12: Perform structured parsing on the DAG application graph to extract subtask node information, dependency edge relationships between tasks, and adjacency matrix to form structured data for task scheduling and computation unloading. S13: Establish a Gauss-Markov trajectory model for the terminal device, and establish a wireless communication model, a task local computing model, a task edge computing model, and a task energy consumption model between the edge server and the terminal device to obtain the transmission latency, execution latency, waiting latency, and energy consumption between the edge server and the terminal device. S14: Based on the structured data and dynamic scheduling status information, construct the multi-DAG task node feature sequence, node depth matrix, task reachability matrix, schedulable mask matrix, and application feature matrix; S15: The feature sequence of the multi-DAG task nodes, the node depth matrix, the task reachability matrix, the schedulable mask matrix, and the application feature matrix are used as inputs to the task node action selection strategy network.
[0007] Preferably, S14 further includes: S141: Based on the structured data and dynamic scheduling status information, extract static features and dynamic features. The static features include the task number, its own data size, the size of intermediate data transmission, and the required computing resources. The dynamic features include dependency, schedulable status, and scheduling position, so as to form the observation feature representation of each task node in the current time slot, wherein nodes in the same row represent subtasks belonging to the same application. S142: The difference between the current system time slot and the application generation time is taken as the application dwell time, and combined with the application deadline and the number of remaining unfinished tasks in the application to form an observation feature representation of the application.
[0008] Preferably, S2 further includes: S21: Combine task node information and computation node information to serve as action representation in reinforcement learning; S22: Generate a Markov decision process based on the quadruple information consisting of state S, action A, reward R, and state transition probability P, and model the scheduling process as a Markov decision model.
[0009] Preferably, in the quadruple information: State S is the current state of the environment. S includes at least the state information after the task node observation features are encoded, the task mask vector, the computing node state information, and the computing node mask vector. Action A is the action performed by the agent, represented by a combination of a task node and a computing node; Reward R is the reward obtained by the agent after performing a certain action in the current state. It is set based on the local execution time and energy consumption of all tasks in the application on the terminal device. The reward is a weighted sum of the difference between the completion time of the current task and the completion time of the task when it is scheduled to be executed on the terminal device, and the difference between the cumulative energy consumption of the current application and the cumulative energy consumption of the application when the task is scheduled to be executed on the terminal device. This is the QoS utility function, and the calculation formula is as follows: ; in, and These are the time and energy consumption assuming that all tasks from the first to the nth are executed locally on the terminal device. and It represents the execution time and energy consumption of all tasks in the DAG on the local device. , , Let α and β represent the generation time of the application, α and β ∈ [0,1], and β represent the relative weights of the system optimization objectives. The other part transforms application-level deadline constraints into task-level immediate rewards or penalties, represented by the difference between the application's deadline and the current time slot, and the estimated remaining workload of the application. The calculation formula is as follows: ; in, The deadline for the application to which the task belongs. Indicates the current time slot of the system. To determine the number of remaining unfinished tasks, It is a fixed balance control factor; The reward function is the sum of the utility value and the slack at the deadline, and the calculation formula is as follows: ; in, To reward the balance coefficient.
[0010] Preferably, S3 further includes: S31: The task node DAG structured encoder uses the task depth position and reachability matrix in the DAG structure to perform structured representation of task nodes in each application. Among them, the DAG depth position encoding represents the depth of each dependent task node through sine and cosine position encoding, and is added to the node features as input to the Transformer layer. The calculation formula of position encoding is as follows: , ; in, Represents a node The depth; DAG reachability attention is based on the reachability matrix. , Indicates from node arrive There are paths that can reach each other. A single layer of multi-head attention and residual connection layer normalization and a feedforward neural network and residual connection layer normalization constitute a DAGTransformer layer, which are stacked and share the same reachability mask. The formula for calculating the attention weights between nodes is as follows: ; ; ; in, It is a node and nodes Attention score It's the number of heads that attract attention. It represents the vector dimension of each head, and exp is an exponential function with base e, used to enhance the effect of the activation function; Based on the obtained attention weights, the formula for calculating the attention of a single-layer DAG structure is as follows: ; ; in, , , , and For learnable parameters, This represents the value vector after the input features of the task node are merged and encoded at the position. For activation function, Representation layer normalization LayerNorm; S32: Establish a feature reconstruction loss function. For each masked task, the reconstruction loss is calculated using the following formula: ; in, It is the set of masked task indices. It is the first output of the encoder. The hidden representation of each task. It is the first The original feature vectors of each task It is a multilayer perceptron network; S33: Use a mask-based autoencoder architecture for multi-batch self-supervised training.
[0011] Preferably, S4 further includes: S41: The task node action selection policy network includes at least a pre-trained task node encoder and a pointer-based network decoder; S42: The task node encoder performs cross-application-level global attention encoding on the schedulable task node features after DAG structure encoding, and inputs the resulting variable-length task state encoding matrix into the decoder. The global attention encoding process applies a multi-layer linear mapping to the feature matrix, raising it to the same dimension as the task representation. Then, it fuses the task features and application features through gating to obtain the layer input, as shown in the following formula: ; Among them, the mixing coefficient From learnable scalar parameters Obtained through Sigmoid activation; Perform multi-head self-attention only on the schedulable set, for the th The layer calculation process is as follows: , , ; ; in, , , These are learnable parameters; For the The attention scores obtained from the layer are then subjected to residual calculation and normalization. The calculation process is as follows: ; ; in, Indicates the first The layer provides an enhanced representation of schedulable tasks, and Dropout is a regularization strategy to prevent overfitting. Representation layer normalization LayerNorm; S43: The pointer network decoder calculates the context vector based on the global task state feature encoding vector. The final query vector is obtained by projecting the pointer vector, and a matching projection is performed on each candidate embedding. , The unscaled attention score is obtained by scaling dot product matching, the final pointing probability distribution is obtained by the Softmax function, and a schedulable task is selected as the action by random sampling. S44: Obtain the original features of the corresponding encoding task node and the corresponding computing node based on the selected task node. After projection, gating fusion is performed to obtain The calculation process is as follows: in, It is the mixing coefficient. It is the selected task embedding; A multi-layer, multi-head, fully connected self-attention algorithm is used to calculate the fused node state features. A pointer network is used to calculate node scores. Unreachable computing nodes and resource-overloaded computing nodes are masked using a mask. A computing node is randomly selected as the action based on its score. The masking strategy is as follows: .
[0012] Preferably, S6 further includes: S61: Initialize the task node action selection strategy network, the compute node action selection strategy network, and the globally shared evaluation network parameters; S62: Input task coding status Enter the task action selection strategy network and input the compute node status. To the computing node action selection strategy network; in, The state features are obtained by passing the feature sequence of multiple DAG task nodes, node depth matrix, task reachability matrix, schedulable mask matrix, and applied feature matrix information through a feature encoder at the current time step t. This involves merging the information of the computational nodes that represent the characteristics of the selected task nodes at the current time step. S63: A complete schedule is constructed by decoding the order at each time step t, based on the local state embedded by the encoder at each time step t. and The task decoder selects a task node action. The compute node decoder selects a compute node action. To form a complete action = ( , ); S64: Collect data by interacting with the environment using actions, the data including state, actions and rewards, and update environmental state information; S65: Store status, action, and reward information in the experience pool, and update it when the experience pool reaches a preset size.
[0013] Preferably, S65 further includes: S66: Input the data from the experience pool into the policy neural network and the evaluation neural network, wherein the weight parameters of the policy neural network and the evaluation neural network are respectively... , Output actions respectively , and state value function The formula for calculating the ratio of new to old strategies is as follows: ; in, Input state under the action selection strategy of task node or compute node Take action in the following circumstances The probability when h=task, To select the set of all parameters in the policy network for the task action, when h = node, The set of all parameters in the policy network for selecting actions of computing nodes includes the weight matrix and bias terms that affect linear changes in the encoder and the attention mechanism weight matrix of the computing node, as well as the weight matrix and bias terms that affect linear changes in the decoder. S67: Calculate the advantage function using Generalized Advantage Estimation (GAE): ; ; in, Used to control the trade-off between bias and variance. It is the TD error term at step size t. In order to calculate the advantage function estimate with reduced variance, a state value network function V(s) with centralized shared parameters is applied. S68: Update the action selection strategy network parameters based on the dominance function and entropy regularization. ; ; ; in, To tailor the target of the agency, For trimming parameters, For the entropy target, The entropy ratio is the coefficient. The entropy of the probability distribution, This represents the average value of the current batch of samples; Among them, when In [1- 1+ When the objective function is within the range of ], it is the conventional policy gradient objective function; when Exceeding [1- ,1+ When the target function is within a certain range, the CLIP function limits its value to restrict policy updates. S69: The target loss for updating the value network to minimize the mean squared error (MSE) is calculated by the following formula: ; ; ; in, In order to reduce the loss of value, For trimming parameters, This is used to prevent significant changes to the output during a single update.
[0014] This application also provides a dynamic scheduling optimization system for DAG applications based on deadline constraints, using the above-described generation method, including: The modeling module is configured to model multi-slot dynamic systems for mobile edge computing; The modeling includes: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling. The decision module is configured to construct a Markov decision process based on the modeling and to set a reward function with respect to the deadline constraint; The encoder module is configured to build a DAG structured encoder and a self-supervised pre-trained model of the encoder, and train the encoder in the cloud. The strategy network module is configured to construct a task node action selection strategy network and a computing node action selection strategy network based on the encoder, and select the corresponding task node and computing node at a preset time step. The scheduling execution module is configured to input the state features obtained after encoding by the encoder into the Markov decision process, and output a scheduling scheme through the Markov decision process. The optimization module is configured to collect execution data at the edge layer, train the task node action selection policy network and the computing node action selection policy network using a near-end policy optimization algorithm, and update the scheduling scheme according to the reward function.
[0015] The technical solution provided in this application can achieve the following beneficial effects: 1. This application employs a DAGTransformer encoder combined with a self-supervised pre-training process. Compared to common graph neural network encoding methods, this encoder can not only capture static information such as the attributes and complex dependencies of task nodes in the DAG, but also extract dynamic features such as constraint variables and execution processes during dynamic scheduling. This provides high-dimensional abstract feature support for task priority selection in multi-DAG application scenarios. By designing a hybrid strategy of random masking, block masking, and depth masking, the encoder possesses good structure-aware initialization capabilities before reinforcement learning, effectively shortening the cold start time and improving the model's convergence quality and scheduling performance.
[0016] 2. This application further proposes a dual-agent pointer network decision-making and centralized Critic collaboration framework based on a gated fusion feature attention network. By adopting a decoupled pointer network structure that prioritizes task selection followed by node selection, and combining it with a centralized value network for unified global value evaluation, constraint awareness and action optimization of the task-node matching strategy are achieved. This scheme can effectively shield unreachable and overloaded nodes during scheduling, avoiding ineffective decisions, while also solving the problems of credit allocation and coupling instability in joint actions. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart of a dynamic scheduling optimization method for DAG applications based on deadline constraints provided in an embodiment of this application; Figure 2 This is an overall framework diagram of the dynamic scheduling optimization method provided in the embodiments of this application; Figure 3 This is a schematic diagram of a mobile edge computing offloading scenario provided in an embodiment of this application; Figure 4 This is an example diagram of a DAG synthesized from low attribute values to high attribute values provided in the embodiments of this application; Figure 5 This is a schematic diagram of the network structure of the DAG structure encoder in the dynamic scheduling optimization method provided in this application embodiment; Figure 6 This is a structural block diagram of a dynamic scheduling optimization system for DAG applications based on deadline constraints, provided in an embodiment of this application. Detailed Implementation
[0019] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0020] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0021] refer to Figure 1 , Figure 2 As shown, in some embodiments, one aspect of this application proposes a heterogeneous multi-farm machinery scheduling method based on deadline constraints, the method comprising the following steps: S1: Abstract modeling of multi-time-slot dynamic systems for mobile edge computing, including: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling.
[0022] like Figure 3 As shown, this Mobile Edge Computing (MEC) system includes multiple base stations supported by edge servers and multiple terminal devices. Wireless communication between the edge servers and terminal devices is achieved through the base stations, such as for task data transmission and status updates. Each terminal device has a communication module and a computing module, enabling communication with the edge servers via the base stations and allowing for local task execution. The edge servers also have communication and computing modules, and each has multiple core processors, with each core processor handling only one task at a time. The scheduling layer collects applications generated by all terminal devices in the current time slot and centrally makes scheduling and offloading decisions. Multiple scheduling operations can be performed in each time slot if resources are sufficient. Each task can be executed locally on the device or offloaded to an edge server for execution, depending on the scheduling policy.
[0023] S11: As Figure 4 As shown, the DAG application dataset construction uses a synthetic DAG generator to generate various DAGs representing heterogeneous applications. Its characteristics are controlled by multiple parameters such as fat, density, and CCR. Fat controls the width and height of the DAG, density determines the number of edges between two layers of the DAG, and CCR represents the ratio of communication cost to computation cost.
[0024] S12: The application data structure parsing process includes extracting key structural data such as subtask node information, dependency relationships between tasks, and adjacency matrices, thereby providing basic support for subsequent task scheduling and computation unloading.
[0025] S13: Establish a Gaussian-Markov trajectory model for the terminal device, and simultaneously establish a wireless communication model between the edge server and the terminal device, a task local computation model on the device, a task computation model on the edge server, and a task computation energy consumption model to obtain the transmission latency, execution latency, waiting latency, and energy consumption between the edge server and the terminal device.
[0026] The embodiment further proposes a preferred specific implementation method for step S13, including the following sub-steps: S131: Equipment movement model; Assume that the velocity and direction of terminal device m at time t are respectively and The update formula for the Gauss-Markov model is generally as follows: ; ; Where 0≤ ≤1 is a coefficient that quantifies the dependence on the previous state, capturing the inertia and orientation persistence of the terminal device. and These represent the average speed and direction of the terminal device, respectively. and It is an independent Gaussian random variable used to introduce random perturbations; Updated location coordinates of each terminal device The calculation is as follows: ; ; in, Representing the duration of a time slot, this model, through its balance of deterministic trends and stochastic variations, provides a robust framework for simulating the subtle and often unpredictable movements of end devices in dynamic and complex environments.
[0027] Step S132: Communication model building; As time slots change, the location of terminal devices may change. Each terminal device in the system communicates with the base station where the edge server is located via wireless communication. Assume that each base station implements an Orthogonal Frequency Division Multiple Access (OFDMA) communication model to manage uplink and downlink communication resources. OFDMA enables the base station to provide K non-interfering orthogonal resource blocks, each with the same bandwidth w. The available resource blocks of each base station are uniformly allocated to associated terminals, with each resource block assigned to a single terminal for communication. The maximum communication range of each base station is known to be range. In the MEC system proposed in this paper, the uplink wireless communication rate is defined as: ; in, This refers to the number of terminal devices within the coverage area of the base station where the edge server is located. The constant baseline transmit power of the terminal equipment, For channel power gain, For noise power, The distance between the terminal device and the base station. This is the path loss index. For fractional-order channel inversion control components; The average wireless transmission delay of the terminal device offloading subtasks and transmitting intermediate data to the edge server in time slot t. , They are as follows: ; ; in, Indicates selecting a task The action, Indicates selecting a task The actions corresponding to the tasks It is a task The precursor mission; The migration delay of a task is defined as the time required to transmit intermediate data along the wired link between base stations to the edge node where the subsequent task resides. This mainly depends on the amount of data to be migrated and the network bandwidth. Let B represent the fiber optic bandwidth, and the migration delay of the subtask... Defined as follows: ; S133: Computational modeling; Assume each edge server has 16 parallel processors with heterogeneous computing capabilities, denoted as... Each terminal device has only one processor, and its computing power is denoted as... Given that each processor can only process one node at a time, and each edge server can continuously receive applications from any terminal device within its coverage area, but it can only execute the same number of subtask nodes as its parallel processors simultaneously. Each node either executes locally on the terminal device or on a processor on the edge server. Therefore, the tasks... The execution time on the compute node can be calculated in two cases: ; in, It is the number of CPU cycles required per bit of task. Subtasks It is scheduled to be executed locally on the terminal device. Subtasks Unload and execute on the edge server; t represents the current time slot, and the task's waiting time can be modeled as follows: ; In this example, Forecast Availability Time (FAT) is used to assess the expected availability of resources. FAT depends on the completion time (CT) of the previous task in the processor. The task was scheduled to be executed on an edge server, and the total execution time was [not specified]. The earliest idle time update calculation for edge servers is as follows: ; ; If task k is scheduled to be executed on a terminal device, the execution start time is calculated as follows: ; ; Subtask The CT completion time calculation formula is: ; S134: Energy consumption modeling; Local processing tasks on terminal devices Energy consumption is determined by formula Given, among which This represents the total energy consumed by the terminal device per cycle, and the energy consumed by the terminal device to offload tasks to the edge server. The energy consumption of the edge server in receiving the task transmission data back to the terminal device is... ,in For the power consumption of terminal equipment transmission, It receives energy consumption.
[0028] S14: Based on the application structured parsing data and time-varying scheduling status information, construct a multi-DAG task node feature sequence, node depth matrix, task reachability matrix, schedulable mask matrix, and application feature matrix.
[0029] Specifically, in step S14, the method includes the following steps: S141: Based on the structured parsing data and time-varying scheduling status information of the application, static and dynamic features are extracted to obtain the sequence number, data size, intermediate data transmission size, required computing resources, dependency, schedulable status, and scheduling position of each task in all applications within the current time slot. These are then used as the observed feature representation of task nodes in the feature sequence of each DAG task node, where nodes in the same row represent subtasks belonging to the same application. S142: The difference between the current system time slot and the application generation time is taken as the application's dwell time, and combined with the application's deadline and the number of remaining unfinished tasks in the application as the application's observation feature representation.
[0030] S15: Use the feature sequence of multi-DAG task nodes, node depth matrix, task reachability matrix, task mask vector, and application feature matrix as input to the task selection strategy network.
[0031] S2: Construct a Markov decision process and define the reward function.
[0032] Specifically, in step S2, the method includes the following steps: S21: Combine task node information and computation node information as action representation in reinforcement learning; S22: Use the quadruplet information to generate a Markov decision process, and then model the scheduling process as a Markov decision process; In actual implementation, the four-tuple information is S, A, R, and P: S represents the current state of the environment, which includes at least the state information after the task node observation features are encoded, the task mask vector, the computing node state information, and the computing node mask vector. A represents the action performed by the intelligent agent, which is represented by a combination of a task node and a computing node; R represents the reward obtained by the agent after performing an action in the current state. It is set based on the local execution time and energy consumption of all tasks in the application on the terminal device. The reward is a weighted sum of the difference between the current task's completion time and the completion time when the task is scheduled to be executed on the terminal device, and the difference between the current application's cumulative energy consumption and the cumulative energy consumption when the task is scheduled to be executed on the terminal device. This sum is the QoS utility function, and the calculation formula is as follows: ; in, and These are the time and energy consumption assuming that all tasks from the first to the nth are executed locally on the terminal device. and It represents the execution time and energy consumption of all tasks in the DAG on the local device. , , The application's generation time. , ∈[0,1] represents the relative weights of the system optimization objectives. In practical scenarios, the values of the tuning parameters can be set according to user preferences and network conditions. For example, when users are more concerned about the real-time performance of the application, the latency weight can be appropriately increased. When device power is limited or network congestion occurs, the energy consumption weight should be increased. When edge server resources are scarce or wireless link quality fluctuates, the system status can be combined to balance the two in order to achieve the optimal trade-off between performance and energy consumption. Another part maps application-level deadline constraints to task-level immediate rewards or penalties. The metric used is the time difference between the application's deadline and the current time slot, minus the estimated remaining workload of the application. This metric characterizes the urgency of task scheduling and dynamically adjusts the reward signal accordingly. The calculation formula is as follows: ; in, The deadline for the application to which the task belongs. Indicates the current time slot of the system. To determine the number of remaining unfinished tasks, It is a fixed balance control factor; The reward function R is the sum of the utility value and the deadline slack. If the difference between the remaining task quantity and the deadline is insufficient, an additional penalty is incurred, causing the scheduler to favor scheduling actions that meet the deadline during policy updates. The calculation formula is as follows: ; in, As a reward balance coefficient, it is used to... Scale the item to make its magnitude the same as Maintaining consistency prevents any one part from becoming overly dominant in the reward function. The value range is positive real numbers, generally set between [0.01, 10]. The specific value can be determined based on the training convergence or the task. The importance of the constraint is determined by its degree; when you want to increase the impact of Slack on rewards, you can appropriately increase it. ,when When used only as a supplementary indicator, it can be reduced. ; P represents the probability of the environment state transition after the action is performed.
[0033] S3: Construct a self-supervised pre-trained model of a DAG structured encoder and train the encoder periodically in the cloud so as to form initial model parameters with topology awareness and structural understanding capabilities before actual scheduling, thereby reducing the cold start cost of online reinforcement learning and improving training stability. Specifically, refer to Figure 5 As shown, in step S3, the method specifically includes the following steps: S31: The task node DAG structured encoder uses the task depth location and reachability matrix in the DAG structure to perform structured representation of task nodes within each application. The encoder can simultaneously capture static information such as task node attributes and structural features such as DAG topological dependencies, providing high-dimensional abstract topology-aware signals for subsequent scheduling. The specific content is as follows: DAG depth position encoding represents the depth of each dependent task node using sine and cosine position encoding. This depth is added to the node features and used as input to the Transformer layer. This method introduces the hierarchical information of tasks in the topology during the initial representation stage, ensuring that the scheduling model can distinguish the priority relationships of tasks on different dependency paths. The calculation formula for the position encoding is as follows: , ; in, Represents a node The depth; DAG reachability attention is based on the reachability matrix. , Indicates from node arrive There are reachable paths between them. A single-layer DAG Transformer layer is formed by normalizing the multi-head attention layer with residual connections and the feedforward neural network layer with residual connections. These layers are stacked and share the same reachability mask. Compared to common graph neural networks (such as GAT), this approach avoids calculating attention for unreachable nodes, thus maintaining the causal relationship of the DAG and improving the rationality and efficiency of the structured representation. The formula for calculating the attention weights between nodes is as follows: ; ; ; in, It is a node and nodes Attention score It's the number of heads that attract attention. It represents the vector dimension of each head, and exp is an exponential function with base e, which enhances the effect of the activation function; Based on the obtained attention weights, the formula for calculating the attention of a single-layer DAG structure is as follows: ; ; in, To encode the input features and location of the task node into a vector The weight matrix mapped to the latent space. , These are the weight matrices for the first and second layers of the feedforward fully connected network, used to implement nonlinear transformations and feature reconstruction. and These are the bias vectors for the corresponding layers. Both the weight matrix and the bias vectors are learnable parameters, and their specific values are adaptively updated during the training process. This represents the value vector after merging and encoding the input of the task node. ReLU is the activation function, and LN represents the layer normalization layerNorm. S32: The design includes a feature reconstruction loss function. The masking method adopts a hybrid strategy of random masking, block masking, and depth masking, enabling the model to learn the deep connection between task features and topological context during training. For each masked task, the reconstruction loss formula is calculated as follows: ; in, It is the set of masked task indices. It is the first output of the encoder. The hidden representation of each task. It is the original feature vector of the i-th task. It is a multilayer perceptron network; S33: Using a mask-based autoencoder architecture for multi-batch self-supervised training, the encoder can learn structured semantics without manual annotation and has strong structure generalization ability. Combined with the mechanism of the mask autoencoder, the model has good initialization parameters during actual scheduling, which can significantly shorten the cold start period and improve convergence efficiency and stability. Before the model officially enters the online reinforcement learning scheduling, it has formed parameter initialization with DAG topology awareness and task feature understanding, so it can adapt to the actual dynamic constraint environment more quickly during online training, effectively decouple encoder training and reinforcement learning training, and reduce the instability caused by the coupling between the two. S4: Construct a task node action selection strategy network and a computing node action selection strategy network, and then select the corresponding task node and computing node at a preset time step; Specifically, refer to Figure 2 In step S4, the method further includes the following steps: S41: The task node action selection policy network includes at least a pre-trained DAG structured attention task encoder and a pointer-based network decoder; S42: The task node encoder network performs cross-application-level global attention encoding on all schedulable task node features after DAG structure encoding. The resulting variable-length task state encoding matrix is input into the decoder, ensuring that scheduling decisions not only depend on the local task structure but also integrate global deadlines and application features. This guarantees the effective transmission of deadline constraints in scheduling at the structural level. Specifically, the global attention encoding performs multi-layer linear mapping on the application feature matrix, elevating it to the same dimension as the task representation, to ensure stable additive fusion with task features. Then, the task features and application features are fused through gating to obtain the layer input. ; Among them, the mixing coefficient The value range of is (0,1), and it is determined by a learnable scalar parameter. Obtained by mapping through the Sigmoid activation function, when When the size is large, more emphasis is placed on application-level information. When the size is smaller, more emphasis is placed on the structural semantics of the task itself. Close to 1 (e.g.) When the value is ≥0.7, it indicates that the model focuses more on application-level information, such as when Close to 0 (e.g.) When the value is ≤0.3, it indicates that the model focuses more on the structural semantics of the task itself, while in the intermediate range (e.g., 0.3 < When <0.7), the model compromises between the two; Perform multi-head self-attention only on the schedulable set, for the th The layer calculation process is as follows: , , ; ; in, , , A trainable weight matrix used to represent the input. Projected onto the query, key, and value spaces for attention computation, the specific dimensions of these weight matrices are jointly determined by the input feature dimension and the attention latent space dimension, and their values are adaptively updated during training through the backpropagation algorithm. The LN representation layer is normalized to LayerNorm. The attention scores obtained from this layer are then subjected to residual calculation and normalization. The calculation process is as follows: ; ; in, The augmented representation of the schedulable task in layer l is represented by Dropout, which is a regularization strategy to prevent overfitting, and LN represents the layer normalization layerNorm. S43: The pointer network decoder calculates the context vector based on the global task state feature encoding vector. Then, the final query vector is obtained by projection onto the pointing vector, and a matching projection is performed on each candidate embedding. , The unscaled attention score is obtained by scaling dot product matching, and the final pointing probability distribution is obtained by the Softmax function. A schedulable task is selected as the action by random sampling. S44: Obtain the original features of the corresponding encoding task node and the corresponding computing node based on the selected task node. After projection, gating fusion is performed to obtain The calculation process is as follows: ; in, This is the mixing coefficient, which ranges from [0,1]. This parameter controls the input representation. With task-level embedding vectors The relative weights in the fusion result, when Close to 0 (e.g.) When the value is ≤0.3), the output representation focuses more on preserving the original features of the node itself. Close to 1 (e.g.) When the value is ≥0.7, the output representation focuses more on incorporating task-level global semantic information; in the intermediate range, the two parts of information are balanced and fused. The specific values can be adaptively updated during training through backpropagation. It is the selected task embedding; Similarly, a multi-layer, multi-head, fully connected self-attention approach is used to calculate the fused node state features. Node scores are calculated through a pointer network. Unlike the task node decoding process, when the pointer network outputs the action selection strategy, a mask is used to block unreachable and resource-intensive computing nodes from the terminal device. Then, a computing node is randomly selected as the action based on the score. This dynamic mask ensures that the model can still provide effective scheduling actions under dynamic resource changes and terminal mobility, avoiding the problem of frequent invalid actions generated by traditional RL methods. The masking strategy is as follows: ; S5: Input the state features obtained by encoding the observed features into the Markov decision process, and then output the scheduling scheme through the Markov decision process.
[0034] S6: Periodically collect execution data at the edge layer, use the near-end policy optimization algorithm to train the multi-discrete action selection policy network, and continuously optimize the task selection policy network and node selection policy network based on the reward of each round. Specifically, refer to Figure 2 As shown, in step S4, the method specifically includes the following steps: S61: Initialize the task node action selection strategy network, the compute node action selection strategy network, and the globally shared evaluation network parameters; S62: Input task coding status To the task action selection strategy network, input computing node status In the network for selecting action strategies for computing nodes, where, The state features are obtained by passing the feature sequence of multiple DAG task nodes, node depth matrix, task reachability matrix, schedulable mask matrix, and applied feature matrix information through a feature encoder at the current time step t. This involves merging the information of the computational nodes that represent the characteristics of the selected task nodes at the current time step. S63: The complete schedule is constructed by decoding the order at each time step t, based on the local state embedded by the encoder at each time step t. and The task decoder selects a task node action. The compute node decoder selects a compute node action. To form a complete action = ( , ); S64: Use actions to interact with the environment, collect the generated data, including {state, action, reward}, and update the environmental state information; S65: Store {status, action, reward} in the experience pool, and update it when the experience pool reaches a preset size.
[0035] In S65, the method further includes the following steps: S66: Input the data from the experience pool into the policy neural network and the evaluation neural network. The weight parameters of the policy neural network and the evaluation neural network are respectively... , Output actions respectively , and state value function Calculate the ratio of the new strategy to the old strategy: ; in, Input state under the action selection strategy of task node or compute node Take action in the following circumstances The probability of , where when h=task, To select the set of all parameters in the policy network for the task action, when h = node, The set of all parameters in the policy network for selecting actions of computing nodes includes the weight matrix and bias terms that affect linear changes in the encoder and the attention mechanism weight matrix of the computing node, as well as the weight matrix and bias terms that affect linear changes in the decoder. S67: Calculate the advantage function using Generalized Advantage Estimation (GAE) : ; ; in, To perform the action at time step t The instant reward obtained afterward ∈[0,1] is a discount factor used to control the degree of influence of future rewards on current decisions. The closer it is to 1, the more emphasis is placed on long-term rewards. Used to control the trade-off between bias and variance. For the cumulative step count index, The TD error term at step size t is used. To calculate the advantage function estimate with reduced variance, a state value network function with centralized shared parameters is applied. , for state The estimated value is used to provide a value benchmark and reduce the variance of the advantage estimate. The environmental state at time step t includes the agent's observation information and environmental information; S68: Update the action selection strategy network parameters based on the dominance function and entropy regularization. ; ; ; in, To tailor the target of the agency, For trimming parameters, For the entropy target, The entropy ratio is the coefficient. The entropy of the probability distribution, This represents the average value of the current batch of samples; Among them, when In [1- ,1+ When the objective function is within the range of ], it is the conventional policy gradient objective function; when Exceeding [1- ,1+ When the target function is within a certain range, the clip function limits its value to restrict policy updates. S69: Update the value network to minimize the mean squared error (MSE). The target loss is calculated by the following formula: ; ; ; in, In order to reduce the loss of value, For the clipping parameters, introduce This is to prevent a significant change in output during a single update, which could disrupt the stability of training.
[0036] S70: Continuously input data until the set iteration requirements are met.
[0037] This joint optimization mechanism ensures that task priority and resource location decisions can be updated collaboratively rather than independently in a dynamic environment, thereby improving overall scheduling performance. The reinforcement learning process involves first running an initial model to obtain a certain amount of data results, then comparing the results obtained with the old data, and updating the parameter information in the network based on the comparison results.
[0038] refer to Figure 6This application also provides a dynamic scheduling optimization system for DAG applications based on deadline constraints, using the above method, including: Modeling module 10 is configured to model a multi-slot dynamic system for mobile edge computing; The modeling includes: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling. Decision module 20 is configured to construct a Markov decision process based on the modeling and to set a reward function with respect to the deadline constraint; Encoder module 30 is configured to build a DAG structured encoder and a self-supervised pre-trained model of the encoder, and train the encoder in the cloud; The strategy network module 40 is configured to construct a task node action selection strategy network and a computing node action selection strategy network based on the encoder, and select the corresponding task node and computing node at a preset time step. The scheduling execution module 50 is configured to input the state features obtained after encoding by the encoder into the Markov decision process, and output a scheduling scheme through the Markov decision process. The optimization module 60 is configured to collect execution data at the edge layer, train the task node action selection strategy network and the computing node action selection strategy network using a near-end policy optimization algorithm, and update the scheduling scheme according to the reward function.
[0039] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A dynamic scheduling optimization method for DAG applications based on deadline constraints, characterized in that, include: S1: Modeling a multi-time-slot dynamic system for mobile edge computing; The modeling includes: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling. S2: Construct a Markov decision process based on the modeling and set a reward function with respect to the deadline constraint; S3: Construct a DAG structured encoder and a self-supervised pre-trained model of the encoder, and train the encoder in the cloud; S4: Construct a task node action selection strategy network and a computing node action selection strategy network based on the encoder, and select the corresponding task node and computing node at a preset time step; S5: Input the state features obtained after encoding by the encoder into the Markov decision process, and output the scheduling scheme through the Markov decision process; S6: Collect execution data at the edge layer, train the task node action selection policy network and the computing node action selection policy network using the near-end policy optimization algorithm, and update the scheduling scheme according to the reward function.
2. The dynamic scheduling optimization method according to claim 1, characterized in that, S1 also includes: S11: Construct a DAG application dataset by synthesizing a DAG generator. The DAG application dataset includes multiple DAG application graphs for representing heterogeneous applications. The parameters controlling the characteristics of the DAG application graphs include fat, density, and CCR. Among them, fat is used to control the width and height of the DAG, density is used to determine the number of edges between adjacent layers, and CCR is used to represent the ratio of communication cost to computation cost; S12: Perform structured parsing on the DAG application graph to extract subtask node information, dependency edge relationships between tasks, and adjacency matrix to form structured data for task scheduling and computation unloading. S13: Establish a Gauss-Markov trajectory model for the terminal device, and establish a wireless communication model, a task local computing model, a task edge computing model, and a task energy consumption model between the edge server and the terminal device to obtain the transmission latency, execution latency, waiting latency, and energy consumption between the edge server and the terminal device. S14: Based on the structured data and dynamic scheduling status information, construct the multi-DAG task node feature sequence, node depth matrix, task reachability matrix, schedulable mask matrix, and application feature matrix; S15: The feature sequence of the multi-DAG task nodes, the node depth matrix, the task reachability matrix, the schedulable mask matrix, and the application feature matrix are used as inputs to the task node action selection strategy network.
3. The dynamic scheduling optimization method according to claim 2, characterized in that, S14 also includes: S141: Based on the structured data and dynamic scheduling status information, extract static features and dynamic features. The static features include the task number, its own data size, the size of intermediate data transmission, and the required computing resources. The dynamic features include dependency, schedulable status, and scheduling position, so as to form the observation feature representation of each task node in the current time slot, wherein nodes in the same row represent subtasks belonging to the same application. S142: The difference between the current system time slot and the application generation time is taken as the application dwell time, and combined with the application deadline and the number of remaining unfinished tasks in the application to form an observation feature representation of the application.
4. The dynamic scheduling optimization method according to claim 1, characterized in that, S2 also includes: S21: Combine task node information and computation node information to serve as action representation in reinforcement learning; S22: Generate a Markov decision process based on the quadruple information consisting of state S, action A, reward R, and state transition probability P, and model the scheduling process as a Markov decision model.
5. The dynamic scheduling optimization method according to claim 4, characterized in that, In the quadruple information: State S is the current state of the environment. S includes at least the state information after the task node observation features are encoded, the task mask vector, the computing node state information, and the computing node mask vector. Action A is the action performed by the agent, represented by a combination of a task node and a computing node; Reward R is the reward obtained by the agent after performing a certain action in the current state. It is set based on the local execution time and energy consumption of all tasks in the application on the terminal device. The reward is a weighted sum of the difference between the completion time of the current task and the completion time of the task when it is scheduled to be executed on the terminal device, and the difference between the cumulative energy consumption of the current application and the cumulative energy consumption of the application when the task is scheduled to be executed on the terminal device. This is the QoS utility function, and the calculation formula is as follows: ; in, and These are the time and energy consumption assuming that all tasks from the first to the nth are executed locally on the terminal device. and It represents the execution time and energy consumption of all tasks in the DAG on the local device. , , Let α and β represent the generation time of the application, α and β ∈ [0,1], and β represent the relative weights of the system optimization objectives. The other part transforms application-level deadline constraints into task-level immediate rewards or penalties, represented by the difference between the application's deadline and the current time slot, and the estimated remaining workload of the application. The calculation formula is as follows: ; in, The deadline for the application to which the task belongs. Indicates the current time slot of the system. To determine the number of remaining unfinished tasks, It is a fixed balance control factor; The reward function is the sum of the utility value and the slack at the deadline, and the calculation formula is as follows: ; in, To reward the balance coefficient.
6. The dynamic scheduling optimization method according to claim 1, characterized in that, S3 also includes: S31: The task node DAG structured encoder uses the task depth position and reachability matrix in the DAG structure to perform structured representation of task nodes in each application. Among them, the DAG depth position encoding represents the depth of each dependent task node through sine and cosine position encoding, and is added to the node features as input to the Transformer layer. The calculation formula of position encoding is as follows: , ; in, Represents a node The depth; DAG reachability attention is based on the reachability matrix. , Indicates from node arrive There are paths that can reach each other. A single layer of multi-head attention and residual connection layer normalization and a feedforward neural network and residual connection layer normalization constitute a DAGTransformer layer, which are stacked and share the same reachability mask. The formula for calculating the attention weights between nodes is as follows: ; ; ; in, It is a node and nodes Attention score It's the number of heads that attract attention. It represents the vector dimension of each head, and exp is an exponential function with base e, used to enhance the effect of the activation function; Based on the obtained attention weights, the formula for calculating the attention of a single-layer DAG structure is as follows: ; ; in, , , , and For learnable parameters, This represents the value vector after the input features of the task node are merged and encoded at the position. For activation function, Representation layer normalization LayerNorm; S32: Establish a feature reconstruction loss function. For each masked task, the reconstruction loss is calculated using the following formula: ; in, It is the set of masked task indices. It is the first output of the encoder. The hidden representation of each task. It is the first The original feature vectors of each task It is a multilayer perceptron network; S33: Use a mask-based autoencoder architecture for multi-batch self-supervised training.
7. The dynamic scheduling optimization method according to claim 1, characterized in that, S4 also includes: S41: The task node action selection policy network includes at least a pre-trained task node encoder and a pointer-based network decoder; S42: The task node encoder performs cross-application-level global attention encoding on the schedulable task node features after DAG structure encoding, and inputs the resulting variable-length task state encoding matrix into the decoder. The global attention encoding process applies a multi-layer linear mapping to the feature matrix, raising it to the same dimension as the task representation. Then, it fuses the task features and application features through gating to obtain the layer input, as shown in the following formula: ; Among them, the mixing coefficient From learnable scalar parameters Obtained through Sigmoid activation; Perform multi-head self-attention only on the schedulable set, for the th The layer calculation process is as follows: , , ; ; in, , , These are learnable parameters; For the The attention scores obtained from the layer are then subjected to residual calculation and normalization. The calculation process is as follows: ; ; in, Indicates the first The layer provides an enhanced representation of schedulable tasks, and Dropout is a regularization strategy to prevent overfitting. Representation layer normalization LayerNorm; S43: The pointer network decoder calculates the context vector based on the global task state feature encoding vector. The final query vector is obtained by projecting the pointer vector, and a matching projection is performed on each candidate embedding. , The unscaled attention score is obtained by scaling dot product matching, the final pointing probability distribution is obtained by the Softmax function, and a schedulable task is selected as the action by random sampling. S44: Obtain the original features of the corresponding encoding task node and the corresponding computing node based on the selected task node. After projection, gating fusion is performed to obtain The calculation process is as follows: in, It is the mixing coefficient. It is the selected task embedding; A multi-layer, multi-head, fully connected self-attention algorithm is used to calculate the fused node state features. A pointer network is used to calculate node scores. Unreachable computing nodes and resource-overloaded computing nodes are masked using a mask. A computing node is randomly selected as the action based on its score. The masking strategy is as follows: 。 8. The dynamic scheduling optimization method according to claim 1, characterized in that, S6 also includes: S61: Initialize the task node action selection strategy network, the compute node action selection strategy network, and the globally shared evaluation network parameters; S62: Input task coding status Enter the task action selection strategy network and input the compute node status. To the computing node action selection strategy network; in, The state features are obtained by passing the feature sequence of multiple DAG task nodes, node depth matrix, task reachability matrix, schedulable mask matrix, and applied feature matrix information through a feature encoder at the current time step t. This involves merging the information of the computational nodes that represent the characteristics of the selected task nodes at the current time step. S63: A complete schedule is constructed by decoding the order at each time step t, based on the local state embedded by the encoder at each time step t. and The task decoder selects a task node action. The compute node decoder selects a compute node action. To form a complete action = ( , ); S64: Collect data by interacting with the environment using actions, the data including state, actions and rewards, and update environmental state information; S65: Store status, action, and reward information in the experience pool, and update it when the experience pool reaches a preset size.
9. The dynamic scheduling optimization method according to claim 8, characterized in that, S65 also includes: S66: Input the data from the experience pool into the policy neural network and the evaluation neural network, wherein the weight parameters of the policy neural network and the evaluation neural network are respectively... , Output actions respectively , and state value function The formula for calculating the ratio of new to old strategies is as follows: ; in, Input state under the action selection strategy of task node or compute node Take action in the following circumstances The probability when h=task, To select the set of all parameters in the policy network for the task action, when h = node, The set of all parameters in the policy network for selecting actions of computing nodes includes the weight matrix and bias terms that affect linear changes in the encoder and the attention mechanism weight matrix of the computing node, as well as the weight matrix and bias terms that affect linear changes in the decoder. S67: Calculate the advantage function using Generalized Advantage Estimation (GAE): ; ; in, Used to control the trade-off between bias and variance. It is the TD error term at step size t. In order to calculate the advantage function estimate with reduced variance, a state value network function V(s) with centralized shared parameters is applied. S68: Update the action selection strategy network parameters based on the dominance function and entropy regularization. ; ; ; in, To tailor the target of the agency, For trimming parameters, For the entropy target, The entropy ratio is the coefficient. The entropy of the probability distribution, This represents the average value of the current batch of samples; Among them, when In [1- ,1+ When the objective function is within the range of ], it is the conventional policy gradient objective function; when Exceeding [1- ,1+ When the target function is within a certain range, the CLIP function limits its value to restrict policy updates. S69: The target loss for updating the value network to minimize the mean squared error (MSE) is calculated by the following formula: ; ; ; in, In order to reduce the loss of value, For trimming parameters, This is used to prevent significant changes to the output during a single update.
10. A dynamic scheduling optimization system for DAG applications based on deadline constraints, using the method described in any one of claims 1-9, characterized in that, include: The modeling module is configured to model multi-slot dynamic systems for mobile edge computing; The modeling includes: DAG application dataset construction, application data structured parsing, terminal mobility process modeling, wireless communication modeling, resource dynamic change modeling, and task completion latency and execution energy consumption modeling. The decision module is configured to construct a Markov decision process based on the modeling and to set a reward function with respect to the deadline constraint; The encoder module is configured to build a DAG structured encoder and a self-supervised pre-trained model of the encoder, and train the encoder in the cloud. The strategy network module is configured to construct a task node action selection strategy network and a computing node action selection strategy network based on the encoder, and select the corresponding task node and computing node at a preset time step. The scheduling execution module is configured to input the state features obtained after encoding by the encoder into the Markov decision process, and output a scheduling scheme through the Markov decision process. The optimization module is configured to collect execution data at the edge layer, train the task node action selection policy network and the computing node action selection policy network using a near-end policy optimization algorithm, and update the scheduling scheme according to the reward function.
Citation Information
Cited By
AI bed separation method, medium and system for optimizing fabric cutting
CN121882594A
Cooperative scheduling method and system for edge computing resources, electronic equipment and storage medium
CN122195582A
AI agent dynamic task scheduling method and system based on multi-model collaborative reasoning
CN122285230A