Multi-DAG satellite edge task scheduling and task unloading method and system
By using Dueling-DDQN and DDPG algorithms in the multi-DAG satellite edge task scheduling and task offloading methods, the task offload location and resource allocation are optimized, and the problem of high task processing delay in multi-node satellite networks is solved, achieving lower average completion delay and higher resource utilization efficiency.
Patent Information
- Application Number
- CN202510056107.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-02
AI Technical Summary
In a multi-node satellite network, how to effectively schedule and unload tasks to reduce task processing delays, especially when user equipment has limited computing power and satellite resources are limited.
A multi-DAG satellite edge task scheduling and task offloading method is proposed. By obtaining the DAG set of user tasks and the resources of satellite edge computing nodes, the Dueling-DDQN algorithm and DDPG algorithm that integrate priority experience playback are used to optimize the task offload location and resource allocation to minimize the average completion delay of the application.
It effectively reduces the average completion latency of the application, improves user experience, optimizes resource utilization efficiency, and shows better performance in complex task environments.
Smart Images

Figure CN119917244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite edge computing task offloading, and in particular to a multi-DAG satellite edge task scheduling and task offloading method and system. Background Art
[0002] Terrestrial Internet services are mainly concentrated in urban areas, while providing high-quality services in remote areas such as islands, oceans and deserts faces many challenges. Especially when natural disasters occur, the instability of terrestrial networks becomes more prominent. In contrast, satellite networks have been applied in emergency communications, navigation and positioning, and smart cities due to their significant features such as wide global coverage and strong disaster resistance. Satellite networks can effectively supplement and expand terrestrial networks to achieve seamless global coverage. For example, low-orbit satellite networks can provide stable and reliable services to densely populated areas and rural areas.
[0003] By deploying edge computing on satellites, data tasks can be processed directly on satellites, thereby reducing the transmission of data between satellites and ground nodes, and significantly reducing communication latency. Especially in areas where ground networks are sparse or remote, satellite edge computing can provide services efficiently and reliably, greatly improving users' communication experience. Therefore, the application of satellite edge computing has enabled satellite networks to play a more important and critical role in modern mobile communication systems, providing users with faster, low-latency services and optimizing user experience.
[0004] However, in a satellite network containing multiple nodes, due to the limited computing power of user mobile devices, when many user devices need to process tasks at the same time, users tend to assign as many tasks as possible to satellite edge computing nodes to reduce processing latency. However, given the limited satellite resources, in the satellite and ground network environment, how to develop efficient task scheduling and task computing offloading algorithms to reduce task processing latency is a major challenge facing current satellite network edge computing research.
[0005] Therefore, achieving efficient task scheduling and task offloading has become a core issue in satellite network edge computing research. Its solution will strongly promote the application and development of satellite edge computing in future mobile communication systems. Summary of the invention
[0006] The purpose of the present invention is to propose a multi-DAG satellite edge task scheduling and task offloading method and system, which can effectively solve the problem of excessive average completion delay of applications due to unreasonable task offloading strategy, thereby causing a decrease in user experience.
[0007] According to a first aspect of an embodiment of the present disclosure, a multi-DAG satellite edge task scheduling and task offloading method is provided, comprising the following steps:
[0008] Get the application DAG set M that the user needs to process and the available resources of the satellite edge computing node
[0009] Arrange all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, arrange them in descending order according to the amount of application DAG data.
[0010] Get the priority of each subtask in the application DAG and sort the subtasks according to the priority;
[0011] For the application DAG processed at the satellite edge node, the Dueling-DDQN algorithm with fusion-priority experience replay is used to obtain the optimal offloading location;
[0012] The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0013] In one embodiment, there are dependencies between the tasks of the application, and the structure of the dependencies is represented as a directed acyclic graph, namely, G i =(V i ,E i ), where V i represents the set of subtask nodes in the i-th application, V i =(v i,1 ,v i,2 …v i,n ); in the figure, for subtask v i,j , v i,j =(C i,j ,D i,j ), C i,j Indicates the CPU cycle required to process the task, in cycles, where D i,j Indicates the size of the current task in bits. The deadline of the application is expressed as E i Represents a set of directed edges between computing tasks The directed edges represent the dependencies between computing tasks. If there is a directed edge from vertex A to vertex B in the directed edge set, it means that the computing task of B must be processed after the computing task of A, that is, the end time of B cannot be earlier than that of A. The satellite edge nodes that provide task processing are S = {1,2...s}. When a task is offloaded, the MEC server will allocate resources to it until the task processing result is completed and then release the resources.
[0014] In one embodiment, the transmission delay of the subtask transmitted in the satellite network to be offloaded to the satellite edge computing server is expressed as:
[0015]
[0016] d 1,s represents the inter-satellite routing distance, represents the intersatellite transmission rate, Indicates the uplink transmission rate (unit: Mbps). It is expressed as:
[0017]
[0018] Subtask v i,j The transmission power is denoted as P i,j , the channel gain is denoted as g i,j , Gaussian white noise power is expressed as N, B i,j The bandwidth resources allocated;
[0019] The processing delay of offloading the subtask to the satellite MEC server is:
[0020]
[0021] Represents the computing resources allocated to a task.
[0022] In one embodiment, the latency of a local computing task includes processing latency and waiting latency. Since only a single task can be processed at a time when the task is processed locally, the waiting time of the task is the latency of the task in the task set v i,j The sum of all previous task calculation times:
[0023]
[0024] F l Indicates the local computing power (unit: cycle / s), Indicates waiting time;
[0025] The total delay for task completion is:
[0026]
[0027] where a i,j For the task processing method, the distance from the mobile device to the satellite node is represented as s, and the speed of light is represented as v.
[0028] Mission v i,j The actual completion delay is:
[0029]
[0030] in, Represents task v i,jThe maximum completion time of the predecessor task. When the task is the exit task of the DAG, its completion time is the completion delay of the entire application.
[0031] To ensure that the dependent tasks are executed first and the dependent tasks are executed later to satisfy the relationship between the tasks; the completion delay of the entire application is:
[0032]
[0033] It is necessary to jointly optimize task offloading decisions, scheduling decisions, and resource allocation:
[0034]
[0035] m is the total number of applications that need to be processed; when application i is within the time limit Inner and return the result, that is , where A and F represent the offloading strategy and resource allocation strategy; A is defined as the set of subtask offloading strategies, F is defined as a set of computing resource allocation strategies It is the computing resources that can be allocated to the satellite edge node s.
[0036] In one embodiment, the priority is expressed as follows:
[0037]
[0038] Among them, succ(v i,j ) represents the set of successor nodes of a node; the priority of the last node is represented by T i,exit Indicates; after obtaining the priorities of all subtasks, sort them to obtain the subtask sequence.
[0039] In one embodiment, the Dueling-DDQN algorithm with fusion priority experience playback is used to obtain the optimal unloading position, specifically:
[0040] Get the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t};
[0041] Randomly select an action A with probability ε t , or choose the best action for the current subtask;
[0042] Execute action A t , enter the new state S t+1 And get R t; The experience tuple (S t ,A t ,R t ,S t+1 ) into the experience pool;
[0043] Experience in the experience pool is calculated by the formula Conduct sampling and sample experience collection;
[0044] By formula The error is obtained by formula P i =σ+ε Update priority;
[0045] According to the formula Get the empirical importance sampling weights;
[0046] Get the Q value corresponding to the target network;
[0047] According to the formula loss = E[(Q tar -Q(s i ,a i ;θ)) 2 ]Get the loss value, and update the estimated network θ parameters through the gradient descent algorithm;
[0048] After each training with stride D, the estimated network parameters θ are assigned to the target network θ - ;
[0049] After training, the optimal unloading decision A is obtained.
[0050] In one embodiment, the resource to be allocated is obtained by using the DDPG algorithm as follows:
[0051] Get the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t};
[0052] Select action a through the critic network t =μ(s t );
[0053] Execute action a t , enter the new state s t+1 and obtain R t ;
[0054] Will (s t ,a t ,s t+1 ,R t ) experience samples are put into the experience pool, and N (st ,a t ,s t+1 ,R t ) as a minibatch experience;
[0055] Based on y t =R t +γQ(s t+1 ,μ(s t+1 )|θ Q ), through the formula L(θ Q )=Ε μ [(y t -Q(s t ,a t |θ Q )) 2 ] Calculate the loss value and get the parameter θ in the critic network Q ;
[0056] By formula Update actor network gradients;
[0057] By formula Update the target network parameters and obtain the optimal resource allocation strategy F.
[0058] According to a second aspect of an embodiment of the present disclosure, a multi-DAG satellite edge task scheduling and task offloading system is provided, including:
[0059] Acquisition module, which obtains the application DAG set M that the user needs to process and the available resources of the satellite edge computing node
[0060] The sorting module sorts all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, they are sorted in descending order according to the amount of application DAG data.
[0061] The priority module obtains the priority of each subtask in the DAG of each application and sorts the subtasks according to the priority;
[0062] The offloading module uses the Dueling-DDQN algorithm with fusion-first experience replay to obtain the optimal offloading location for the application DAG processed at the satellite edge node;
[0063] Resource allocation module,The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0064] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the method for multi-DAG satellite edge task scheduling and task offloading is implemented.
[0065] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for multi-DAG satellite edge task scheduling and task offloading is implemented.
[0066] The above technical solutions adopted by the present invention have the advantages of being compared with the prior art: the present invention focuses on the task scheduling problem in a multi-satellite environment, in which the tasks can be transmitted through intersatellite links. In view of the task requirements of multiple users, the present invention constructs an optimization model to minimize the average completion delay of the application in a multiple satellite node environment, and jointly optimizes the offloading strategy and the resource allocation strategy. On this basis, the present invention incorporates the completion delay into the definition of the subtask priority, and proposes a priority-based scheduling strategy to more reasonably determine the execution order of the subtasks. Further, in order to realize the computational offloading scheme, the present invention designs a Dueling-DDQN (i.e. Dueling Double Deep Q-Network) algorithm combined with priority experience replay. In view of the continuity of resource allocation actions, the present invention adopts the DDPG (i.e. Deep Deterministic Policy Gradient) algorithm to formulate the resource allocation strategy of the task. Compared with traditional schemes such as random processing and DQN algorithm, the algorithm proposed by the present invention shows better performance in reducing the average completion delay of the application. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0068] Figure 1 A schematic diagram of a satellite network architecture based on SDN provided by the present invention;
[0069] Figure 2 Different satellite computing resources provided for the present invention Comparison of the average completion delay of DAG in different scenarios;
[0070] Figure 3 Different local device computing resources F provided by the present invention l Comparison of the average completion delay of DAG in different scenarios;
[0071] Figure 4This is a comparison chart of the average completion delay of DAGs in different scenarios of the number of DAGs to be processed provided by the present invention;
[0072] Figure 5 This is a comparison chart of the average DAG completion delay in scenarios with different numbers of DAG subtasks provided by the present invention. DETAILED DESCRIPTION
[0073] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0074] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0075] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0076] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0077] Embodiment 1:
[0078] Satellite ground network architecture Figure 1As shown. The present invention focuses on considering the dependencies between tasks, scheduling subtasks based on the in-degree of the tasks, and the network architecture adopts a satellite ground network structure based on SDN. In this architecture, three GEO satellites are placed with SDN controllers. Multiple LEO satellites can deploy edge computing nodes to directly provide services to end users, which can shorten the response delay of terminal devices. The network scenario includes a variety of devices such as user terminals, data monitoring devices, and IoT devices connected to the LEO satellite network, and LEO satellites undertake computing services. Because each DAG task has a different priority. The processing order of tasks is derived according to different priorities.
[0079] Based on this idea, this embodiment provides a multi-DAG satellite edge task scheduling and task offloading method, including the following steps:
[0080] Step 1: Obtain the application DAG set that the user needs to process and the available resources of the satellite edge computing node;
[0081] The application DAG set consists of a set of tasks with dependency constraints and application deadlines. Assume that each application can be divided into multiple dependent tasks. Model the application as a DAG with an entry task. The entry task is simple and does not have any direct predecessor tasks, and the exit task does not have any direct successor tasks. Model a set of tasks in an application, and the application should be completed before the application deadline. There are a total of M mobile devices in the system, and the device number set is represented as M = {1,2,3…m}, indexed by i. At the beginning of each time slot, the mobile device will generate an application that needs to be processed. The application generated on a mobile device can be divided into multiple interrelated tasks, and the set of computing tasks is represented as N = {1,2,3…n}, indexed by j. Since there are dependencies between the tasks of the application, its structure can be represented as a directed acyclic graph, usually called a directed acyclic graph (DAG), that is, G i =(V i ,E i ), where V i represents the set of subtask nodes in the i-th application, V i =(v i,1 ,v i,2 …v i,n ). In the DAG graph, for subtask v i,j , v i,j =(C i,j ,D i,j ), C i,j Indicates the CPU cycle required to process the task, in cycles, where D i,j Indicates the size of the current task in bits. The deadline of the application is expressed as E iRepresents a set of directed edges between computing tasks The directed edges represent the dependencies between computing tasks. If there is a directed edge from vertex A to vertex B in the directed edge set, it means that the computing task of B must be processed after the computing task of A, that is, the end time of B cannot be earlier than the time of A. The satellite edge nodes that can provide task processing are S = {1,2......s}. When a task is unloaded, the MEC server will allocate resources to it until the task processing results are completed and then release the resources.
[0082] The time required for the subtask to be transmitted from the mobile device to the satellite edge computing server s. This transmission delay takes into account the time it takes for data to be transmitted in the satellite network. The transmission delay of the subtask offloaded to the satellite edge computing server can be expressed as:
[0083]
[0084] d 1,s Indicates the inter-satellite routing distance (single hop or multi-hop), represents the intersatellite transmission rate, Indicates the uplink transmission rate (unit: Mbps). It is expressed as:
[0085]
[0086] Subtask v i,j The transmission power is denoted as P i,j , the channel gain is denoted as g i,j , Gaussian white noise power is expressed as N, B i,j The allocated bandwidth resource (unit: MHz).
[0087] The processing delay of offloading the subtask to the satellite MEC server is:
[0088]
[0089] Indicates the computing resources allocated to the task (unit: cycle / s).
[0090] The processing delay of local computing tasks consists of processing delay and waiting delay. Since only one task can be processed at a time when the task is processed locally, the waiting time of the task is the task v in the task collection. i,j The sum of all previous task calculation times. The expression is:
[0091]
[0092] F l Indicates the local computing power (unit: cycle / s), Indicates the waiting time.
[0093] The total delay for task completion is:
[0094]
[0095] where a i,j For the task processing method, the distance from the mobile device to the satellite node is represented as s, and the speed of light is represented as v.
[0096] Mission v i,j The actual completion delay is:
[0097]
[0098] in, Represents task v i,j The maximum completion time of the predecessor task. When the task is the exit task of DAG, its completion time is the completion delay of the entire DAG application.
[0099] To ensure that the dependent tasks are executed first and the dependent tasks are executed later to satisfy the relationship between tasks. The completion delay of the entire DAG application is:
[0100]
[0101] Task offloading decisions, scheduling decisions, and resource allocation need to be jointly optimized.
[0102]
[0103] m is the total number of applications that need to be processed; when application i is within the time limit Inner and return the result, that is , where A and F represent the offloading strategy and resource allocation strategy. A is defined as the set of subtask offloading strategies, F is defined as a set of computing resource allocation strategies It is the computing resources that can be allocated to the satellite edge node s.
[0104] Step 2: Arrange all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, arrange them in descending order according to the amount of application DAG data.
[0105] Step 3: Get the priority of each subtask in the DAG of each application and sort the subtasks according to the priority;
[0106] There may be dependencies between each task of each DAG application. The execution of such tasks will be affected by the tasks associated with it. Therefore, the dependencies between tasks need to be considered before scheduling subtasks. Different types of applications processed by user requests have different requirements for completion time. Therefore, in order to reduce the processing delay of applications in the system and ensure that all applications in the system can be completed within their respective deadlines, it is necessary to first build a priority queue at the beginning. After receiving the application information, the applications are sorted according to priority. The priority of the application is expressed by the deadline delay. The shorter the deadline delay, the higher the priority. Since there are dependencies between subtasks, the tasks also need to be sorted to describe the dependencies between the current tasks. In order to meet the dependency constraints between subtasks in the application, we sort the tasks in descending order according to the order value of the tasks. The subtask priorities are as follows:
[0107]
[0108] succ(v i,j ) represents the set of successor nodes of a node. The priority of the last node can be expressed by T i,exit After obtaining the priorities of all subtasks, they are sorted to obtain the subtask priority sequence.
[0109] Step 4: For the application DAG processed at the satellite edge node, the Dueling-DDQN algorithm with fusion-first experience replay is used to obtain the optimal offloading location;
[0110] Since the state of the satellite-to-ground link in the satellite edge computing scenario is changing, the unloading strategy should be updated in time to reduce the latency cost of computing offloading. In the present invention, there are a total of s+1 unloading locations, and each satellite edge computing node can be allocated certain computing resources, and there may be different processing methods for the generated tasks. Deep reinforcement learning combines the capabilities of deep learning and reinforcement learning, and can solve the sequential decision-making problems of complex systems very well. Based on the Dueling-DDQN network, the task computing offloading problem is modeled as an MDP problem, and the action space A, state space S and reward function R are defined.
[0111] (1) State space
[0112] In the scenario studied in this embodiment, the state will change along with time, and the state can be defined as S t :
[0113] S t = {Da t ,Pr t ,F t}
[0114] Among themt Indicates the information of the DAG application currently being processed, including the subtask structure of the application, the CPU cycles required to process the task, and the size of the task. t is the priority sequence information of the application, F t ={F t,1 ,F t,2 …F t,s} is the available computing resources.
[0115] (2) Action Space
[0116] The action space should include the unloading location of the user's current task and the resource allocation result. The resource allocation can be obtained by the DDPG algorithm, and the action of the subtask is defined as if 0, indicating local processing. If the task v i,j Processed on the satellite, is 1. The set of all actions is A t
[0117] (3) Reward Function
[0118] The reward function is an important factor in obtaining the optimal Q network and is also the reason for the convergence of the algorithm. It is mainly aimed at minimizing the processing delay of the application within a period of time. The set reward function should be inversely proportional to the delay, so the reward R t It can be expressed as:
[0119]
[0120] The traditional DQN algorithm uses the maximization step, which leads to overestimation. The DDQN algorithm decouples the action selection and target value generation, making the model learning faster and more stable. m is defined as:
[0121]
[0122] θ is the parameter of the neural network.
[0123] The target Q value is expressed as:
[0124] Q tar =R t +γQ(s t+1 ,a m θ - )
[0125] θ - are the parameters of the target neural network, where γ represents the discount factor of the reward importance.
[0126] The Dueling Double DQN network structure is based on the Double DQN algorithm. Dueling DQN divides the Q value into two parts: the value function and the advantage function. The value represents the importance of the current state. This structure is suitable for dynamically changing environments. The Q value consists of the value function and the advantage function:
[0127] Q(s t ,a t ;θ)=V(s t θ V )+A(s t ,a t θ A )
[0128] The Dueling Network structure changes the way the Q value is calculated, which is an improvement on the DQN algorithm. The loss function obtained by DuelingDDQN is as follows:
[0129] loss=E[(Q tar -Q(s i ,a i ;θ)) 2 ]
[0130] The TD error represents the difference between the current estimated value and the target value, so the TD error can be expressed as:
[0131]
[0132] where θ and θ - Represent the parameter values of the estimated network and the target network respectively. The larger the empirical error, the greater the priority. The priority is defined as:
[0133] P i =|σ|+ε,ε>0
[0134] The larger the absolute value of the TD error of a sample, the higher the probability of it being sampled. The TD error of a sample determines the probability of being sampled. In order to solve problems such as low learning efficiency and overfitting, we can combine the random sampling method of pure greedy sampling and uniform distribution sampling, and ensure that the probability of sampling in the priority of training data is monotonic. Then the probability of sampling is defined as the priority sampling probability of the sample, which can be expressed as:
[0135]
[0136] in, is a hyperparameter used to adjust the priority. The sampling becomes uniform sampling.
[0137] The introduction of priority will cause the data distribution of samples to change, resulting in training bias or overfitting. In order to reduce this bias, the priority experience replay algorithm uses the importance sampling weight method to correct the bias. The importance sampling weight of the sample is defined as follows
[0138]
[0139] Therefore, the Dueling-DDQN algorithm is used to determine the unloading decision process:
[0140] Input the subtask sequence to be processed, the total amount of resources that can be allocated, and initialize the experience pool capacity to C; initialize the parameter values θ and θ of the estimation network and the target network - ; Initialize parameters ε, γ, ρ; Initialize the number of traversals X. and the update stride D of the target network;
[0141] S4.1 Obtain the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t};
[0142] S4.2 Randomly select an action A with probability ε t , or choose the best action for the current subtask;
[0143] S4.3 Execute action A t , enter the new state S t+1 And get R t ; The experience tuple (S t ,A t ,R t ,S t+1 ) into the experience pool;
[0144] S4.4 Experience of the experience pool through the formula Conduct sampling and sample experience collection;
[0145] S4.5 By formula The error is obtained by formula P i =σ+ε Update priority;
[0146] S4.6 According to the formula Get the empirical importance sampling weights;
[0147] S4.7 obtains the Q value corresponding to the target network;
[0148] S4.8 According to the formula loss = E[(Q tar -Q(s i ,ai ;θ)) 2 ]Get the loss value, and update the estimated network θ parameters through the gradient descent algorithm;
[0149] S4.9 After each training with a step length of D, the estimated network parameters θ are assigned to the target network θ - ;
[0150] S4.10 obtains the optimal unloading decision A after training.
[0151] Step 5: The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0152] The DDPG algorithm has symmetric characteristics. It follows the Actor-Critic architecture and can effectively handle problems with continuous action spaces by using deep neural networks to approximate strategies. The DDPG algorithm evaluates the quality of state-action pairs through policy functions and Q-value functions. The actor network makes action decisions based on the observed state, while the critic network is responsible for evaluating the actor's behavior. Each DDPG strategy μ defines an action-state pair value function Q μ (s t ,a t ), which means that given a state s t The expected return of the action when executed, the Q value is:
[0153] Q μ (s t ,a t )=Ε[R t +γQ(s t+1 ,μ(s t+1 ))]
[0154] The update parameters of the actor target network and the critic target network are θ μ and θ Q . Similar to the structure of DQN, the loss function of the critic network can be calculated as follows:
[0155] L(θ Q )=Ε μ [(y t -Q(s t ,a t |θ Q )) 2 ]
[0156] y t =R t +γQ(s t+1 ,μ(s t+1 )|θ Q )
[0157] The policy gradient can be calculated as follows:
[0158]
[0159] To minimize the loss function, the critic network Q will be updated by a given optimizer. After that, the actor network will act on the critic network in small batches, thereby achieving the gradient change of action a. The parameters (which can be derived by their own optimizer) use these two gradients and can be used to update the actor network with the formula
[0160]
[0161] The DDPG algorithm uses the current actor and critic network θ μ and θ Q , and soft-update it as follows:
[0162] θ Q' ←—θ Q +(1-τ)θ Q'
[0163] θ μ' ←—θ μ +(1-τ)θ μ'
[0164] The DDPG algorithm is used to solve the resource allocation problem with continuous variables. The process is as follows:
[0165] Initialize the parameters θ of the critic network and the actor network Q ,θ μ Initialize the target network parameter values Q' and μ' parameters, θ Q' ←—θ Q ,θ μ' ←—θ μ Initialize experience pool C;
[0166] S5.1 Obtain the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t};
[0167] S5.2 Selecting action a through the critic network t =μ(s t );
[0168] S5.3 Execute action a t , enter the new state s t+1 and obtain R t ;
[0169] S5.4 will (s t ,a t ,s t+1 ,R t ) experience samples are put into the experience pool, and N (s t ,a t ,s t+1 ,R t ) as a minibatch experience;
[0170] S5.5 based on y t =R t +γQ(s t+1 ,μ(s t+1 )|θ Q ), through the formula L(θ Q )=Ε μ [(y t -Q(s t ,a t |θ Q )) 2 ] Calculate the loss value and get the parameter θ in the critic network Q ;
[0171] S5.6 By formula Update actor network gradients;
[0172] S5.7 By formula Update the target network parameters and obtain the optimal resource allocation strategy F.
[0173] The number of CEO satellites is set to 3, the number of LEO satellites is set to 66, the computing resources of the LEO satellite edge computing server are set to 10Gcycles, the local computing resources are set to 2.5Gcycles, the data volume of the subtask is set between 50 and 100kb, the number of subtasks of each DAG is between 10 and 15, the deadline of the application is based on between 0.5 and 1.5s, and in order to verify the performance of the satellite edge resource allocation algorithm based on task dependence, in the simulation experiment, multiple random DAGs are generated in the satellite network for testing, and the DQN offloading method, all-on-satellite edge processing, all-on-ground processing, average resource allocation and random processing are respectively experimented with the relationship between the average completion delay of DAG and the changes in resources, processing methods, the number of DAGs and the number of DAG subtasks.
[0174] like Figure 2As shown in the figure, with the increase of satellite computing resources, the application completion delay of all schemes shows a downward trend. The task completion delay of DQN-based task processing (deep Q network and deep deterministic policy gradient, D-DDPG) and random processing (random processing, RA) is low, while the method proposed in the present invention (task processing combining Dueling-Double Deep Q-Network and deep deterministic policy gradient, D3-DDPG) is superior to other algorithms in completion delay. When the satellite computing resources gradually increase, the delay gap between D3-DDPG and RA is reduced from 0.37 seconds to 0.11 seconds. This is because as the satellite computing resources increase, the computing resources allocated to the task will also increase accordingly, thereby reducing the computing delay. Random decision-making is to randomly select the processing location, which fails to achieve the optimal effect, so the completion delay of the algorithm is relatively high. The method proposed in the present invention has a lower processing delay than D-DDPG, indicating that the improved DQN algorithm is effective. In addition, compared with the RA method, the D3-DDPG proposed in the present invention also shows a lower completion delay in processing delay, emphasizing the importance of appropriate task offloading decisions. Finally, the average allocation algorithm (AV) has a higher latency, indicating the importance of proper resource allocation for the tasks offloaded to the satellite.
[0175] like Figure 3 As shown in the figure, with the increase of local computing resources, the application processing delay of all schemes shows a downward trend. The completion delay of the task processing method (D-DDPG) based on the deep Q network (DQN) and the random processing (RA) is low, while the method (D3-DDPG) proposed in the present invention performs better in terms of completion delay, mainly due to the more computing resources on the user side, which effectively reduces the computing delay of the task. Random decision-making randomly selects the processing position, and its completion delay is higher because it does not pursue the optimal solution. The method of the present invention is better than D-DDPG in processing delay, indicating that the improved DQN algorithm has achieved significant improvement. However, with the increase of local computing resources, the delay difference between D3-DDPG and D-DDPG gradually narrows. This is because the improvement of this study is mainly aimed at satellite edge computing nodes, which enhances the ability of users to autonomously calculate tasks and reduces dependence on edge computing. In addition, the completion delay of the method of the present invention is lower than that of RA, emphasizing the importance of reasonable task offloading decisions to improve system efficiency. The solution that chooses to offload all tasks to the local (AL) has an increased completion delay compared with the method proposed in the present invention, which further proves the necessity of reasonable offloading in task processing.
[0176] like Figure 4As shown in the figure, as the number of applications to be processed increases, the application processing delay of all schemes shows an upward trend. In addition, the task processing method (D-DDPG) based on the deep Q network (DQN) performs better than the random processing (RA) in task processing delay. Although the completion delay of the method (D3-DDPG) proposed in the present invention also gradually increases, it always remains below other algorithms. This is because as the number of directed acyclic graph (DAG) tasks to be processed increases, the resources of the satellite edge nodes become more and more limited. Since the random decision does not seek the optimal solution, the completion delay of the algorithm is relatively high. The processing delay of the method proposed in the present invention is lower than that of D3-DDPG, which further proves the effectiveness of the improved algorithm based on DQN. At the same time, the completion delay of the method is also lower than that of RA, emphasizing the importance of appropriate task offloading decisions in improving system performance. In addition, the delay of the average allocation algorithm (AV) is higher than the average completion delay of the method proposed in the present invention, indicating that it is crucial to make reasonable resource allocation when offloading tasks to satellites.
[0177] like Figure 5 As shown in the figure, as the number of subtasks increases, the application processing delay of all schemes shows an upward trend. This phenomenon reflects the direct impact of the increase in task complexity on system performance, especially when resources are limited. The completion delay of the task processing method (D-DDPG) based on the deep Q network (DQN) and the random processing (RA) also increases when the number of tasks increases, showing the limitations of these two methods. However, the method (D3-DDPG) proposed in the present invention always maintains a low completion delay, which is better than other algorithms. This is mainly due to the fact that as the number of directed acyclic graph (DAG) subtasks increases, the resources of satellite edge nodes are increasingly limited, resulting in an increase in task calculation delay. D3-DDPG optimizes task offloading decisions and makes more efficient use of available resources, thereby reducing the overall calculation delay. The random decision method relies on randomly selecting the task processing location and does not seek the optimal solution, so the completion delay is high, further emphasizing the importance of reasonable decision-making in improving system performance. Compared with D-DDPG, the method of the present invention has lower processing delay, which verifies the effectiveness of the improved DQN algorithm, especially in resource-constrained environments. In addition, the completion delay of this method is lower than that of RA, which highlights the importance of appropriate task offloading in improving system performance. Reasonable task offloading can not only reduce delays, but also improve resource utilization efficiency. Finally, the delay of the average allocation algorithm (AV) is higher than that of the method proposed in this invention, which further shows the importance of reasonable configuration in resource allocation. By comparing different algorithms, it is found that optimizing resource allocation and task offloading strategies are crucial when dealing with complex tasks, which significantly improves the overall efficiency and response speed of the system.
[0178] Embodiment 2:
[0179] This embodiment provides a multi-DAG satellite edge task scheduling and task offloading system, including:
[0180] Acquisition module, which obtains the application DAG set M that the user needs to process and the available resources of the satellite edge computing node
[0181] The sorting module sorts all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, they are sorted in descending order according to the amount of application DAG data.
[0182] The priority module obtains the priority of each subtask in the DAG of each application and sorts the subtasks according to the priority;
[0183] The offloading module uses the Dueling-DDQN algorithm with fusion-first experience replay to obtain the optimal offloading location for the application DAG processed at the satellite edge node;
[0184] Resource allocation module,The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0185] Embodiment three:
[0186] An electronic device includes a memory, a processor, and a computer program stored and running on the memory, wherein the processor implements the above-mentioned multi-DAG satellite edge task scheduling and task offloading method when executing the program, including:
[0187] Get the application DAG set M that the user needs to process and the available resources of the satellite edge computing node
[0188] Arrange all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, arrange them in descending order according to the amount of application DAG data.
[0189] Get the priority of each subtask in the application DAG and sort the subtasks according to the priority;
[0190] For the application DAG processed at the satellite edge node, the Dueling-DDQN algorithm with fusion-priority experience replay is used to obtain the optimal offloading location;
[0191] The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0192] Embodiment 4:
[0193] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned multi-DAG satellite edge task scheduling and task offloading method, including:
[0194] Get the application DAG set M that the user needs to process and the available resources of the satellite edge computing node
[0195] Arrange all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, arrange them in descending order according to the amount of application DAG data.
[0196] Get the priority of each subtask in the application DAG and sort the subtasks according to the priority;
[0197] For the application DAG processed at the satellite edge node, the Dueling-DDQN algorithm with fusion-priority experience replay is used to obtain the optimal offloading location;
[0198] The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
[0199] Those skilled in the art should understand that the modules or steps of the present disclosure can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present disclosure is not limited to any specific combination of hardware and software.
[0200] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0201] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.
Claims
1. A multi-DAG satellite edge task scheduling and task offloading method, characterized in that: The following steps are involved: Get the application DAG set M that the user needs to process and the available resources of the satellite edge computing node Arrange all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, arrange them in descending order according to the amount of application DAG data. Get the priority of each subtask in the application DAG and sort the subtasks by priority; For the application DAG processed at the satellite edge node, the Dueling-DDQN algorithm with fusion-priority experience replay is used to obtain the optimal offloading location; The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
2. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: There are dependencies between the tasks of an application, and its structure is represented as a directed acyclic graph, G i =(V i ,E i ), where V i represents the set of subtask nodes in the i-th application, V i =(v i,1 ,v i,2 …v i,n ); In the figure , for subtask v i,j , v i,j =(C i,j ,D i,j ), C i,j Indicates the CPU cycle required to process the task, in cycles, where D i,j Indicates the size of the current task in bits. The deadline of the application is expressed as E i Represents a set of directed edges between computing tasks The directed edges represent the dependencies between computing tasks. If there is a directed edge from vertex A to vertex B in the directed edge set, it means that the computing task of B must be processed after the computing task of A, that is, the end time of B cannot be earlier than that of A. The satellite edge nodes that provide task processing are S = {1,2...s}. When a task is offloaded, the MEC server will allocate resources to it until the task processing result is completed and then release the resources.
3. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: The transmission delay of the subtask transmitted in the satellite network to the satellite edge computing server is expressed as: d 1,s represents the inter-satellite routing distance, represents the intersatellite transmission rate, Indicates the uplink transmission rate (unit: Mbps). It is expressed as: Subtask v i,j The transmission power is denoted as P i,j , the channel gain is denoted as g i,j , Gaussian white noise power is expressed as N, B i,j The bandwidth resources allocated; The processing delay of offloading the subtask to the satellite MEC server is: Represents the computing resources allocated to a task.
4. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: The latency of local computing tasks includes processing latency and waiting latency. Since only one task can be processed at a time when the task is processed locally, the waiting time of the task is the latency of the task in the task set v i,j The sum of all previous task calculation times: F l Indicates the local computing power (unit: cycle / s), Indicates waiting time; The total delay for task completion is: where a i,j For the task processing method, the distance from the mobile device to the satellite node is represented as s, and the speed of light is represented as v. Mission v i,j The actual completion delay is: in, Represents task v i,j The maximum completion time of the predecessor task. When the task is the exit task of the DAG, its completion time is the completion delay of the entire application. To ensure that the dependent tasks are executed first and the dependent tasks are executed later to satisfy the relationship between the tasks; the completion delay of the entire application is: It is necessary to jointly optimize task offloading decisions, scheduling decisions, and resource allocation: m is the total number of applications that need to be processed; when application i is within the time limit Inner and return the result, that is , where A and F represent the offloading strategy and resource allocation strategy; A is defined as the set of subtask offloading strategies, F is defined as a set of computing resource allocation strategies It is the computing resources that can be allocated to the satellite edge node s.
5. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: The priority is expressed as: Among them, succ(v i,j ) represents the set of successor nodes of a node; the priority of the last node is represented by T i,exit Indicates; after obtaining the priorities of all subtasks, sort them to obtain the subtask sequence.
6. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: The Dueling-DDQN algorithm with fusion priority experience playback is used to obtain the optimal unloading position, which is as follows: Get the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t }; Randomly select an action A with probability ε t , or choose the best action for the current subtask; Execute action A t , enter the new state S t+1 And get R t ; The experience tuple (S t ,A t ,R t ,S t+1 ) into the experience pool; The experience of the experience pool is calculated by the formula Conduct sampling and sample experience collection; By formula The error is obtained by formula P i =σ+ε Update priority; According to the formula Get the empirical importance sampling weights; Get the Q value corresponding to the target network; According to the formula loss = E[(Q tar -Q(s i ,a i ;θ)) 2 ]Get the loss value, and update the estimated network θ parameters through the gradient descent algorithm; After each training with stride D, the estimated network parameters θ are assigned to the target network θ - ; After training, the optimal unloading decision A is obtained.
7. A multi-DAG satellite edge task scheduling and task offloading method according to claim 1, characterized in that: The way to obtain the resources that need to be allocated through the DDPG algorithm is as follows: Get the sequence of subtasks to be processed according to the subtask scheduling priority; initialize the environment and obtain the current satellite edge computing network status S through the SDN controller t = {Da t ,Pr t ,F t }; Select action a through the critic network t =μ(s t ); Execute action a t , enter the new state s t+1 and obtain R t ; Will (s t ,a t ,s t+1 ,R t ) experience samples are put into the experience pool, and N (s t ,a t ,s t+1 ,R t ) as a minibatch experience; Based on y t =R t +γQ(s t+1 ,μ(s t+1 )θ Q ), through the formula L(θ Q )=Ε μ [(y t -Q(s t ,a t θ Q )) 2 ] Calculate the loss value and get the parameter θ in the critic network Q ; By formula Update actor network gradients; By formula Update the target network parameters and obtain the optimal resource allocation strategy F.
8. A multi-DAG satellite edge task scheduling and task offloading system, characterized in that: include: Acquisition module, which obtains the application DAG set M that the user needs to process and the available resources of the satellite edge computing node The sorting module sorts all application DAGs that need to be processed in ascending order according to their deadlines. If the deadlines are the same, they are sorted in descending order according to the amount of application DAG data. The priority module obtains the priority of each subtask in the DAG of each application and sorts the subtasks according to the priority; The offloading module uses the Dueling-DDQN algorithm with fusion-first experience replay to obtain the optimal offloading location for the application DAG processed at the satellite edge node; Resource allocation module,The tasks offloaded to the satellite edge nodes obtain the resources that need to be allocated through the DDPG algorithm.
9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the multi-DAG satellite edge task scheduling and task offloading method is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the multi-DAG satellite edge task scheduling and task offloading method is implemented.
Citation Information
Cited By
Priority scheduling method and system for on-orbit data processing tasks
CN121144051A
Prioritization scheduling method and system for on-orbit data processing tasks
CN121144051B