Multi-agent-based satellite edge task unloading and resource allocation method and device

By adopting the multi-agent deep reinforcement learning method in the satellite edge computing environment, the master-slave agent structure is built, and the problems of satellite communication and resource capacity limitation are solved, achieving more efficient task processing and resource allocation.

CN120104205APending Publication Date: 2025-06-06XIHUA UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510050105.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art fails to take into account the constraints of satellite communications and the resource capacity limitations of satellite edge computing servers, resulting in low task processing efficiency.

Method used

The deep reinforcement learning method based on multiple agents is adopted to build a master-slave agent structure, where the master agent is deployed on the satellite edge computing server, responsible for considering resource constraints and communication information, and the slave agent is distributed in user equipment, responsible for specific task offload decisions and resource allocation.

Benefits of technology

By balancing the overestimation and underestimation of the target network of multiple agents, learning efficiency is improved, communication and resource constraints of satellite edge computing servers can be effectively considered, service timeouts or resource waste can be avoided, and task processing efficiency can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104205A_ABST
    Figure CN120104205A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent-based satellite edge task unloading and resource allocation method. The method comprises the following steps: acquiring a task; constructing a system model, wherein the system model comprises user equipment, a satellite edge computing server and a cloud server; respectively calculating the delay and energy consumption of the user equipment, the satellite edge computing server and the cloud server for processing the task according to the system model; constructing a target function; constructing an unloading model based on multi-agent deep reinforcement learning; and according to the unloading model based on the multi-agent deep reinforcement learning and a preset constraint condition, optimizing the objective function, and obtaining an optimized task unloading and resource allocation strategy. According to the method, a master-slave multi-agent structure is adopted, communication and resource constraints of the star edge computing server are considered at the same time, and service timeout or resource waste is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network resource allocation, and in particular to a satellite edge task unloading and resource allocation method and device based on multi-agent. Background Art

[0002] Currently, MEC servers carried by LEO satellites have become a research hotspot in academia and industry. Considering the deployment of MEC platforms with computing and storage resources on low-orbit satellites, studies have solved the problems of joint service request scheduling and service placement. In response to the allocation problem in the task offloading process, many scholars have proposed various solutions, such as traditional optimization algorithms and scheduling algorithms based on deep reinforcement learning, to achieve better offloading decisions and effects. In the process of user equipment offloading tasks to satellite edge computing servers, if the server is overloaded, the task processing speed may be slow. In this case, introducing cloud data centers for cloud computing can significantly improve the task processing speed.

[0003] In order to efficiently process user equipment tasks in the satellite edge-cloud environment, many researchers have studied from two directions: traditional methods and deep reinforcement learning algorithms. For example: Computation offloading in satellite edge-cloud computing: A large number of researchers have studied computation offloading in satellite edge-cloud computing from traditional methods such as heuristic algorithms and search algorithms. These methods can solve long-term optimization problems, but cannot obtain prior information in dynamic environments, and can only obtain approximate optimal solutions to complex non-convex problems, and have limitations in obtaining optimal offloading decisions; Satellite edge-cloud computing based on DRL: In recent years, the field of machine learning has developed rapidly, and the decision-making ability of reinforcement learning and the perception ability of deep learning have been combined to form deep reinforcement learning. In view of the limitations of traditional optimization methods in obtaining optimal offloading decisions, more and more researchers have turned to DRL to solve the problem of satellite edge-cloud computing offloading. However, when using a single-agent reinforcement learning algorithm for task offloading, many researchers have failed to consider multiple constraints at the same time. Therefore, some scholars have proposed a multi-agent-based DRL task offloading method, which may obtain better global optimization effects to achieve optimal task processing efficiency.

[0004] Although the above studies have shown effectiveness in task processing and can quickly complete offloading decisions, they fail to simultaneously consider the constraints of satellite communications and the resource capacity limitations of satellite edge computing servers. Summary of the invention

[0005] The present invention proposes a multi-agent-based satellite edge task offloading and resource allocation method and device to solve the problem that the prior art fails to simultaneously consider the constraints of satellite communications and the resource capacity limitations of satellite edge computing servers.

[0006] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0007] A satellite edge task offloading and resource allocation method based on multi-agent, comprising:

[0008] Get the task;

[0009] Constructing a system model, wherein the system model includes a user device, a satellite edge computing server, and a cloud server;

[0010] According to the system model, respectively calculate the delay and energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task;

[0011] Constructing an objective function, wherein the objective function is to minimize the total delay and total energy consumption of the user equipment, wherein the total delay and the total energy consumption are the sum of the delay and the sum of the energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task;

[0012] Constructing an offloading model based on multi-agent deep reinforcement learning, the offloading model includes a master agent and a slave agent, the master agent is used to feed back the number of tasks offloaded by the user equipment, the resource capacity and communication information on the satellite edge computing server to the slave agent, and the slave agent is used to make task offloading decisions, calculate resources, allocate transmission rates, and adjust task offloading decisions of the user equipment according to the data fed back by the master agent;

[0013] The objective function is optimized according to the offloading model based on multi-agent deep reinforcement learning and preset constraints to obtain an optimized task offloading and resource allocation strategy.

[0014] Specifically, according to the system model, respectively calculating the delay and energy consumption of the user equipment in performing the task and the delay and energy consumption of the satellite edge computing server and the cloud server in unloading the task, including:

[0015] The delay and energy consumption of the user equipment in performing the task are calculated, and the calculation formula is as follows:

[0016]

[0017] The delay and energy consumption of the satellite edge computing server to offload the task are calculated using the following formula:

[0018]

[0019] The delay and energy consumption of the cloud server unloading the task are calculated using the following formula:

[0020]

[0021] Among them, D n Expressed as task data size, c n represents the number of CPU cycles used to process each bit of task for the user equipment, f n Indicates the computing resource capabilities of the user device, represents the delay of the user device in executing task n, represents the energy consumption of the user device when executing task n, d n Indicates the data transmission rate of a single channel of a wireless network, p n represents the transmission power budget, B represents the system bandwidth, K is the number of channels of the satellite edge computing server, represents the background noise variance, g n Indicates that the channel gain is affected by many factors, f e It is represented by the computing resource capacity of the satellite edge computing server, Indicates the transmission time, Indicates the task execution time. represents the energy consumption of the satellite edge computing server offloading task n, represents the delay of the satellite edge computing server unloading task n, f c represents the computing resource capability of the cloud server, represents the delay of offloading task n from the cloud server, Represents the energy consumption of offloading task n to the cloud server.

[0022] Specifically, an objective function is constructed, and the objective function is as follows:

[0023]

[0024] L n =αE n +(1-α)T n

[0025]

[0026] Where t represents the time, L n represents the system cost function, T n represents the total delay, E n represents the total energy consumption, Indicates the execution of task n on the user device. and represents the satellite edge computing server unloading task n, and represents the cloud server offloading task n, N is a natural number, and α is the weight coefficient of the preset system cost function.

[0027] Furthermore, the objective function is optimized according to the offloading model based on multi-agent deep reinforcement learning and preset constraints, including:

[0028] Training the master agent and the slave agent based on a multi-agent twin delayed deep deterministic policy gradient algorithm to obtain the trained master agent and the slave agent;

[0029] According to the trained slave agent, four continuous value actions in the range of [0, 1] are generated, and the action space is expressed as:

[0030]

[0031] in is the task offloading decision of user n, p n (t) is the user action that determines the transmission power, f n (t) is the action that determines the allocation of local computing resources, and the subordinate agent action p n and f n is determined as follows:

[0032]

[0033] Acquiring the state and action of the trained slave agent through the trained master agent, wherein the state of the trained slave agent includes: task state, channel gain state, power transmission state, local resource allocation state and battery capacity state;

[0034] The main agent selects actions and outputs the optimized action A m (t), the optimized action is the optimized task offloading and resource allocation strategy, and the selection formula is as follows:

[0035]

[0036] Furthermore, the main agent selects actions and outputs optimized actions, including:

[0037] like Then assign task n locally;

[0038] like and Then, task n is offloaded to the cloud server, and task n is processed by the cloud server;

[0039] Otherwise, A(t) is sent to the master agent through the slave agent for decision making, and the master agent outputs the optimized action based on the binary decision;

[0040] Among them, is the task offloading decision of the user at time t, p n (t) is the user action that determines the transmission power at time t, f n (t) is the action that determines the allocation of local computing resources at time t.

[0041] Further, the master agent and the slave agent are trained based on a multi-agent twin delayed deep deterministic policy gradient algorithm, including:

[0042] Initialization: setting parameters and initializing the network parameters of the master agent and the slave agent, including the learning rate, discount factor μ, and weight coefficient β;

[0043] Data sampling: select batches of data of size M from the experience pool, including state (S), action (A), target action (Amas), reward (r), next state (S') and termination flag (done);

[0044] Target action generation: Leverage the policy network π of each slave agent n , generate target action A' according to the next state S' n ;

[0045] Target state and action extraction: From the target action A' n Extract the binary uninstall decision x' (1) n and x' (2) n ;

[0046] Relative Q value calculation: If the target action x' (1) n and x' (2) n If both are 1, the two Q values ​​Q' are offloaded to the satellite edge computing server. n,1 and Q' n,2 ;

[0047] Q value list update: Q' n, The smaller value in the list is added to Q' N list.

[0048] Update the next Q value: if Q' N If the list is empty, calculate the two Q values ​​nextQ of the local processing task i,1 , nextQ i,2 , and add its minimum value to the next step nextQ i Otherwise, calculate the average Q value And Q' N The maximum value is added to nextQ i List;

[0049] The target Q value calculation formula is: i =r i +μβnextQ i,j +μ(1-β)Q′ avg ;

[0050] Error calculation: Calculate the gradient δ of the main agent;

[0051] Parameter update: Update the parameters θ of the master agent and the target master agent network parameters θ';

[0052] For each slave agent n, generate a new action A new n ;

[0053] If Q' N If the list is empty, calculate the Q value Q' of the local processing task loc and add it to Q' N List;

[0054] Otherwise, Q' N The maximum value in the list is added to the targQ list.

[0055] Calculate the gradient δ from the agent n .

[0056] Specifically, the constraints include:

[0057]

[0058] The formulas of the above constraints are expressed in sequence: whether the task is processed locally or uploaded to the server; the task is offloaded to the SEC or CC server; the transmission power should be within the power allocation budget; ensure that the local computing resources allocated to each task do not exceed the preset minimum and maximum limits; ensure that the battery power does not fall below the low power level; ensure that the processing time of each task does not exceed its specified processing period; establish the principle that a channel can only be used by one task at a time to prevent the number of offloaded tasks from exceeding the number of available sub-channels; ensure that the total amount of data of the offloaded tasks does not exceed the storage capacity of the satellite edge computing server.

[0059] Specifically, the penalty function of the offloading model based on multi-agent deep reinforcement learning is:

[0060]

[0061] Among them, λ 1 +λ 2 =1,λ 1 ,λ 2 are weight coefficients; τ nIndicates the deadline for processing task n. If the deadline is not exceeded, no penalty will be given. n The power of the user's device; The minimum power level is set to ensure that the battery power will not fall below the low power level. If it does not fall below the minimum power level, no penalty will be given.

[0062] Specifically, the reward function of the offloading model based on multi-agent deep reinforcement learning is:

[0063]

[0064] L′ n (t) is the penalty function at time t in the above penalty function.

[0065] The present invention also provides a multi-agent-based satellite edge task unloading and resource allocation device, which is characterized by comprising:

[0066] An acquisition module, wherein the acquisition module is used to acquire tasks;

[0067] A first building module, the first building module is used to build a system model, the system model includes a user device, a satellite edge computing server and a cloud server;

[0068] A calculation module, the calculation module is used to calculate the delay and energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task according to the system model;

[0069] A second construction module, the second construction module is used to construct an objective function, the objective function is to minimize the total delay and total energy consumption of the user equipment, the total delay and the total energy consumption are the sum of the delay and the sum of the energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task;

[0070] A third building block, the third building block is used to build an offloading model based on multi-agent deep reinforcement learning, the offloading model includes a master agent and a slave agent, the master agent is used to feed back the number of task offloading of the user equipment, the resource capacity and communication information on the satellite edge computing server to the slave agent, and the slave agent is used to make task offloading decisions, calculate resources, allocate transmission rates, and adjust task offloading decisions of the user equipment according to the data fed back by the master agent;

[0071] An optimization module, wherein the optimization module is used to optimize the objective function according to the offloading model based on multi-agent deep reinforcement learning and preset constraints to obtain an optimized task offloading and resource allocation strategy.

[0072] The beneficial effects of the present invention are:

[0073] The present invention adopts a master-slave multi-agent structure, in which the master agent is deployed on the satellite edge computing server, responsible for considering the constraints of the satellite edge computing server and controlling the task offloading and resource allocation of the slave agents. The slave agents are distributed in user devices and perform specific task offloading operations. By balancing the overestimation and underestimation problems of the updated target network obtained by multi-agents, the learning efficiency of the multi-agents is improved. At the same time, the communication and resource constraints of the satellite edge computing server are considered to avoid service timeout or resource waste. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 A diagram showing a satellite edge-cloud environment task offloading process in an embodiment of the present application;

[0075] Figure 2 A diagram showing the training process of the master agent and the slave agent in an embodiment of the present application;

[0076] Figure 3 Represents a reward convergence diagram under different algorithms in the embodiments of the present application;

[0077] Figure 4 A diagram showing a comparison of energy consumption and delay of different algorithms in an embodiment of the present application;

[0078] Figure 5 The performance diagrams of each algorithm under different numbers of UEs in the embodiments of the present application are shown as follows: (a) average cost, (b) average energy consumption, and (c) average delay;

[0079] Figure 6 The performance diagrams of each algorithm under different data amounts in the embodiments of the present application are shown as (a) average cost, (b) average energy consumption, and (c) average delay. DETAILED DESCRIPTION

[0080] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0081] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0082] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0083] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.

[0084] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.

[0085] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0086] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0087] like Figure 1 As shown, it is a diagram of the satellite edge-cloud environment task offloading process. The system model in this application is constructed according to this process. Figure 1 Here, SEC server refers to the SEC server, task1...taskn refers to task 1...task n, UserEquipment refers to user equipment, and Cloud Data Center refers to the cloud server:

[0088] User Equipment Model (UE)

[0089] Once the task is selected to be executed on the local UE, the present invention has Compute tasks locally n The desired delay can be described as:

[0090]

[0091] Where D nExpressed as task data size, c n Represented as UE n The number of CPU cycles to process each bit of task, f n Represented as UE n computing resource capabilities.

[0092] Compute tasks on UE n The required energy is given by formula (2):

[0093]

[0094] Satellite Edge Computing Server (hereinafter referred to as SEC Server)

[0095] In this mode, once you select a task n Offload to the SEC server, i.e. and The task will be offloaded to the SEC server and executed by the processing unit of the SEC server. When the UE submits the offload task and the master agent accepts the task, the corresponding decision will be triggered. In order to process the task on the SEC server, it is necessary to first allocate transmission resources to the user, and its transmission power is determined by the UE transmission power budget p n Then, the Shannon formula is used to calculate the data transmission rate d of a single channel of a wireless network. n , the formula is as follows:

[0096]

[0097] Where B is the system bandwidth, K is the number of SEC server channels, and the background noise variance and channel gain g n It is affected by many factors, including distance, satellite movement, etc. To simplify the calculation, the present invention assumes that the SEC server and UE are both in a stationary state and have a stable channel gain g n The background noise variance is also considered constant because it is assumed that each task occupies only one channel, thus ignoring the interference between multiple UEs.

[0098] When the task is offloaded to the server, the offloaded data does not need to be returned to the UE in the downlink, only the result needs to be transmitted back. Therefore, the response time and energy consumption in the downlink are much smaller than those in the uplink, so the response time and energy consumption in the downlink can be ignored. Once the data transmission rate is determined, the task n The response time is mainly affected by the data transmission time and task execution time, among which the transmission time The calculation formula is:

[0099]

[0100] In low-orbit satellite edge nodes, the propagation delay is relatively small and can be ignored. Therefore, the task execution time calculation formula is:

[0101]

[0102] where f e It is represented by the computing resource capacity of the SEC server. The energy consumption of task offloading is only calculated for ground UEs, which are all battery-powered. Then the energy consumed by UE in the satellite edge computing model is expressed as:

[0103]

[0104] Because the tasks have deadline constraints, the delay in processing the tasks on the server is important. The total delay in processing the tasks on the server is determined by the transmission time, the earliest available time of the processing unit on the server, and the time required to process the tasks on the server. The processing time of task n on the server is calculated as However, when a task arrives at the server, it is not processed immediately because the processing units on the server can only process one task at a time. Tasks transmitted to the server are processed in the order they arrive, determined by the order in which they are received. The task to be processed is assigned to the earliest idle processing unit. So the start of processing task n depends on the earliest availability of a processing unit, which is determined by the number of processing units on the server and Therefore, the total delay of task offloading task n to the SEC server is calculated as:

[0105]

[0106] in It indicates the estimated time of the first available processing unit in the server after task n arrives. It ensures that after the task is successfully unloaded, an idle processing unit can be found quickly and task processing can start. It is calculated based on the earliest time that task n can be unloaded after other tasks on the server have been accepted and completed.

[0107] Cloud Server (hereinafter referred to as CC Server)

[0108] When UE chooses to set task n Unload to CC server, i.e. and Similarly, at this time task n The energy consumption and delay offloaded to the CC server can be expressed as:

[0109]

[0110] where f c Indicates the computing resource capabilities of the CC server.

[0111] In this paper, the process of offloading tasks to CC servers involves first offloading tasks to SEC servers and then forwarding tasks to cloud computing centers for processing. Once again, this study focuses on the energy consumed by ground UEs, so when choosing to offload to CC servers, the energy consumption required by the device is expressed as:

[0112]

[0113] 2. Problem Definition

[0114] The goal of this paper is to minimize the energy consumption and delay of ground UE at the same time. The total delay and total energy consumption consumed by the user offloading task during the whole offloading process can be expressed as:

[0115]

[0116] in, is the total delay when task n chooses local processing, is the total delay when task n chooses to perform cloud computing, and is the total delay of task n when performing SEC. Similarly, E n is the total energy consumption when task n is processed in different places. The specific calculation methods are given in the previous section.

[0117] Therefore, the present invention defines the system cost function of processing tasks as L n , which is the weighted sum of the total delay and total energy consumption when processing tasks, and is specifically expressed as:

[0118] L n =αE n +(1-α)T n (12)

[0119] Where α is the weight coefficient of the system cost function. The optimization problem that this paper aims to solve is to minimize the energy consumption and delay of all UEs while satisfying different constraints of UE and server, as follows:

[0120]

[0121] Formula (13a) describes whether the task is processed locally or uploaded to the server. Formula (13b) further expresses whether the task is offloaded to the SEC or CC server. Formula (13c) indicates that the transmission power should be within a power allocation budget. Formula (13d) ensures that the local computing resources allocated to each task do not exceed the preset minimum and maximum limits. Formula (13e) ensures that the battery power does not fall below the low power level. Formula (13f) ensures that the processing time of each task does not exceed its specified processing period. Formula (13g) establishes the principle that a channel can only be used by one task at a time to prevent the number of offloaded tasks from exceeding the number of available sub-channels. Finally, formula (13h) ensures that the total amount of data of the offloaded tasks does not exceed the storage capacity of the SEC server.

[0122] 3. Algorithm

[0123] Since the optimization problem of formula (13) is NP-hard, the present invention restates it as a Markov decision process, and then proposes an offloading model based on MADRL to solve this problem. The present invention uses a master-slave multi-agent, in which the slave agent makes user offloading decisions and allocates computing resources and transmission rates, while the master agent adjusts the offloading decision based on the number of UE task offloading, the resource capacity on the SEC server, and the communication situation. The specific conditions of the state, user and master agent actions, and the reward function will be introduced below.

[0124] 3.1 Status

[0125] UE status S at time t n (t) and the SEC server state S(t), where S n (t) represents the set of UE states, and UE status S n (t) consists of five parts: task status Channel gain status Power transfer status Local resource allocation status and battery capacity status As defined in formula (14):

[0126]

[0127] in is the set of channel gains, It represents the UE transmission power p n gather, is the UE resource allocation f n gather, It indicates the battery capacity status b in the UEn gather.

[0128] 3.2 Action

[0129] In each step of performing task offloading and resource allocation,

[0130] Each user device formulates its resource allocation strategy with the help of slave agents.

[0131] Subsequently, the SEC server collects the status and action information of the UE and performs one of the following three processes according to its master agent: (1) For UEs that choose local processing or cloud computing, the SEC server does not intervene, that is, it allows them to freely execute decisions. (2) If the number of UEs that decide to offload to the SEC server exceeds the capacity of the communication channel, or the total amount of offloaded data exceeds the storage capacity of the SEC server, the master agent of the SEC server will evaluate and decide which offloading requests should be approved and which should be rejected to ensure the efficient use of system resources. (3) If the number of requests does not exceed the constraint, the SEC server will accept all offloading requests and assign them to the corresponding processing units for processing. In the task offloading and processing process, each channel is limited to a single UE at a time, but the same channel can be assigned to different UEs multiple times. In this scenario, the storage capacity and communication limitations of the SEC server become the main constraints for action selection. Therefore, the decision-making process of the slave agent and the master agent needs to fully consider these constraints to ensure the stability and efficiency of the system.

[0132] Slave agent actions: In each round, each slave agent produces four actions, which are all continuous valued actions in the range [0, 1]. The action space can be represented as:

[0133]

[0134] in is the task offloading decision of user n, p n (t) is the user action that determines the transmission power, f n (t) is the action that determines the allocation of local computing resources. Then the subordinate agent action p n and f n is determined as follows:

[0135]

[0136] Master Agent Action: The master agent obtains the status and actions of the slave agents and provides binary output for decision making, including which of them should be assigned locally and which should be processed by the SEC or CC server.

[0137]

[0138] The action selection in the master agent is decided by modifying the Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm. The slave agent represents the policy of the UE, the master agent represents the policy on the SEC server, and the environment represents the resource allocation on the UE and the SEC server. After the user produces its output, it executes as follows, if Then start local allocation; if and Then it will offload the task to the cloud service for processing; otherwise, it will It is forwarded to the master agent for decision making. The master agent then produces a binary decision and applies it to the UE and the server. Finally, the total reward is calculated and provided to the master agent to train its value function, while the slave agents are also trained using the master agent's calculation error as feedback.

[0139] 3.3 Rewards

[0140] To calculate the reward, the present invention uses the negative value of the objective function and calculates the penalty function for the task timeout and the battery capacity constraint of the UE.

[0141]

[0142] Among them, λ 1 +λ 2 =1,λ 1 ,λ 2 are weight coefficients; τ n Indicates the deadline for processing task n. If the deadline is not exceeded, no penalty will be given. n The power of the user's device; The minimum power is set to ensure that the battery power will not fall below the low power level. If it does not fall below the minimum power level, no penalty will be given. In the experiment of the present invention, if the total delay of the current unloading task exceeds the set task processing deadline or the total energy consumption exceeds the preset minimum battery threshold, a corresponding penalty will be given.

[0143] Since the design of the reward formula is based on the cost minimization problem in formula (13), the system reward function is equal to the negative value of the system cost function and the penalty function. The system reward function is expressed as:

[0144]

[0145] As mentioned above, is a discrete variable, p n , f n is a continuous variable, L ′n (t) is the penalty function at time t in the penalty function. The problem defined in formula (20) is a mixed integer nonlinear programming problem, i.e., an NP-hard problem. It is difficult to find the optimal solution for this problem using traditional optimization methods. Therefore, the present invention proposes a reinforcement learning algorithm to solve this problem.

[0146] like Figure 2 As shown, Figure 2 In the figure, slave agent represents the slave agent, and master agent represents the master agent. The slave agent represents the strategy of the UE, while the master agent represents the strategy on the SEC server. In the figure, rectangular symbols are used to represent the network structure, circular symbols represent state values, and dotted boxes identify sets or processes. Each slave agent has a Q network with the same architecture, which is able to access shared information and is used to estimate the value of the current action state. In addition, the Q networks of other slave agents act as external critics to evaluate the Q value of agent n and obtain the Q value by calculating the average value. avg Agent n passes min(Q 1,1 , Q n,1 ) and Q avg The weighted sum of the Q network is used to obtain the updated target y n .

[0147] Specifically, in this embodiment, the master agent and the slave agent are trained based on the multi-agent twin delayed deep deterministic policy gradient algorithm (hereinafter referred to as MS_M2ATD3), including:

[0148] Initialization: setting parameters and initializing the network parameters of the master agent and the slave agent, including the learning rate, discount factor μ, and weight coefficient β;

[0149] Data sampling: select batches of data of size M from the experience pool, including state (S), action (A), target action (Amas), reward (r), next state (S') and termination flag (done);

[0150] Target action generation: Leverage the policy network π of each slave agent n , generate target action A' according to the next state S' n ;

[0151] Target state and action extraction: From the target action A' n Extract the binary uninstall decision x' (1) n and x' (2) n ;

[0152] Relative Q value calculation: If the target action x'(1) n and x' (2) n If both are 1, the two Q values ​​Q' are offloaded to the satellite edge computing server. n,1 and Q' n,2 ;

[0153] Q value list update: Q' n, The smaller value in the list is added to Q' N list.

[0154] Update the next Q value: if Q' N If the list is empty, calculate the two Q values ​​nextQ of the local processing task i,1 , nextQ i,2 , and add its minimum value to the next step nextQ i Otherwise, calculate the average Q value And Q' N The maximum value is added to nextQ i List;

[0155] The target Q value calculation formula is: i =r i +μβnextQ i,j +μ(1-β)Q′ avg ;

[0156] Error calculation: Calculate the gradient δ of the main agent;

[0157] Parameter update: Update the parameters θ of the master agent and the target master agent network parameters θ';

[0158] For each slave agent n, generate a new action A new n ;

[0159] If Q' N If the list is empty, calculate the Q value Q' of the local processing task loc and add it to Q' N List;

[0160] Otherwise, Q' N The maximum value in the list is added to the targQ list.

[0161] Calculate the gradient δ from the agent n .

[0162] The present invention is further described below by means of specific examples. The test parameters of this example are shown in Table 1 below:

[0163] Table 1

[0164]

[0165] Experimental setup:

[0166] Due to economic, technical and hardware limitations, the proposed task offloading method is difficult to apply in real satellite scenarios. Therefore, the present invention establishes a simulated satellite edge-cloud environment and uses simulation experiments to verify the performance of the proposed method. In this experiment, all UEs are set to have the same battery capacity and resource allocation threshold, and the values ​​of these parameters are randomly generated from a uniform distribution. In order to simplify the experimental design, this study chooses to set the storage capacity limit of the SEC server in MB to more clearly examine the effect of task offloading. The design of the simulation experiment includes 40 slave agents and 1 master agent. The experiment runs 10 times in a row, and each run contains 500 iterations. In each iteration, the number of interactions between the agent and the environment is set to 10 times. After completing 10 runs, the present invention will draw a graph of the final result based on a 95% confidence interval. All experiments are conducted on a Windows device equipped with an Inteli5-12400F@2.5GHz, 32GBRAM and NVIDIAGeForceGTX4060GPU. The proposed simulation experiment is implemented with Python3.10 and torch1.11.0. In addition to the above settings, other parameter settings in the satellite edge-cloud environment are set according to the configurations in Table 1

[0167] Training process:

[0168] like Figure 2 As shown in the figure, the training process between the master agent and the slave agent of the MS_M2ATD3 algorithm is demonstrated. Among them, the slave agent represents the strategy of the UE, and the master agent represents the strategy on the SEC server. In this figure, rectangular symbols are used to represent the network structure, circular symbols represent state values, and dotted boxes identify sets or processes. Each slave agent has a Q network with the same architecture, which can access shared information and is used to estimate the value of the current action state. In addition, the Q networks of other slave agents act as external critics to evaluate the Q value of agent n and obtain Q by calculating the average value. avg Agent n passes min(Q 1,1 , Q n,1 ) and Q avg The weighted sum of the Q network is used to obtain the updated target y n .

[0169] Comparison algorithms:

[0170] The present invention compares the proposed MS_M2ATD3 algorithm with the following four algorithms:

[0171] (1) MATD3: When the server selects to offload, the main difference between this baseline method and the MS_M2ATD3 algorithm is that it assigns tasks to subchannels according to the order of task offloading, and tasks that are not assigned to any channel will be discarded. For discarded tasks, the agent in the user device will assign a timeout as a penalty. For fairness, the present invention compares the following adaptation method to ensure that the unaccepted tasks are equivalent to the MS_M2ATD3 algorithm in decision making.

[0172] (2) MATD3 with shortest offload time priority (OF_MATD3): This method is similar to MATD3, except that tasks not assigned to a subchannel or storage are designated as local processing instead of being directly discarded.

[0173] (3) MATD3 with size priority (SF_MATD3): This method is different from OF_MATD3 because it uses the increasing order of task size as the priority instead of offloading time.

[0174] (4) Multi-agent CCM_MADRL algorithm: This algorithm is also a multi-agent algorithm of master-slave agents, but it differs from the MS_M2ATD3 algorithm proposed in the present invention in that the MS_M2ATD3 algorithm takes into account the constraints on the SEC server during task offloading, and at the same time improves the learning efficiency of the multi-agent by balancing the overestimation and underestimation of the Q network to update the target network.

[0175] In addition, the present invention also compares the random unloading scheme in terms of convergence performance.

[0176] Comparison and analysis of experimental results

[0177] In order to verify the effectiveness of the algorithm proposed in this invention, a performance comparison experiment was conducted between this invention and the existing method. The experimental results are as follows:

[0178] (1) Experimental convergence and performance comparison:

[0179] Figure 3It clearly shows that during the 500 iterations of the experiment, the MS_M2ATD3 algorithm proposed in this study has achieved a significant improvement in the reward convergence speed compared to the best performing CCM_MADRL algorithm, with an increase of 50%. Specifically, the MS_M2ATD3 algorithm has reached convergence in the first 100 iterations, while the CCM_MADRL algorithm needs 200 iterations to achieve the same level of convergence. In addition, compared with the traditional MATD3 algorithm and its variants, the MS_M2ATD3 algorithm has a more significant advantage in reward convergence speed. This result fully demonstrates that the MS_M2ATD3 algorithm has efficient and rapid convergence capabilities in task offloading and resource allocation. Figure 3 The vertical axis represents the reward, and the horizontal axis represents the number of iterations (training episodes).

[0180] Figure 4 The average energy consumption and latency under different algorithms are compared. The MS_M2ATD3 algorithm proposed in this study performs well in reducing total energy consumption and latency, which effectively demonstrates its high efficiency in task offloading decisions. Compared with the MATD3 algorithm and its related algorithms, the MS_M2ATD3 algorithm shows obvious advantages. Specifically, the MATD3 algorithm and the OF_MATD3 algorithm perform the worst in energy consumption and latency, respectively. The MATD3 algorithm lags behind in minimizing energy consumption, while the OF_MATD3 algorithm has difficulty in effectively optimizing latency. The reason for this phenomenon is that when the offloading request exceeds the SEC server limit, the MATD3 algorithm chooses to directly discard the task, resulting in task processing timeout, while the algorithm still focuses on latency optimization and ignores energy consumption optimization. On the other hand, the OF_MATD3 algorithm sorts tasks according to the offloading time, thereby prioritizing energy consumption optimization. Both algorithms fail to process tasks based on the optimal decision. However, it can be seen from the figure that the MS_M2ATD3 algorithm still has shortcomings. Its average delay is slightly higher than that of the CCM_MADRL algorithm, but in terms of average energy consumption, the MS_M2ATD3 algorithm is significantly lower than that of the CCM_MADRL algorithm. Therefore, the algorithm of the present invention achieves a significant reduction in energy consumption by sacrificing a small part of the delay, thereby achieving a balance and improvement in overall performance.

[0181] (2) Performance of each algorithm under different numbers of UEs:

[0182] Figure 5 The performance comparison of the five algorithms is shown as the number of UEs changes. Figure 5As shown in (a), when the number of UEs is small, the average cost difference of the five algorithms is small (here the cost is the weighted sum of energy consumption and delay). This is because when the number of UEs is small, the number of tasks offloaded to the SEC server is also small, and the SEC server resource capacity and communication constraints are rarely exceeded, so the superiority of the MS_M2ATD3 algorithm is not fully demonstrated. Further analysis Figure 5 (b) and Figure 5 (c), the present invention observes that the average energy consumption and delay of all algorithms increase with the increase in the number of UEs. This phenomenon can be attributed to the increase in the number of UEs, that is, the increase in the number of agents, which slows down the convergence speed of each algorithm, thereby prolonging the time to obtain the optimal offloading decision, thereby increasing the average cost. Correspondingly, the energy consumption and delay of each algorithm also increase. It is worth noting that with the increase in the number of UEs, the delay growth rate of the MATD3 algorithm and its variants gradually accelerates. This is because with the increase in the number of UEs, the number of tasks offloaded to the SEC server also increases, and the situation of exceeding the resource capacity and communication constraints of the SEC server occurs frequently, and these algorithms do not fully consider these constraints, thereby triggering the penalty mechanism. In particular, in the MATD3 algorithm, the penalty caused by the direct abandonment of the task is more severe, and its delay growth rate is more significant. In contrast, the MS_M2ATD3 algorithm proposed in the present invention takes into account the resource capacity and communication constraints of the SEC server, so it can more effectively handle the task offloading problem when the number of UEs increases. This feature makes the MS_M2ATD3 algorithm show obvious advantages when dealing with large-scale UE scenarios.

[0183] Performance of each algorithm under different data volumes:

[0184] Figure 6 The effect of different task input data sizes on the performance of the five algorithms is demonstrated. Figure 6 The results in (a) show that as the amount of task data increases, the average cost of all algorithms shows an upward trend. Among them, the MS_M2ATD3 algorithm proposed in the present invention has a lower average cost than all other algorithms, thanks to its faster convergence speed, which can quickly determine the optimal offloading and thus reduce costs. In contrast, the MATD3 algorithm performs the worst. Its slow convergence speed and the task discarding phenomenon on the SEC server lead to an increase in cost. This discarding phenomenon is due to the fact that the MATD3 algorithm does not take into account the SEC server limitations. Further, as Figure 6 As shown in (b), with the increase of data volume, the advantage of MS_M2ATD3 algorithm in energy consumption optimization becomes more and more prominent, and its energy consumption is lower than that of other algorithms. Figure 6As shown in (c), the performance of this algorithm in terms of delay may not be as good as some algorithms, which is due to the sacrifice of delay when optimizing energy consumption. In general, the increase in the amount of task data leads to an increase in the demand for computing and communication resources. Under resource-constrained conditions, reasonable task offloading decisions are crucial to improving system performance. The MS_M2ATD3 algorithm of the present invention has significant advantages in task offloading decisions and can quickly and effectively determine the optimal offloading, so its average cost and energy consumption are lower than those of the other four algorithms.

[0185] The above-mentioned multi-agent twin-delayed deep deterministic policy gradient algorithm is an algorithm obtained by the applicant after improving the traditional MATD3 algorithm according to the needs. The improved algorithm effectively balances the overestimation and underestimation of the target network, thereby significantly improving the learning efficiency and overall performance of the multi-agent. On this basis, this study designed a master-slave multi-agent algorithm based on deep reinforcement learning. The algorithm adopts a master-slave architecture that combines centralized control with distributed execution. The master agent is located on the SEC server, responsible for handling server constraints and controlling the slave agent to perform task offloading and resource allocation; and multiple slave agents are deployed in user devices, responsible for implementing specific task offloading.

[0186] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A satellite edge task offloading and resource allocation method based on multi-agent, characterized in that: include: Get the task; Constructing a system model, wherein the system model includes a user device, a satellite edge computing server, and a cloud server; According to the system model, respectively calculate the delay and energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task; Constructing an objective function, wherein the objective function is to minimize the total delay and total energy consumption of the user equipment, wherein the total delay and the total energy consumption are the sum of the delay and the sum of the energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task; Constructing an offloading model based on multi-agent deep reinforcement learning, the offloading model includes a master agent and a slave agent, the master agent is used to feed back the number of tasks offloaded by the user equipment, the resource capacity and communication information on the satellite edge computing server to the slave agent, and the slave agent is used to make task offloading decisions, calculate resources, allocate transmission rates, and adjust task offloading decisions of the user equipment according to the data fed back by the master agent; The objective function is optimized according to the offloading model based on multi-agent deep reinforcement learning and preset constraints to obtain an optimized task offloading and resource allocation strategy.

2. According to the multi-agent-based satellite edge task offloading and resource allocation method of claim 1, it is characterized in that: The delay and energy consumption of the user equipment in executing the task and the delay and energy consumption of the satellite edge computing server and the cloud server in unloading the task are calculated according to the system model, including: The delay and energy consumption of the user equipment in performing the task are calculated, and the calculation formula is as follows: The delay and energy consumption of the satellite edge computing server to offload the task are calculated using the following formula: The delay and energy consumption of the cloud server unloading the task are calculated using the following formula: Among them, D n Expressed as task data size, c n represents the number of CPU cycles used to process each bit of task for the user equipment, f n Indicates the computing resource capabilities of the user device. represents the delay of the user device in executing task n, represents the energy consumption of the user device when executing task n, d n Indicates the data transmission rate of a single channel of a wireless network, p n represents the transmission power budget, B represents the system bandwidth, K is the number of channels of the satellite edge computing server, represents the background noise variance, g n Indicates that the channel gain is affected by many factors, f e It is represented by the computing resource capacity of the satellite edge computing server, Indicates the transmission time, Indicates the task execution time. represents the energy consumption of the satellite edge computing server offloading task n, represents the delay of the satellite edge computing server unloading task n, f c represents the computing resource capability of the cloud server, represents the delay of offloading task n from the cloud server, Represents the energy consumption of offloading task n to the cloud server.

3. The satellite edge task offloading and resource allocation method based on multi-agent according to claim 2 is characterized in that: Construct an objective function, which is as follows: Where t represents the time, L n represents the system cost function, T n represents the total delay, E n represents the total energy consumption, Indicates the execution of task n on the user device. and represents the satellite edge computing server unloading task n, and represents the cloud server offloading task n, N is a natural number, and α is the weight coefficient of the preset system cost function.

4. The satellite edge task offloading and resource allocation method based on multi-agent according to claim 3 is characterized in that: Optimizing the objective function according to the offloading model based on multi-agent deep reinforcement learning and preset constraints includes: Training the master agent and the slave agent based on a multi-agent twin delayed deep deterministic policy gradient algorithm to obtain the trained master agent and the slave agent; According to the trained slave agent, four continuous value actions in the range of [0, 1] are generated, and the action space is expressed as: in is the task offloading decision of user n, p n (t) is the user action that determines the transmission power, f n (t) is the action that determines the allocation of local computing resources, and the subordinate agent action p n and f n is determined as follows: Acquiring the state and action of the trained slave agent through the trained master agent, wherein the state of the trained slave agent includes: task state, channel gain state, power transmission state, local resource allocation state and battery capacity state; The main agent selects actions and outputs the optimized action A m (t), the optimized action is the optimized task offloading and resource allocation strategy, and the selection formula is as follows:

5. The multi-agent-based satellite edge task offloading and resource allocation method according to claim 4, characterized in that: The main agent selects actions and outputs optimized actions, including: like Then assign task n locally; like and Then task n is offloaded to the cloud server, and task n is processed by the cloud server; Otherwise, A(t) is sent to the master agent through the slave agent for decision making, and the master agent outputs the optimized action based on the binary decision; Among them, is the task offloading decision of the user at time t, p n (t) is the user action that determines the transmission power at time t, f n (t) is the action that determines the allocation of local computing resources at time t.

6. According to the multi-agent-based satellite edge task offloading and resource allocation method of claim 4, the master agent and the slave agent are trained based on the multi-agent twin delayed deep deterministic policy gradient algorithm, comprising: Initialization: setting parameters and initializing the network parameters of the master agent and the slave agent, including the learning rate, discount factor μ, and weight coefficient β; Data sampling: select batches of data of size M from the experience pool, including state (S), action (A), target action (Amas), reward (r), next state (S') and termination flag (done); Target action generation: Leverage the policy network π of each slave agent n , generate target action A' according to the next state S' n ; Target state and action extraction: From the target action A' n Extract the binary uninstall decision x' (1) n and x' (2) n ; Relative Q value calculation: If the target action x' (1) n and x' (2) n If both are 1, the two Q values ​​Q' are offloaded to the satellite edge computing server. n,1 and Q' n,2 ; Q value list update: Q' n, The smaller value in the list is added to Q' N List; Update the next Q value: if Q' N If the list is empty, calculate the two Q values ​​nextQ of the local processing task i,1 , nextQ i,2 , and add its minimum value to the next step nextQ i list, otherwise, calculate the average Q value And Q' N The maximum value is added to nextQ i List; The target Q value calculation formula is: i =r i +μβnextQ i,j +μ(1-β)Q′ avg ; Error calculation: Calculate the gradient δ of the main agent; Parameter update: Update the parameters θ of the master agent and the target master agent network parameters θ'; For each slave agent n, generate a new action A new n ; If Q' N If the list is empty, calculate the Q value Q' of the local processing task loc and add it to Q' N List; Otherwise, Q' N The maximum value in the list is added to the targQ list; Calculate the gradient δ from the agent n .

7. The satellite edge task offloading and resource allocation method based on multi-agent according to claim 4 is characterized in that: The constraints include, among others: The formulas of the above constraints are expressed in order: whether the task is processed locally or uploaded to the server; the task is offloaded to the SEC or CC server; the transmission power should be within the power allocation budget; ensure that the local computing resources allocated to each task do not exceed the preset minimum and maximum limits; ensure that the battery power does not fall below the low power level; ensure that the processing time of each task does not exceed its specified processing period; establish the principle that a channel can only be used by one task at a time to prevent the number of offloaded tasks from exceeding the number of available sub-channels; ensure that the total amount of data of the offloaded tasks does not exceed the storage capacity of the SEC server.

8. The multi-agent-based satellite edge task offloading and resource allocation method according to claim 7, characterized in that: The penalty function of the unloading model based on multi-agent deep reinforcement learning is: Among them, λ1+λ2=1, λ1, λ2 are weight coefficients; τ n Indicates the deadline for processing task n. If the deadline is not exceeded, no penalty will be given. n The power of the user's device; The minimum power level is set to ensure that the battery power will not fall below the low power level. If it does not fall below the minimum power level, no penalty will be given.

9. The multi-agent-based satellite edge task offloading and resource allocation method according to claim 7, characterized in that: The reward function of the offloading model based on multi-agent deep reinforcement learning is: L′ n (t) is the penalty function at time t in the above penalty function.

10. A satellite edge task unloading and resource allocation device based on multi-agent, characterized in that: include: An acquisition module, wherein the acquisition module is used to acquire tasks; A first building module, the first building module is used to build a system model, the system model includes a user device, a satellite edge computing server and a cloud server; A calculation module, the calculation module is used to calculate the delay and energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task according to the system model; A second construction module, the second construction module is used to construct an objective function, the objective function is to minimize the total delay and total energy consumption of the user equipment, the total delay and the total energy consumption are the sum of the delay and the sum of the energy consumption of the user equipment, the satellite edge computing server, and the cloud server in processing the task; A third building block, the third building block is used to build an offloading model based on multi-agent deep reinforcement learning, the offloading model includes a master agent and a slave agent, the master agent is used to feed back the number of task offloading of the user equipment, the resource capacity and communication information on the satellite edge computing server to the slave agent, and the slave agent is used to make task offloading decisions, calculate resources, allocate transmission rates, and adjust task offloading decisions of the user equipment according to the data fed back by the master agent; An optimization module, wherein the optimization module is used to optimize the objective function according to the offloading model based on multi-agent deep reinforcement learning and preset constraints to obtain an optimized task offloading and resource allocation strategy.

Citation Information

Cited By

  • Satellite internet constellation task unloading and resource allocation method based on multi-agent reinforcement learning

    CN121664260A