Satellite internet constellation task unloading and resource allocation method based on multi-agent reinforcement learning
By constructing a satellite internet system architecture and adopting the self-attention improved MATD3 algorithm, the resource optimization problem of multi-beam low-orbit satellite systems was solved, achieving more efficient user services and system optimization with low resource consumption, and supporting the development of 6G satellite internet.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-03-13
AI Technical Summary
Existing multi-beam low-Earth orbit satellite systems face challenges in resource optimization, including scarce spectrum resources, inter-beam co-frequency interference, and limitations in transmit power and computing capabilities. Furthermore, current research has failed to effectively address the joint optimization problem of uplink and downlink.
A satellite internet system architecture is constructed using a multi-agent reinforcement learning approach. The multi-agent dual-delay deep deterministic policy gradient (MATD3) algorithm, improved with a self-attention mechanism, is used to optimize task offloading and resource allocation. Efficient resource allocation is achieved through a partially observable Markov decision process (POMDP) model.
It provides more user services with lower resource consumption, improves the system's convergence and robustness, and supports the development of 6G satellite internet.
Smart Images

Figure CN121664260A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of satellite communication technology, specifically relating to a method for task offloading and resource allocation in satellite internet constellations based on multi-agent reinforcement learning. Background Technology
[0002] Satellite internet has become a key technology in industry and academia, providing seamless global connectivity, bridging the digital divide, and supporting a wide range of applications. Unlike terrestrial networks, satellite internet covers remote areas, oceans, and airspace, providing reliable high-speed data services for education, healthcare, entertainment, emergency communications, and Internet of Things (IoT) applications. With the surge in demand for broadband access, satellite internet plays a crucial role in global connectivity and is an important component of the integration of fifth-generation (5G) and sixth-generation (6G) non-terrestrial networks (NTNs).
[0003] Multi-beam low-Earth orbit (LEO) satellite systems serve as core infrastructure for global broadband connectivity, providing low latency, high throughput, and flexible coverage to support global user access. LEO satellite constellations, including Starlink, employ multi-beam technology, utilizing narrow beams and high-gain antennas to achieve frequency reuse and regional coverage, thereby enhancing system capacity. Multi-beam design allows satellites to dynamically adjust beam direction and coverage to adapt to uneven user distribution and diverse needs, while Orthogonal Frequency Division Multiple Access (OFDMA) technology supports multi-user access. However, the limited resources of LEO satellite systems constrain their performance. Scarce spectrum resources necessitate efficient frequency reuse, which increases co-channel interference between beams, while limited transmit power and computing power restrict dynamic scheduling. Therefore, the complexity of multi-beam LEO satellite systems, combined with uneven user distribution and diverse needs, presents challenges for resource optimization.
[0004] Current research has made some progress in optimizing the resources of multi-beam satellite internet constellations. However, most of these studies only consider the optimization scenario of the downlink communication between satellites and users. Occasionally, they consider the joint optimization problem of uplink and downlink communication links, but they do not involve beam-level allocation schemes. Therefore, it is of great significance to study and explore the joint optimization problem of uplink and downlink resource allocation in multi-satellite, multi-beam low-Earth orbit satellite systems based on OFDMA technology. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention provides a satellite internet constellation task offloading and resource allocation method based on multi-agent reinforcement learning. First, the architecture of the satellite internet system is established, involving multiple satellites, multiple beams, and multiple ground users. Communication, computation, and task models for the uplink and downlink are modeled separately. Then, an optimization objective is defined: the satellite system supports multiple user services through appropriate task offloading and resource allocation while minimizing resource consumption. Next, the optimization problem is modeled as a partially observable Markov decision process (POMDP), and a multi-agent dual-delay deep deterministic policy gradient (MATD3) algorithm based on a self-attention mechanism is proposed to solve this problem. This invention can provide more user services with lower resource consumption and outperforms traditional multi-agent reinforcement learning methods in convergence and robustness tests, providing technical support for the development, design, and optimization of 6G satellite internet.
[0006] The technical solution adopted by this invention to solve its technical problem is as follows:
[0007] Step 1: Construct the network architecture of the satellite internet system, including multiple satellites, multiple beams and multiple ground users. Model the uplink and downlink communication links between satellites and users, satellite spectrum resources, power resources and computing resources respectively.
[0008] Step 2: Considering the diversity of user service needs in practical applications of satellite internet, three different types of user service needs are modeled: uplink computing task upload requests, downlink high communication rate requests, and combined uplink and downlink high communication rate requests.
[0009] Step 3: Derive the optimization objective from steps 1 and 2, namely, to enable the satellite system to support multiple user services while reducing resource consumption through task offloading and resource allocation;
[0010] Step 4: Model the optimization problem as a partially observable Markov decision process, and sample the improved multi-agent dual-delay deep deterministic policy gradient algorithm based on self-attention to solve the optimization problem; the improved multi-agent dual-delay deep deterministic policy gradient algorithm based on self-attention introduces an 8-head, 32-dimensional self-attention network before the hidden layer of the actor network of the multi-agent dual-delay deep deterministic policy gradient algorithm to improve the model's adaptability to dynamic environments.
[0011] Furthermore, step 1 specifically includes:
[0012] Step 1-1: Construct the network architecture of the satellite internet system, using multiple low-Earth orbit satellites to form a satellite system to provide internet services to users in remote areas;
[0013] The satellite system consists of a high-throughput satellite and multiple base satellites; the satellites provide internet service to multiple ground user units (GUs) in the form of multiple beams, and each GU is covered by one beam of a base satellite, so that all base satellites can cover the entire ground user base.
[0014] The basic satellite set is Each basic satellite is configured to have the same number of beams. The beam is from Index, use Indicates high-throughput satellite , This represents the entire set of satellites; the number of beams for high-throughput satellites is... , > The beam is from Index; by satellite The user set covered is The satellite system's service time is divided into multiple time slots, namely... ;
[0015] Assuming each GU covered by a satellite has one service request per time slot, and the service request varies across time slots, consider binary task offloading, meaning a user's service task can only be offloaded to one satellite—either the base satellite or a high-throughput satellite covering it—within a single beam. The unloading decision variable is represented as , , Indicates user The task is offloaded to the satellite ;
[0016] Steps 1-2: Construct the uplink and downlink communication model between the satellite and the GU under the OFDMA multiple access scheme;
[0017] Establish a channel model between ground users and base satellites; assume that each base satellite uses the same communication frequency band and has the same uplink and downlink bandwidth. , The beams employ full-frequency reuse technology, with a bandwidth of [missing information]. , All can be divided into The subcarrier bandwidth is [number], therefore the uplink subcarrier bandwidth is [number]. The downlink subcarrier bandwidth is Both uplink and downlink subcarriers are generated by index;
[0018] Let the downlink subcarrier decision variables be... , , Indicates subcarrier Assigned to satellite Beam Users under coverage ;
[0019] satellite Beam Users below The downlink transmission rate is expressed as:
[0020]
[0021] in, User Downlink subcarrier Signal-to-interference-plus-noise ratio (SINR) on the surface; This represents the subcarrier bandwidth of the downlink;
[0022] Similarly, let the subcarrier decision variables of the uplink be... Assume the total transmission power of all ground users is a constant. ,satellite Beam Users below Set the transmit power of each allocated subcarrier to :
[0023]
[0024] Then satellite Beam Users below The uplink transmission rate is expressed as:
[0025]
[0026] in, User uplink subcarrier The signal-to-interference-to-noise ratio (SIR) is... This represents the uplink subcarrier bandwidth;
[0027] Next, establish connections between ground users and high-throughput satellites. Channel model; uplink subcarrier bandwidth Downlink subcarrier bandwidth The uplink and downlink subcarriers are respectively composed of and index; , These represent the total bandwidth of the uplink and downlink, respectively. , These represent the number of uplink and downlink subcarriers, respectively.
[0028] Uplink and downlink subcarrier decision variables are defined as follows: and And there are and downlink subcarrier Transmit power is defined as , If subcarrier If not occupied by a user, then ,user The downlink transmission rate is expressed as:
[0029]
[0030] user The uplink transmission rate can be expressed as:
[0031]
[0032] in, These represent the uplink and downlink subcarrier bandwidths, respectively. , They represent the uplink and downlink subcarriers, respectively. Signal-to-interference-plus-noise ratio;
[0033] Steps 1-3: Establish the normalized transmit power of subcarriers in each satellite beam;
[0034] Satellite The upper limit of the transmission power is , , and have To obtain basic satellites and satellites Normalized transmit power of subcarriers in the beam:
[0035] ,
[0036] ,
[0037] in, Indicates the basic satellite subcarrier The transmission power, Indicates high-throughput satellite subcarriers The transmission power; These represent the basic satellite and high-throughput satellite subcarriers, respectively. Normalized transmit power;
[0038] Steps 1-4: Establish computational resource and mission latency models for each satellite;
[0039] The satellite's computing resources are its maximum achievable computing frequency, i.e., the number of CPU cycles per second. Assuming... For satellite The maximum achievable computation frequency, then For users If there is a computational task, it is represented as , This indicates the amount of data that needs to be transferred for the computation task. This indicates the amount of computing resources required to complete the task. The maximum tolerable delay for the mission; for offloading the mission to the satellite Beam users The total delay for completing the computation task is expressed as:
[0040]
[0041] in, This indicates the proportion of total computing power resources allocated. Indicates the physical distance between the satellite and the user. It is the speed of light.
[0042] Furthermore, step 2 specifically includes:
[0043] Step 2-1: Based on the typical 5G service types proposed by the International Telecommunication Union (ITU), the services provided by satellite internet are designed into three types: URLLC type service, eMBB1 type service, and eMBB2 type service.
[0044] First, the user Service requirements modeled as quintuples , These represent the user's uplink and downlink communication rate requirements thresholds, respectively; the user set required for URLLC type services is set to... The user set that requires eMBB1 type services and eMBB2 type services is set as follows: ;
[0045] URLLC type services: , The latency of the user's computation task needs to be less than the task's tolerance latency, that is: , , Indicates the tolerable latency of the task;
[0046] eMBB1 type service: high bandwidth and high throughput of satellite-to-ground downlink, service requirements modeled as follows , For offloading tasks to satellites Beam users The downlink transmission rate should be greater than the threshold. ,Right now: ;
[0047] eMBB2 type service: simultaneously meets the high throughput requirements of both uplink and downlink; service requirements are modeled as follows: , For offloading tasks to satellites Beam users Uplink and downlink transmission rates should be greater than the threshold. ,Right now: , ;
[0048] Step 2-2: Construct the expression for the resource cost and optimization problem of the satellite system;
[0049] In time slot t, the system resource cost is defined as follows:
[0050]
[0051] in, , These represent the spectrum resource cost, transmit power resource cost, and computing resource cost of the satellite system, respectively.
[0052] The service breach cost of the system is defined as:
[0053]
[0054] in, , , These represent the default costs for users' URLLC type services, eMBB1 type services, and eMBB2 type services, respectively.
[0055] The cost of a satellite system is defined as: ;
[0056] The optimization problem thus established is expressed as follows:
[0057]
[0058] in, For unloading decision set, For the allocation of spectrum resources, For the set of transmit power resource allocation, Allocate sets for computing resources; Indicates the total number of satellite service time slots. Indicates user The unloading decision variable, Indicates satellite Covers the user set.
[0059] Furthermore, step 3 specifically includes:
[0060] Step 3-1: Transform the optimization problem into a POMDP and define the key elements of the POMDP: state, observation, action, and reward function;
[0061] Status: In time slot The state is represented as:
[0062]
[0063] This includes service demand models and user geographic locations; , ;
[0064] Observation: Since communication between satellites is not considered, each satellite can only obtain its local information; therefore, for satellites... The observations are as follows:
[0065]
[0066] Action: Satellite The action is:
[0067]
[0068]
[0069] in, Indicates high-throughput satellite information about users The unloading decision variable, Indicates high-throughput satellite for users Allocated computing resources;
[0070] award: ;
[0071] Step 3-2: Update and model the actor and critic networks of the MATD3 algorithm;
[0072] For the critic network, updates are performed by minimizing the loss function. The loss functions for the two critic networks are as follows:
[0073] ,
[0074] in, To determine the target Q-value, the lower Q-value of the two target CRT networks is selected. Indicates the critic network parameters. Indicates the amount of sample data. ) represents the Q function, Indicates the sample state. Indicates the sample action, Indicates the specific updated sample;
[0075] The policy gradient of the actor network is:
[0076]
[0077] in, Represents the state components of the sampled data. Represents the action component of the sampled data. Indicates actor network parameters, Represents the Q function, This represents the first critic network parameter. Represents a deterministic policy network;
[0078] Since the actor network is not affected by the overestimation problem, the Q value of the first critic network is used;
[0079] The target network updates its parameters using a soft update method to stabilize the training process.
[0080]
[0081] =
[0082] =
[0083] in, This is a soft update coefficient; Indicates the target network parameters of the actor. This represents the first critic target network parameter. This represents the second critic target network parameter;
[0084] Step 3-3: Add a self-attention mechanism to the actor network of the MATD3 algorithm;
[0085] A self-attention mechanism is introduced before the hidden layer of the actor network in the TD3 algorithm to dynamically capture the contextual relationship between input features, improve the model's adaptability to dynamic environments and resource allocation efficiency. The self-attention layer receives local observation input, generates query, key and value vectors through linear transformation, calculates the similarity between Q and K, and applies scaling and Softmax operations to dynamically weight features.
[0086] A multi-head mechanism is used to execute multiple attention heads in parallel to capture the interaction patterns between features; the outputs are concatenated and mapped back to a unified feature space through linear projection, combined with residual connections and layer normalization.
[0087] In the output layer design, task offloading and subcarrier allocation decisions are continuous variables, and the Sigmoid activation function is used to limit the output to the [0,1] interval; the transmit power and computing resource allocation variables are normalized using the Softmax activation function to ensure that the sum of the allocation ratios does not exceed 1.
[0088] Steps 3-4: Design the MATD3 algorithm framework;
[0089] The improved MATD3 algorithm based on self-attention consists of multiple agents, each making decisions using the improved TD3 algorithm. It employs a centralized training and distributed execution framework. During the centralized training phase, the computing power of the ground station is used to aggregate observation data from all agents to simulate the global state. When updating the actor and critic networks, the actor network selects actions based on local observations, while the critic network evaluates the value of actions using information from other agents.
[0090] During the distributed execution phase, the trained actor network is deployed to various satellites, and the agent makes real-time decisions based solely on local observations.
[0091] Preferably, the devices accessing satellite services include smartphones, computers, and vehicles.
[0092] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device performs the above-mentioned task offloading and computing resource allocation method.
[0093] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described task offloading and computing resource allocation method.
[0094] A chip includes a processor for retrieving and running a computer program from a memory, causing a device equipped with the chip to perform the aforementioned task offloading and computing resource allocation method.
[0095] A computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the above-described task offloading and computing resource allocation method.
[0096] The beneficial effects of this invention are as follows:
[0097] This invention conducts a detailed study of the network model, uplink and downlink communication link model, computing resource model, satellite beam transmit power model, task completion delay model, and user diversity service demand model under a satellite internet constellation system, and obtains suitable modeling representations. Based on this, the optimization objective is constructed, and an improved MATD3 algorithm is proposed to efficiently solve it. The improved MATD3 algorithm proposed in this invention can provide support for the construction and optimization of 6G low-Earth orbit satellite internet systems. Attached Figure Description
[0098] Figure 1 This is a schematic diagram of a satellite internet system;
[0099] Figure 2 These are task completion rate curves of the proposed algorithms under different learning rates;
[0100] Figure 3 These are the system cost curves of the proposed algorithms under different learning rates;
[0101] Figure 4 This is a graph showing the task completion rate of the proposed algorithm and the benchmark algorithm;
[0102] Figure 5 This is a system cost curve comparing the proposed algorithm with the benchmark algorithm;
[0103] Figure 6 This is a graph showing the average task completion rate of different algorithms for different maximum number of people covered by different single beams;
[0104] Figure 7 This is a graph showing the average task completion rate of satellite S0 with different uplink bandwidths and different algorithms.
[0105] Figure 8 This is a graph showing the average task completion rate of satellite S0 with different downlink bandwidths and different algorithms.
[0106] Figure 9 This is a graph showing the average mission completion rate of satellite S0 with different transmission powers and different algorithms.
[0107] Figure 10 This is a graph showing the average task completion rate of satellite S0 with different computing resources and different algorithms. Detailed Implementation
[0108] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0109] This invention provides a method for task offloading and resource allocation in satellite internet constellations based on multi-agent reinforcement learning. First, the architecture of the satellite internet system is established, involving multiple satellites, multiple beams, and multiple ground users. Communication, computation, and task models for the uplink and downlink are modeled separately. Then, an optimization objective is defined: the satellite system supports multiple user services through appropriate task offloading and resource allocation while minimizing resource consumption. Next, the optimization problem is modeled as a partially observable Markov decision process (POMDP), and a multi-agent dual-delay deep deterministic policy gradient (MATD3) algorithm based on a self-attention mechanism is proposed to solve this problem. This invention can provide more user services with lower resource consumption and outperforms traditional multi-agent reinforcement learning methods in convergence and robustness tests, providing technical support for the development, design, and optimization of 6G satellite internet.
[0110] The technical solution adopted by this invention to solve its technical problem is as follows:
[0111] Step 1: The network architecture of the satellite internet system was constructed, involving multiple satellites, multiple beams and multiple ground users. The uplink and downlink communication links between satellites and users, satellite spectrum resources, power resources and computing resources were modeled respectively.
[0112] Step 1-1: The network architecture of the satellite internet system was constructed, using multiple low-Earth orbit satellites to form a satellite system to provide internet services to users in remote areas;
[0113] The satellite system consists of a high-throughput satellite and multiple base satellites. Ground users (GUs) can be devices such as smartphones, computers, and vehicles that access satellite services. The satellites provide internet service to multiple ground users in a multi-beam configuration. Each GU is covered by one beam from a base satellite, and all base satellites can cover the entire ground user base.
[0114] Each basic satellite is configured to have the same number of beams. The beam is from The index shows that the number of beams for high-throughput satellites is... , ( > ), the beam is index
[0115] Steps 1-2: Construct the uplink and downlink communication model between the satellite and the GU under the OFDMA multiple access scheme;
[0116] Establish a channel model between ground users and the base satellites. Assume that each base satellite uses the same communication frequency band and has the same uplink and downlink bandwidth. , The beams employ full-frequency reuse technology, with a bandwidth of [missing information]. , They can all be divided into The subcarrier bandwidth is [number], therefore the uplink subcarrier bandwidth is [number]. The downlink subcarrier bandwidth is The uplink and downlink subcarriers are respectively composed of and index.
[0117] Let the downlink subcarrier decision variables be... , , Indicates subcarrier Assigned to satellite Beam Users under coverage .
[0118] satellite Beam Users below The downlink transmission rate can be expressed as:
[0119]
[0120] User Downlink subcarrier Signal-to-interference-plus-noise ratio (SINR) on the surface. This indicates that the beam is assigned to a specific beam. Neutron carrier The transmission power, User Downlink channel gain.
[0121] Similarly, we can assume that the subcarrier decision variables for the uplink are... Assume the total transmission power of all ground users is a constant. ,satellite Beam Users below Set the transmit power of each allocated subcarrier to :
[0122]
[0123]
[0124] Then satellite Beam Users below The uplink transmission rate can be expressed as:
[0125] User uplink subcarrier The signal-to-interference-to-noise ratio (SIR) is... On behalf of users Uplink channel gain.
[0126] Next, establish connections between ground users and high-throughput satellites. The channel model. Uplink subcarrier bandwidth. Downlink subcarrier bandwidth The uplink and downlink subcarriers are respectively composed of and index.
[0127] Uplink and downlink subcarrier decision variables are defined and And there are and downlink subcarrier Transmit power is defined as , If subcarrier If not occupied by a user, then .user The downlink transmission rate can be expressed as:
[0128] user The uplink transmission rate can be expressed as:
[0129]
[0130] Steps 1-3: Establish the normalized transmit power of subcarriers in each satellite beam.
[0131] Satellite The upper limit of the transmission power is , , and have It is possible to obtain basic satellites and satellites. Normalized transmit power of subcarriers in the beam:
[0132] ,
[0133] ,
[0134] Steps 1-4: Establish computing resource and mission latency models for each satellite.
[0135] The satellite's computing resources are the maximum achievable computing frequency (CPU cycles per second), assuming... For satellite The maximum achievable computation frequency, then For users If there is a computational task, it can be represented as , This indicates the amount of data that needs to be transferred for the computation task. This indicates the amount of computing resources required to complete the task. This represents the maximum tolerable delay for the mission. For offloading the mission to the satellite... Beam users The total delay for completing the computation task can be expressed as:
[0136]
[0137] Step 2: Considering the diversity of user service needs in practical applications of satellite internet, three different demand models are used to model user service needs: uplink computing task upload requests, downlink high communication rate requests, and combined uplink and downlink high communication rate requests.
[0138] Step 2-1: Based on the typical 5G service types proposed by the International Telecommunication Union (ITU), the services provided by satellite internet are designed into three types;
[0139] First, the user Service requirements modeled as quintuples .
[0140] URLLC type services: , The latency of the user's computation task needs to be less than the task's tolerance latency, that is: ,
[0141] eMBB1 type service: Emphasizes high bandwidth and high throughput in satellite-to-ground downlink, with service requirements modeled as follows: , For offloading tasks to satellites Beam users The downlink transmission rate should be greater than the threshold. ,Right now:
[0142] eMBB2 type service: emphasizes simultaneously meeting the high throughput requirements of both uplink and downlink. Service requirement modeling is as follows: , For offloading tasks to satellites Beam users Uplink and downlink transmission rates should be greater than the threshold. ,Right now: ,
[0143] Step 2-2: Construct the expression for the resource cost and optimization problem of the satellite system;
[0144] In time slot t, the system resource cost is defined as follows:
[0145]
[0146] The service breach cost of the system is defined as:
[0147]
[0148] The cost of a satellite system is defined as:
[0149] The research objective of this invention is to minimize the resource consumption of the satellite system while meeting user needs. The optimization problem established thereby is expressed as follows:
[0150]
[0151] For unloading decision set, For the allocation of spectrum resources, For the set of transmit power resource allocation, Allocate a set of computing resources.
[0152] Step 3: Based on steps 1 and 2, the optimization objective of this invention is derived, namely, to enable the satellite system to support multiple user services while minimizing resource consumption through appropriate task offloading and resource allocation.
[0153] Step 3-1: Transform the optimization problem into a POMDP, and then define the key elements of the POMDP: state, observation, action, and reward function.
[0154] Status: In time slot The state can be represented as:
[0155]
[0156] This includes service demand models and user geographic locations.
[0157] Observation: Since communication between satellites is not considered, each satellite can only obtain its local information; therefore, for satellites... The observations are as follows:
[0158]
[0159] Action: Satellite The action is:
[0160]
[0161]
[0162] award:
[0163] Step 3-2: Update and model the actor and critic networks of the MATD3 algorithm.
[0164] For the critic network, updates are performed by minimizing the loss function. The loss functions for the two critic networks are as follows:
[0165] ,
[0166] To mitigate the overestimation problem, the lower Q value of the two target CRT networks is selected as the target Q value.
[0167] The policy gradient of the actor network is:
[0168] Since the actor network is not affected by the overestimation problem, the Q-value of the first critic network is used here.
[0169] The target network updates its parameters using a soft update method to stabilize the training process.
[0170]
[0171] =
[0172] =
[0173] This is the soft update coefficient.
[0174] Step 3-3: Add a self-attention mechanism to the actor network of the MATD3 algorithm.
[0175] In task offloading and resource allocation, the system needs to handle high-dimensional and heterogeneous local observation data. The complex dependencies of these characteristics and the spatiotemporal dynamics of the network pose challenges to the modeling capabilities of traditional fully connected networks.
[0176] To address this issue, a self-attention mechanism is introduced before the hidden layer of the actor network in the TD3 algorithm to dynamically capture the contextual relationships between input features, thereby improving the model's adaptability to dynamic environments and resource allocation efficiency. The self-attention layer receives local observation inputs, generates query, key, and value vectors through linear transformations, calculates the similarity between Q and K, applies scaling and Softmax operations, and dynamically weights features to prioritize critical tasks.
[0177] To enrich feature representations, a multi-head mechanism is employed to execute multiple attention heads in parallel, capturing diverse interaction patterns between features. The outputs are concatenated and mapped back to a unified feature space via linear projection. Residual connections and layer normalization are then combined to enhance model stability.
[0178] In the output layer design, task offloading and subcarrier allocation decisions are continuous variables, and the Sigmoid activation function is used to limit the output to the [0,1] interval. Transmit power and computational resource allocation variables are normalized using the Softmax activation function to ensure that the sum of the allocation ratios does not exceed 1.
[0179] Steps 3-4: Design the MATD3 algorithm framework.
[0180] The improved MATD3 algorithm based on self-attention consists of multiple agents, each making decisions using the improved TD3 algorithm. To address the resource constraints of satellite networks, the algorithm employs a centralized training and distributed execution framework. During the centralized training phase, the computing power of ground stations is used to aggregate observation data from all agents to simulate the global state. Taking satellite agent 𝑠0 as an example, all satellite states and actions are stored in 𝑠0's replay buffer.
[0181] During sampling, the d-th sample {𝑠0 (d), 𝑎0 (d), 𝑟0 (d), 𝑠0 (d + 1)} is constructed with 𝑠0 (d) and 𝑠0 (d + 1) as observations and actions of all agents. When updating the actor and critic networks, the actor network selects actions based on local observations, and the critic network evaluates the value of actions using information from other agents.
[0182] In the distributed execution phase, the trained actor network is deployed to various satellites, and the agents make decisions in real time based solely on local observations, reducing communication overhead and energy consumption, making it suitable for resource-constrained satellite networks.
[0183] Step 4: The optimization problem is modeled as a partially observable Markov decision process, and an improved multi-agent dual-delay deep deterministic policy gradient algorithm based on self-attention is proposed to solve the problem.
[0184] Example:
[0185] This invention proposes a satellite internet resource optimization allocation scheme based on the improved MATD3 algorithm with a self-attention mechanism. The technical solution adopted to solve its technical problems is as follows.
[0186] Step S1: The network architecture of the satellite internet system was constructed, involving multiple satellites, multiple beams and multiple ground users. The uplink and downlink communication links between satellites and users, satellite spectrum resources, power resources and computing resources were modeled respectively.
[0187] Step S101: The network architecture of the satellite internet system was constructed, using multiple low-orbit satellites to form a satellite system to provide internet services to users in remote areas.
[0188] like Figure 1 As shown, the satellite system consists of one high-throughput satellite and multiple basic satellites. Ground users can be devices such as smartphones, computers, and vehicles that access satellite services. It is assumed that the service time provided by the satellite system is divided into multiple time slots, i.e. ,make In time slot The set of GUs covered by a satellite system, and having | |= A quasi-static model is adopted, in each time slot Within this framework, the satellite and user topology and channel characteristics are assumed to remain unchanged. Let the basic satellite set be... Satellites provide internet service to multiple ground users in a multi-beam configuration. Each GU is covered by one beam from a base satellite, and all base satellites can cover the entire ground user base.
[0189] Each basic satellite is configured to have the same number of beams. The beam is from Index, in time slot By satellite Beam The user set covered is , , by satellite The user set covered is .
[0190] High-throughput satellites It shares the same orbital altitude and communication technology as the basic satellite, but is equipped with richer resources and independently covers all ground areas. (This is used...) express , This represents the set of all satellites. The number of beams is , ( > ), the beam is Indexes, analogous to basic satellites, are located in time slots. By satellite Beam The user set covered is , By satellite The user set covered is .
[0191] Assuming each GU covered by a satellite has one service request per time slot, we consider binary task offloading, meaning a user's service task can only be offloaded to one beam of a single satellite. This means the user... The unloading decision variable is represented as , , Indicates user The task is offloaded to the satellite .
[0192] Step S102: Construct a communication model for the uplink and downlink between the satellite and the GU. The multiple access scheme for the link is OFDMA.
[0193] First, a channel model is established between ground users and the base satellites. It is assumed that each base satellite uses the same communication frequency band and has the same uplink and downlink bandwidth. , The beams employ full-frequency reuse technology, with a bandwidth of [missing information]. , They can all be divided into The subcarrier bandwidth is [number], therefore the uplink subcarrier bandwidth is [number]. The downlink subcarrier bandwidth is The uplink and downlink subcarriers are respectively composed of and index.
[0194] Let the downlink subcarrier decision variables be... , , Indicates subcarrier Assigned to satellite Beam Users under coverage .
[0195] satellite Beam Users below The downlink transmission rate can be expressed as:
[0196]
[0197] User Downlink subcarrier The signal-to-interference-plus-noise ratio (SINR) on the surface can be expressed as:
[0198]
[0199] This indicates that the beam is assigned to a specific beam. Neutron carrier The transmit power, if the subcarrier If not occupied by a user, then . User Downlink channel gain.
[0200] Similarly, we can assume that the subcarrier decision variables for the uplink are... Assume the total transmission power of all ground users is a constant. ,satellite Beam Users below Set the transmit power of each allocated subcarrier to :
[0201]
[0202]
[0203] Then satellite Beam Users below The uplink transmission rate can be expressed as:
[0204] User uplink subcarrier The signal-to-interference-to-noise ratio (SIR) is... On behalf of users Uplink channel gain.
[0205] Next, establish connections between ground users and high-throughput satellites. The channel model. Therefore, the uplink subcarrier bandwidth. Downlink subcarrier bandwidth The uplink and downlink subcarriers are respectively composed of and index.
[0206] Uplink and downlink subcarrier decision variables are defined and And there are and downlink subcarrier Transmit power is defined as , If subcarrier If not occupied by a user, then .user The downlink transmission rate can be expressed as:
[0207] user The uplink transmission rate can be expressed as:
[0208]
[0209] Step S103: Establish the normalized transmit power of subcarriers in each satellite beam.
[0210] Satellite The upper limit of the transmission power is , , and have It is possible to obtain basic satellites and satellites. Normalized transmit power of subcarriers in the beam:
[0211] ,
[0212] ,
[0213] Step S104: Establish the computing resource model for each satellite.
[0214] Assuming the satellite's computing resources are limited to its maximum achievable computing frequency (CPU cycles per second), assuming... For satellite The maximum achievable computation frequency, then For users If there is a computational task, it can be represented as , This indicates the amount of data that needs to be transferred for the computation task. This indicates the amount of computing resources required to complete the task. This represents the maximum tolerable delay for the mission. For offloading the mission to the satellite... Beam users The total delay for completing the computation task can be expressed as:
[0215]
[0216] Indicates satellite With users physical distance, It represents the speed of light. , representing a satellite For users The ratio of allocated computing resources.
[0217] Step S2: Considering the diversity of user service needs in practical applications of satellite internet, three different demand models are used to model user service needs: uplink computing task upload requests, downlink high communication rate requests, and combined uplink and downlink high communication rate requests.
[0218] Step S201: Drawing on the typical 5G service types proposed by the International Telecommunication Union, the services provided by satellite internet are designed into three types: one is URLLC type service, and the other two are eMBB1 and eMBB2 type services.
[0219] Due to the heterogeneity of user tasks, for ease of analysis, users are... The task requirements are modeled as a quintuple. .
[0220] URLLC type services: Emphasizing extremely low latency and high reliability, user task requirements are modeled as follows , The latency of the user's computation task needs to be less than the task's tolerance latency, that is: ,
[0221] eMBB1 type service: Emphasizing high bandwidth and high throughput in satellite-to-ground downlink, the task requirements of eMBB1 type service users are modeled as follows: , For offloading tasks to satellites Beam users The downlink transmission rate should be greater than the threshold. ,Right now:
[0222] eMBB2 type service: emphasizes simultaneously meeting the high throughput requirements of both uplink and downlink. The task requirements of users of eMBB2 type service are modeled as follows: , For offloading tasks to satellites Beam users Uplink and downlink transmission rates should be greater than the threshold. ,Right now: ,
[0223] Step S202: Constructing the resource cost and optimization problem of the satellite system.
[0224] In the time slot The system resource cost is defined as follows:
[0225]
[0226]
[0227]
[0228]
[0229] In the time slot The service breach cost of the system is defined as:
[0230]
[0231]
[0232]
[0233]
[0234] In the time slot The cost of a satellite system is defined as:
[0235] The research objective of this invention is to minimize the resource consumption of the satellite system while meeting user needs. The optimization problem established thereby is expressed as follows:
[0236]
[0237] For unloading decision set, For the allocation of spectrum resources, For the set of transmit power resource allocation, Allocate a set of computing resources.
[0238] Step S3: Based on steps 1 and 2, the optimization objective in the research context of this invention is derived, namely, to enable the satellite system to support multiple user services while minimizing resource consumption through appropriate task offloading and resource allocation.
[0239] Step S301: Transform the optimization problem into a POMDP, and then define the key elements of the POMDP—state, observation, action, and reward function.
[0240] Status: In time slot The state can be represented as:
[0241]
[0242] This includes service demand models and user geographic locations.
[0243] Observation: Since communication between satellites is not considered, each satellite can only obtain its local information; therefore, for satellites... The observations are as follows:
[0244]
[0245] Action: Satellite The action is:
[0246]
[0247]
[0248] award:
[0249] Step S302: Update the actor network and critic network of the MATD3 algorithm.
[0250] satellite The actor network is defined as , parameters are The corresponding target network is , parameters are Similarly, we can assume the crtic network is... and The parameters are and .
[0251] The corresponding target network is and The parameters are and .
[0252] To enhance the agent's exploration capabilities and avoid local optima, noise is added to the actor network, and it is pruned as follows. =clip( + ,0,1) . The variance of the exploration noise is represented, and the clip function constrains the action output range to [0,1]. Each time agent j interacts with the environment, it receives a set of data as samples and stores them in the experience replay pool. During training, D samples are randomly selected from the experience replay pool. This is used as a mini-batch for network updates. For the critic network, updates are performed by minimizing the loss function. The loss functions for the two critic networks are as follows:
[0253] ,
[0254] To mitigate the overestimation problem, the lower Q value from the two target critic networks is selected as the target Q value.
[0255] The policy gradient of the actor network is:
[0256] Since the actor network is not affected by the overestimation problem, the Q-value of the first critic network is used here.
[0257] The target network updates its parameters using a soft update method to stabilize the training process.
[0258]
[0259] =
[0260] =
[0261] Step S303: Add a self-attention mechanism to the actor network of the MATD3 algorithm.
[0262] In task offloading and resource allocation scenarios, the system needs to process high-dimensional and heterogeneous local observation data, including the geographical locations and task requirements of multiple ground users. The complex dependencies between these features and the spatiotemporal dynamics of the network pose significant challenges to the modeling capabilities of traditional fully connected networks.
[0263] To address these limitations, a self-attention mechanism is introduced before the hidden layers of the actor network in TD3. This mechanism dynamically captures the contextual relationships between input features, thereby enhancing the model's adaptability to dynamic environments and improving resource allocation efficiency. The self-attention layer first receives local observation input and then generates query, key, and value vectors through linear transformations, capturing diverse representations of input features for attention computation. The scaled dot product attention mechanism computes the similarity between queries and keys, applies scaling and softmax operations, and dynamically weights features to prioritize critical tasks.
[0264] To further enrich the feature representation, a multi-head mechanism is employed, as shown in the parallel structure in the figure. This method executes multiple attention heads in parallel to capture diverse interaction patterns between features. The outputs of each head are concatenated and mapped back to a unified feature space via linear projection, ensuring consistency in information fusion. Residual connections and layer normalization work together to enhance model stability.
[0265] In the design of the final output layer, considering that task offloading decisions and subcarrier allocation decisions are relaxed to continuous variables, specifically... ∈ [0, 1]、 ∈ [0, 1] and For values ∈ [0, 1], the Sigmoid activation function is used to ensure that the output value is confined to the interval [0, 1]. For the transmit power resource allocation variable... and calculate resource allocation variables Given the need to meet total resource constraints, the Softmax activation function is used for normalization to ensure that the sum of the allocation ratios is less than or equal to 1.
[0266] During task unloading and resource allocation, the system needs to process high-dimensional and heterogeneous local observation data (such as the geographical location of ground users and task requirements). The complex dependencies of these characteristics and the spatiotemporal dynamics of the network pose challenges to the modeling capabilities of traditional fully connected networks.
[0267] To address this issue, a self-attention mechanism is introduced before the hidden layer of the TD3 actor network to dynamically capture the contextual relationships between input features, improving the model's adaptability to dynamic environments and resource allocation efficiency. The self-attention layer receives local observation input, generates query, key, and value vectors through linear transformation, calculates the similarity between Q and K, applies scaling and softmax operations, and dynamically weights features to prioritize critical tasks.
[0268] To enrich feature representations, a multi-head mechanism is employed to execute multiple attention heads in parallel, capturing diverse interaction patterns between features. The outputs are concatenated and mapped back to a unified feature space via linear projection. Residual connections and layer normalization are then combined to enhance model stability.
[0269] In the output layer design, task offloading and subcarrier allocation decisions are continuous variables, and the Sigmoid activation function is used to limit the output to the [0,1] interval. Transmit power and computational resource allocation variables are normalized using the Softmax activation function to ensure that the sum of the allocation ratios does not exceed 1.
[0270] Step S304: Design the MATD3 algorithm framework.
[0271] The improved MATD3 resource allocation algorithm based on self-attention consists of multiple agents, each making decisions using an improved TD3 algorithm. Considering the resource-constrained nature of satellite networks, the proposed algorithm employs a centralized training and distributed execution framework to adapt to this environment. Specifically, in the centralized training phase, the powerful computing capabilities of ground stations are used to aggregate observation data from all agents and simulate the global state. Taking satellite agent 𝑠0 as an example, the states and actions of all satellites are stored in agent 𝑠0's replay buffer.
[0272] During sampling, for the d-th sample {𝑠0 (d), 𝑎0 (d), 𝑟0 (d), 𝑠0 (d + 1)}, 𝑠0 (d) and 𝑠0 (d + 1) are constructed as 𝑠0 (d) = {𝑜 𝑗(d), 𝑎 𝑗(d) | 𝑗 ∈ S} and 𝑠0 (d + 1) = {𝑜 𝑗(d + 1), 𝑎 𝑗(d) | 𝑗 ∈ S}. When updating the parameters of the actor network and the critic network, the actor network selects actions based on local observations, while the critic network uses information from other agents to estimate the value of that action.
[0273] In the distributed execution phase, the trained actor network is deployed across various satellites, enabling each agent to make real-time decisions based solely on its local observations. This design significantly reduces inter-satellite communication overhead and energy consumption, making it ideal for resource-constrained environments in satellite networks.
[0274] Given the common objective of the optimization problem, each agent should cooperate to minimize the system cost. Assume each agent receives the same immediate reward, 𝑟(𝑡) = −𝑊(𝑡).
[0275] Step S4: The optimization problem is modeled as a partially observable Markov decision process, and an improved MATD3 algorithm based on self-attention is proposed to solve the problem.
[0276] Example:
[0277] The improved MATD3 algorithm based on the self-attention mechanism (SA-MATD3) was simulated and compared with the benchmark algorithm, verifying the convergence and effectiveness of the proposed algorithm. Under resource-abundant conditions, the proposed algorithm SA-MATD3 reduces resource consumption by 22.3% compared with the best-performing benchmark algorithm MATD3. Under resource-constrained conditions, the task completion rate of the SA-MATD3 algorithm is higher than that of the benchmark algorithm.
[0278] Without compromising universality, the satellite system is configured to operate at an altitude of 600 kilometers, comprising three base satellites and satellite 𝑠0. Each base satellite is equipped with 5 beams, each beam containing 5 subcarriers. Ground users are randomly distributed within each beam of the base satellites, with a maximum of 5 users covered by each beam. Satellite 𝑠0 is configured with 8 beams, each beam containing 10 subcarriers, and each beam can serve a maximum of 10 ground users.
[0279] First, the convergence of the SA-MATD3 algorithm is tested. Figure 2 and Figure 3The impact of different learning rates on the convergence of the proposed algorithm is demonstrated, with learning rates set to (1e-3, 1e-2), (1e-4, 1e-3), and (1e-5, 1e-4) for the actor network and the critic network, respectively. Figure 2 In the training, the system's Task Completion Rate (TCR) increases with the number of training epochs and converges to 1 across all learning rate settings. Specifically, for the learning rate pair (1e-3, 1e-2), the TCR begins to converge at approximately epoch 60; for (1e-4, 1e-3), convergence begins at approximately epoch 74; and for (1e-5, 1e-4), convergence occurs at approximately epoch 83. Figure 3 In the study, the system cost (average over 100 steps per epoch) of all algorithms gradually converges to its achievable lower bound as the number of training epochs increases, with slight fluctuations around this lower bound. The convergence epochs of system cost under the three learning rates are related to... Figure 2 The TCRs shown are highly consistent. However, the learning rate (1e-4, 1e-3) leads to a better convergence of system cost, followed by (1e-5, 1e-4). Therefore, in subsequent simulation experiments, the learning rate of the SA-MATD3 algorithm is uniformly set to (1e-4, 1e-3) to obtain better performance.
[0280] like Figure 4 As shown, the TCR curves of the proposed algorithm are very close to those of the benchmark algorithm, increasing from approximately 0.15 to nearly 1.0. This is because a higher task default cost parameter was set in the simulation, prioritizing task completion under conditions of sufficient resources. Figure 5 Further evidence shows that the SA-MATD3 algorithm has the lowest system cost, followed by MATD3, MADDPG, TD3, and DDPG. Combined with... Figure 4 This indicates that, under sufficient resource conditions, the TCR of both the proposed algorithm and the benchmark algorithm are close to 1.0, but the proposed algorithm has a lower system cost. This demonstrates the superior efficiency of the proposed algorithm in task offloading and resource allocation, enabling it to meet user service needs with lower resource consumption while reducing overall system cost. Multi-agent reinforcement learning algorithms, through collaborative decision-making among multiple agents, outperform single-agent deep reinforcement learning algorithms in resource allocation; therefore, TD3 and DDPG... Figure 5 The TD3 algorithm performed the worst among them. As an improved version of DDPG, the TD3 algorithm optimizes resource allocation, making MATD3 and TD3 outperform MADDPG and DDPG respectively in simulation results. Figure 5 The results show that the SA-MATD3 algorithm reduces system cost by 22.3% compared to the best-performing benchmark algorithm MATD3, and by 56.4% compared to the worst-performing DDPG.
[0281] Figure 6The performance of the proposed algorithms under resource-constrained conditions was evaluated by varying the maximum number of users covered by a single beam. When the maximum number of users was small, the average TCR of all algorithms approached 1 over 1000 rounds. However, as the maximum number of users increased, the resources allocated to each user decreased, potentially leading to task failure. Starting with a maximum number of users of 6, the average TCR of DDPG and TD3 decreased significantly, while SA-MATD3, MATD3, and MADDPG only showed a significant decrease with 7 users. Figure 6 The results show that the average TCR of the SA-MATD3 algorithm is consistently higher than all benchmark algorithms, demonstrating its superior performance under resource-constrained conditions. At a maximum user count of 10, the average TCR of the SA-MATD3 algorithm remains at approximately 61%, while the TCR of DDPG drops to approximately 31%.
[0282] Figures 7-10 The performance of satellite 𝑠0 under different resource levels is demonstrated. As resource availability increases, the average TCR of the SA-MATD3 algorithm gradually rises, eventually approaching 1. Under the four resource types and different levels of satellite 𝑠0, the proposed SA-MATD3 algorithm consistently outperforms all benchmark algorithms, exhibiting strong robustness. The average TCR curves of MATD3 and MADDPG are very similar, with MATD3 generally performing better. Figure 6 Consistently, TD3 and DDPG outperform multi-agent reinforcement learning algorithms. Furthermore, Figure 10 As shown, the average TCR change of computing resources is more gradual than that of other resources because computing resources only affect users who require URLLC type services, while uplink bandwidth, downlink bandwidth, and transmit power resources affect users who require both eMBB1 and eMBB2 type services.
Claims
1. A method for task offloading and resource allocation in a satellite internet constellation based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct the network architecture of the satellite internet system, including multiple satellites, multiple beams and multiple ground users. Model the uplink and downlink communication links between satellites and users, satellite spectrum resources, power resources and computing resources respectively. Step 2: Considering the diversity of user service needs in practical applications of satellite internet, three different types of user service needs are modeled: uplink computing task upload requests, downlink high communication rate requests, and combined uplink and downlink high communication rate requests. Step 3: Derive the optimization objective from steps 1 and 2, namely, to enable the satellite system to support multiple user services while reducing resource consumption through task offloading and resource allocation; Step 4: Model the optimization problem as a partially observable Markov decision process, and use an improved multi-agent dual-delay deep deterministic policy gradient algorithm based on self-attention to solve the optimization problem; The improved multi-agent dual-delay deep deterministic policy gradient algorithm based on self-attention introduces an 8-head, 32-dimensional self-attention network before the hidden layer of the actor network in the multi-agent dual-delay deep deterministic policy gradient algorithm, thereby improving the model's adaptability to dynamic environments.
2. The method for task offloading and resource allocation of a satellite internet constellation based on multi-agent reinforcement learning according to claim 1, characterized in that, Step 1 specifically includes: Step 1-1: Construct the network architecture of the satellite internet system, using multiple low-Earth orbit satellites to form a satellite system to provide internet services to users in remote areas; The satellite system consists of a high-throughput satellite and multiple base satellites; the satellites provide internet service to multiple ground user units (GUs) in the form of multiple beams, and each GU is covered by one beam of a base satellite, so that all base satellites can cover the entire ground user base. The basic satellite set is Each basic satellite is configured to have the same number of beams. The beam is from Index, use Indicates high-throughput satellite , This represents the entire set of satellites; the number of beams for high-throughput satellites is... , > The beam is from Index; by satellite The user set covered is The satellite system's service time is divided into multiple time slots, namely... ; Assuming each GU covered by a satellite has one service request per time slot, and the service request varies across time slots, consider binary task offloading, meaning a user's service task can only be offloaded to one satellite—either the base satellite or a high-throughput satellite covering it—within a single beam. The unloading decision variable is represented as , , Indicates user The task is offloaded to the satellite ; Steps 1-2: Construct the uplink and downlink communication model between the satellite and the GU under the OFDMA multiple access scheme; Establish a channel model between ground users and base satellites; assume that each base satellite uses the same communication frequency band and has the same uplink and downlink bandwidth. , The beams employ full-frequency reuse technology, with a bandwidth of [missing information]. , All can be divided into The subcarrier bandwidth is [number], therefore the uplink subcarrier bandwidth is [number]. The downlink subcarrier bandwidth is Both uplink and downlink subcarriers are generated by index; Let the downlink subcarrier decision variables be... , , Indicates subcarrier Assigned to satellite Beam Users under coverage ; satellite Beam Users below The downlink transmission rate is expressed as: ; in, User Downlink subcarrier Signal-to-interference-plus-noise ratio (SINR) on the surface; This represents the subcarrier bandwidth of the downlink; Similarly, let the subcarrier decision variables of the uplink be... Assume the total transmission power of all ground users is a constant. ,satellite Beam Users below Set the transmit power of each allocated subcarrier to : ; Then satellite Beam Users below The uplink transmission rate is expressed as: ; in, User uplink subcarrier The signal-to-interference-to-noise ratio (SIR) is... This represents the uplink subcarrier bandwidth; Next, establish connections between ground users and high-throughput satellites. Channel model; uplink subcarrier bandwidth Downlink subcarrier bandwidth The uplink and downlink subcarriers are respectively composed of and index; , These represent the total bandwidth of the uplink and downlink, respectively. , These represent the number of uplink and downlink subcarriers, respectively. Uplink and downlink subcarrier decision variables are defined as follows: and And there are and downlink subcarrier Transmit power is defined as , If subcarrier If not occupied by a user, then ,user The downlink transmission rate is expressed as: ; user The uplink transmission rate can be expressed as: ; in, These represent the uplink and downlink subcarrier bandwidths, respectively. , They represent the uplink and downlink subcarriers, respectively. Signal-to-interference-plus-noise ratio; Steps 1-3: Establish the normalized transmit power of subcarriers in each satellite beam; Satellite The upper limit of the transmission power is , , and have To obtain basic satellites and satellites Normalized transmit power of subcarriers in the beam: , ; , ; in, Indicates the basic satellite subcarrier The transmission power, Indicates high-throughput satellite subcarriers The transmission power; These represent the basic satellite and high-throughput satellite subcarriers, respectively. Normalized transmit power; Steps 1-4: Establish computational resource and mission latency models for each satellite; The satellite's computing resources are its maximum achievable computing frequency, i.e., the number of CPU cycles per second. Assuming... For satellite The maximum achievable computation frequency, then For users If there is a computational task, it is represented as , This indicates the amount of data that needs to be transferred for the computation task. This indicates the amount of computing resources required to complete the task. The maximum tolerable delay for the mission; for offloading the mission to the satellite Beam users The total delay for completing the computation task is expressed as: ; in, This indicates the proportion of total computing power resources allocated. Indicates the physical distance between the satellite and the user. It is the speed of light.
3. The method for task offloading and resource allocation of a satellite internet constellation based on multi-agent reinforcement learning according to claim 2, characterized in that, Step 2 specifically includes: Step 2-1: Based on the typical 5G service types proposed by the International Telecommunication Union (ITU), the services provided by satellite internet are designed into three types: URLLC type service, eMBB1 type service, and eMBB2 type service. First, the user Service requirements modeled as quintuples , These represent the user's uplink and downlink communication rate requirements thresholds, respectively; the user set required for URLLC type services is set to... The user set that requires eMBB1 type services and eMBB2 type services is set as follows: ; URLLC type services: , The latency of the user's computation task needs to be less than the task's tolerance latency, that is: , , Indicates the tolerable latency of the task; eMBB1 type service: high bandwidth and high throughput of satellite-to-ground downlink, service requirements modeled as follows , For offloading tasks to satellites Beam users The downlink transmission rate should be greater than the threshold. ,Right now: ; eMBB2 type service: simultaneously meets the high throughput requirements of both uplink and downlink; service requirements are modeled as follows: , For offloading tasks to satellites Beam users Uplink and downlink transmission rates should be greater than the threshold. ,Right now: , ; Step 2-2: Construct the expression for the resource cost and optimization problem of the satellite system; In time slot t, the system resource cost is defined as follows: ; in, , These represent the spectrum resource cost, transmit power resource cost, and computing resource cost of the satellite system, respectively. The service breach cost of the system is defined as: ; in, , , These represent the default costs for users' URLLC type services, eMBB1 type services, and eMBB2 type services, respectively. The cost of a satellite system is defined as: ; The optimization problem thus established is expressed as follows: ; in, For unloading decision set, For the allocation of spectrum resources, For the set of transmit power resource allocation, Allocate sets for computing resources; Indicates the total number of satellite service time slots. Indicates user The unloading decision variable, Indicates satellite Covers the user set.
4. The method for task offloading and resource allocation of a satellite internet constellation based on multi-agent reinforcement learning according to claim 3, characterized in that, Step 3 specifically involves: Step 3-1: Transform the optimization problem into a POMDP and define the key elements of the POMDP: state, observation, action, and reward function; Status: In time slot The state is represented as: ; This includes service demand models and user geographic locations; , ; Observation: Since communication between satellites is not considered, each satellite can only obtain its local information; therefore, for satellites... The observations are as follows: ; Action: Satellite The action is: ; ; in, Indicates high-throughput satellite information about users The unloading decision variable, Indicates high-throughput satellite for users Allocated computing resources; award: ; Step 3-2: Update and model the actor and critic networks of the MATD3 algorithm; For the critic network, updates are performed by minimizing the loss function. The loss functions for the two critic networks are as follows: , 。 5. Among them, To determine the target Q-value, the lower Q-value of the two target CRT networks is selected. Indicates the critic network parameters. Indicates the amount of sample data. ) represents the Q function, Indicates the sample state. Indicates the sample action, Indicates the specific updated sample; The policy gradient of the actor network is: ; in, Represents the state components of the sampled data. Represents the action component of the sampled data. Indicates actor network parameters, Represents the Q function, This represents the first critic network parameter. Represents a deterministic policy network; Since the actor network is not affected by the overestimation problem, the Q value of the first critic network is used; The target network updates its parameters using a soft update method to stabilize the training process. ; = ; = ; in, This is a soft update coefficient; Indicates the target network parameters of the actor. This represents the first critic target network parameter. This represents the second critic target network parameter; Step 3-3: Add a self-attention mechanism to the actor network of the MATD3 algorithm; A self-attention mechanism is introduced before the hidden layer of the actor network in the TD3 algorithm to dynamically capture the contextual relationship between input features, improve the model's adaptability to dynamic environments and resource allocation efficiency. The self-attention layer receives local observation input, generates query, key and value vectors through linear transformation, calculates the similarity between Q and K, and applies scaling and Softmax operations to dynamically weight features. A multi-head mechanism is used to execute multiple attention heads in parallel to capture the interaction patterns between features; the outputs are concatenated and mapped back to a unified feature space through linear projection, combined with residual connections and layer normalization. In the output layer design, task offloading and subcarrier allocation decisions are continuous variables, and the Sigmoid activation function is used to limit the output to the [0,1] interval; the transmit power and computing resource allocation variables are normalized using the Softmax activation function to ensure that the sum of the allocation ratios does not exceed 1. Steps 3-4: Design the MATD3 algorithm framework; The improved MATD3 algorithm based on self-attention consists of multiple agents, each making decisions using the improved TD3 algorithm. It employs a centralized training and distributed execution framework. During the centralized training phase, the computing power of the ground station is used to aggregate observation data from all agents to simulate the global state. When updating the actor and critic networks, the actor network selects actions based on local observations, while the critic network evaluates the value of actions using information from other agents. During the distributed execution phase, the trained actor network is deployed to various satellites, and the agent makes real-time decisions based solely on local observations.
6. The method for task offloading and resource allocation of a satellite internet constellation based on multi-agent reinforcement learning according to claim 4, characterized in that, The devices that access satellite services include smartphones, computers, and vehicles.
7. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
9. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, The computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the method as described in any one of claims 1 to 5.