Task offloading and resource allocation method and system for ocean-surface stereo satellite network
By optimizing resource scheduling through a multi-layer offloading architecture and the MAPPO algorithm, the task offloading and resource allocation problems of satellite networks in marine environments are solved, efficient, low-latency and low-energy task processing is achieved, and system performance is optimized.
Patent Information
- Application Number
- CN202510068102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Task offloading and resource allocation of satellite networks in marine environments face problems of high latency, low efficiency, high energy consumption and high cost. Especially under dynamic and complex ocean conditions, traditional multi-agent reinforcement learning algorithms such as MADQN and MAA2C perform poorly and have difficulty in effectively coordinating multi-layer networks.
A multi-layer offloading architecture model is combined with the multi-agent proximal policy optimization (MAPPO) algorithm to achieve collaborative decision-making among multiple agents by optimizing resource scheduling and task offloading strategies. A heterogeneous agent system consisting of LEO satellites, HAPS, ground edge servers, and sea surface terminal users is used to build a communication, latency, energy consumption, and cost model to optimize task offloading and resource allocation.
Achieve efficient task processing in complex marine environments, reduce system latency and energy consumption, improve resource utilization, optimize costs, and enhance system performance.
Smart Images

Figure CN119967486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite application technology in sea surface scenarios, and in particular to a technical solution for optimizing task offloading efficiency and resource scheduling based on a multi-layer offloading architecture and a multi-agent deep reinforcement learning algorithm (MARL). The solution is suitable for task offloading, computing resource allocation, delay optimization, energy consumption reduction, and cost control in sea surface stereo satellite networks. Background Art
[0002] With the growing demand for marine communications, especially in offshore and oceanic applications, satellite networks have become a key solution to overcome the inherent limitations of traditional communication methods. However, maritime users in remote areas often face significant challenges, including high latency, limited bandwidth, and constraints on energy and computing power, which hinder the effective execution of complex services. These challenges highlight the dual problems of communication inefficiency and computational burden in marine networks, necessitating innovative strategies for reliable service delivery under harsh conditions. Although task offloading techniques have shown great promise in terrestrial networks, their application in marine environments presents unique difficulties. The unique characteristics of marine networks, including the variability of communication links, the scarcity of computing resources, and the dynamic nature of marine weather, pose significant obstacles to the effective implementation of task offloading and warrant further research and development.
[0003] A multi-layered satellite-drone-user architecture has been proposed as a promising solution to enhance task offloading and resource allocation in ocean networks. This architecture offers key advantages, including expanded communication coverage, reduced latency through drone relay, and support for distributed task processing. However, significant challenges hinder its practical deployment. The reliance on drones presents significant weaknesses, such as limited flight time due to energy constraints, susceptibility to adverse weather conditions, and difficulty maintaining stable connectivity with satellites and users. Furthermore, the architecture requires complex coordination across multiple network layers, which often leads to inefficiencies and increased latency in dynamic ocean environments. Addressing continuous availability, meeting quality of service requirements, and adapting to environmental changes further exacerbate deployment difficulties. Among these challenges, the operational limitations of drones, a key component of the architecture, require thorough investigation.
[0004] Specifically, environmental factors, including distance, inclement weather, and dynamic ocean conditions, often degrade the communication link between satellites and drones, leading to interference, increased latency, and reduced reliability. In high-traffic scenarios, satellite bandwidth limitations exacerbate inefficient resource allocation and degrade service quality, especially for data-intensive missions such as video streaming or emergency transmissions. Drones, constrained by limited battery capacity, payload limitations, and sensitivity to environmental conditions, often experience connectivity interruptions during extended operations due to energy depletion, insufficient bandwidth, or mission overload. Such interruptions pose a significant challenge to maintaining stable and reliable communications, especially in dynamic ocean environments. Given the limitations of drones in meeting these demands, alternative solutions, including energy-efficient designs, hybrid communication systems, and advanced satellite technologies, deserve further exploration.
[0005] To address the inherent limitations of drones, existing research has proposed the Sky-Ground-Sea Integrated Network (SAGSIN) as a robust and scalable solution for maritime communications. This multi-layer architecture leverages the collaborative operation of low-Earth orbit (LEO) satellites and high-altitude platform systems (HAPS) to improve communication relay, computation offloading, and resource optimization. LEO satellites are equipped with onboard computing servers to enhance computing power, while HAPS can provide longer operation time and wider coverage compared to drones, making them particularly suitable for wide-area maritime surface communications. The integration of HAPS and LEO satellites in SAGSIN not only significantly expands network coverage in remote ocean areas, but also alleviates the latency and reliability challenges associated with drone-based architectures. In addition, the architecture facilitates efficient task offloading by optimizing resource scheduling, making it a promising solution to address the complex requirements of maritime communication networks.
[0006] However, despite these advantages, implementing SAGSIN, particularly its LEO satellite-HAPS-user architecture, faces significant challenges due to the complexity of managing complex communication links and allocating heterogeneous server resources across multiple layers. As a key component of SAGSIN, the architecture must adapt to a variety of offloaded task sources, including ships, buoys, and offshore platforms, which vary significantly in task types and computational requirements. This variability complicates resource allocation and necessitates efficient conversion of tasks into unified binary data through feature extraction for consistent processing across the network. Furthermore, the heterogeneous computing capabilities of nodes at different network layers require advanced scheduling algorithms to optimize resource utilization and maintain system performance. These challenges highlight the need for further research to improve the operational efficiency of the SAGSIN framework in real-world ocean scenarios.
[0007] Building on the challenges of task offloading and resource optimization in the SAGSIN architecture, deep reinforcement learning (DRL) has shown great potential for managing dynamic and uncertain conditions in multi-layer networks. However, traditional multi-agent reinforcement learning (MARL) algorithms, such as the Multi-Agent Deep Q-Network (MADQN) and the Multi-Agent Advantage Actor-Critic (MAA2C), exhibit limitations in scalability, adaptability, and coordination efficiency in highly dynamic ocean environments. Specifically, MADQN suffers from instability and inefficiency when handling large state-action spaces, while MAA2C suffers from high computational complexity and suboptimal coordination under rapidly changing conditions. In contrast, Multi-Agent Proximal Policy Optimization (MAPPO) demonstrates superior sample efficiency, stability, and improved inter-agent coordination through centralized training and decentralized execution, making it particularly well-suited for dynamic and complex environments. MAPPO enhances training stability in heterogeneous scenarios using clipping updates and implements a centralized training, decentralized execution (CTDE) framework to balance global information sharing with real-time agent autonomy. Furthermore, it effectively promotes cooperation and competition among agents, enabling efficient task allocation and resource optimization in complex scenarios. These properties make MAPPO a powerful approach to addressing computational and communication challenges in marine applications.
[0008] To this end, the present invention proposes a multi-layer offloading architecture model for ocean surface stereo satellite networks. Combined with the MAPPO method, by optimizing resource scheduling and task offloading strategies, it effectively improves the task offloading efficiency in ocean surface stereo satellite networks, reduces system latency, saves energy and optimizes costs. Summary of the Invention
[0009] The main purpose of this invention is to provide a task offloading and resource allocation method and system for a 3D ocean-surface satellite network. By designing a multi-layer offloading architecture model and combining it with the MAPPO algorithm, and through the collaboration of multiple intelligent agents, this method optimizes task offloading decisions and resource allocation strategies in the 3D ocean-surface satellite network in real time, effectively improving system performance and minimizing latency, energy consumption, and cost.
[0010] Technical Solution: A task offloading and resource allocation system for a 3D ocean-surface satellite network. This system is based on a multi-layer offloading model for the 3D ocean-surface satellite network. The multi-layer offloading model includes a space-based network, an air-based network, and a ground-sea integrated network.
[0011] The space-based network includes LEO satellites equipped with onboard computing services;
[0012] The air-based network includes a high-altitude platform system (HAPS) equipped with airborne computing services, and the ground-sea integrated network includes ground edge servers and sea-surface terminal users;
[0013] The servers in each layer of the network are regarded as an agent, and the agents in each layer have different capabilities, resources and goals, forming a heterogeneous agent system;
[0014] The task offloading and resource allocation method of this system is:
[0015] The tasks generated by the sea surface terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}. The set of agents in the system is N, n l Represents the intelligent agents in the lth layer. The intelligent agents in each layer can monitor the communication, load, and resource distribution within the layer, calculate the transmission rate, transmission and computing delay, energy consumption and cost based on the task size assigned to each layer, and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observation.
[0016] Furthermore, the system adopts a multi-layer offloading method for a three-dimensional ocean surface satellite network, which includes:
[0017] S1. Build a multi-layered heterogeneous architecture with at least four main layers, including surface end users, ground edge servers, HAPS, and LEO satellites. The agents in each layer can provide appropriate resource support and offloading strategies based on the task, and the offloading models at different layers can cooperate with each other.
[0018] S2. Various tasks generated by end users are converted into binary data through feature extraction to achieve consistent processing across the entire network, including the construction of communication models, latency models, energy consumption models, and cost models;
[0019] Communication model: The wireless transmission rate is expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the transmission power and channel gain of the sea surface terminal user, respectively, and σ(t) 2 is the power of additive white Gaussian noise;
[0020] Delay model: The computational delay of each layer is expressed as The transmission delay models are Among them, u l The CPU cycles required for each layer of agents to calculate 1 bit of task, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, f l is the maximum computing resource of each layer of agents, β lis the percentage of computing resources allocated to each layer of agents; the total delay T(t) is the maximum value of each layer, that is,
[0021] Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is Among them, the computing power of each layer of intelligent agent is Where μ≥0 is the effective switching capacitance, the total energy consumption generated by the computing task is expressed as: The transmission energy consumption model of task offloading to each layer is: Where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t; the total energy consumption generated by the transmission task is expressed as: Therefore, the total energy consumption of the system is E(t) = E comp (t)+E tran (t);
[0022] Cost model: The model of the computational cost generated by the computational tasks in each layer is: in is the cost of processing each unit of data by each layer of agents in each time slot t; the transmission cost model of task offloading to each layer is: The cost of transmitting each unit of data in the link for each layer of agents in each time slot t; the total cost of the system is C(t) = C comp (t)+C tran (t);
[0023] S3. Each server in each layer is considered as an intelligent agent. The state space and action space of each intelligent agent are defined, including information such as task allocation, computing resources, and communication bandwidth. The action of each intelligent agent is the proportion of computing tasks in each time slot, and it is defined as a partial Markov decision process.
[0024] S4. After performing task offloading and resource allocation actions, each agent updates its state based on environmental feedback. Multiple rounds of training based on the MAPPO algorithm are used to update the agent's strategy, so that each agent can continuously optimize the task offloading and resource allocation strategies, gradually improving the overall performance of the system. Finally, the task offloading decision and corresponding performance indicators are output. The performance indicators include stable cumulative rewards, optimized delay, energy consumption and cost.
[0025] Furthermore, in the multi-layer heterogeneous architecture described in step S1, the functional characteristics of each layer and their collaborative relationships are defined as follows:
[0026] Sea surface terminal user layer: This layer includes terminal devices deployed on the sea surface, which are used for basic processing including data collection and user interaction;
[0027] Ground edge server layer: This layer includes edge servers deployed on the ground, which are used to handle tasks that are more complex or require higher computing power than the surface terminal user layer, and also includes processing computing tasks offloaded from the surface terminal user layer;
[0028] HPAS layer: This layer is deployed on high-altitude platforms on the sea surface, including balloons, airships, or high-altitude base stations. It is used to handle tasks offloaded from the sea surface terminal user layer and / or the ground edge server layer, and handles computing tasks with computing power and latency between the sea surface terminal user layer and the ground edge server layer.
[0029] LEO satellite layer: This layer consists of a group of satellites in low Earth orbit to provide global communication and data transmission services.
[0030] In the above scheme, the partial Markov decision process described in step S3 is as follows:
[0031] State space: In each time slot t, the environment state is expressed as: S(t) = {α l M(t),B(t),f l ,β l ,u l |l∈L}; where the amount of task data per layer is α l M(t), the communication wireless bandwidth is B(t), and the maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The number of CPU cycles required to compute a 1-bit task for each layer of the agent;
[0032] Action space: In each time slot t, the action of the agent is expressed as: A(t) = {α l ,β l |l∈L}; Make decisions based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents;
[0033] Reward function: In each time slot t, the future rewards are weighted summed using the discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t = -(η1T(t)+η2E(t)+η3C(t)), where η1, η2, and η3 are the weighting factors of the total system delay, energy consumption, and cost, respectively; the total reward R(t) is the cumulative reward value obtained during the entire decision-making process;
[0034] The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem is defined as:
[0035]
[0036] Among them, the constraint C1 means to ensure the computing delay of the task at each layer and the transmission time of each layer It should be within the maximum delay range that the system can tolerate; Constraint C2 ensures that the sum of the task data allocated to each layer should be equal to the total task data volume; Constraint C3 represents the cost of processing each unit of data by the server at each layer, from low to high: marine terminal users, ground edge servers, HAPS with airborne computing servers, and LEO satellites with airborne computing servers; Constraint C4 indicates that both the task data volume and the processing cycle are positive.
[0037] The strategy for updating the agent based on multiple rounds of training of the MAPPO algorithm in step S4 is as follows:
[0038] S41, the algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively;
[0039] S42, initializing a resource-constrained environment and a memory buffer U for storing experience data generated during the training process;
[0040] The main loop of S43 and MAPPO algorithm starts from each training cycle E, firstly initializes the environment state and agent action, and then iterates according to the maximum number of steps; each step operation is carried out in a multi-level architecture, by traversing each layer, the agent n in each layer is initialized. l Perform task assignment and action selection;
[0041] S44. In the specific operation of each layer, each agent first observes its current local state (i.e., partial observation), and then θ , select an optimal action a;
[0042] S45, MAPPO algorithm through value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t ,
[0043] The MAPPO algorithm involves accumulating the total reward R of all steps using a discount factor γ to optimize long-term returns;
[0044] S46, the environment updates the state S according to the agent's actions and provides new input for the next iteration;
[0045] S47. After each cycle, the MAPPO algorithm stores the accumulated experience data, including the stable cumulative reward, optimized delay, energy consumption, and cost, into the memory buffer U.
[0046] S48. After completing all training cycles, the MAPPO algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration.
[0047] Beneficial effects: Compared with the prior art, the substantial features and significant improvements of the solution of the present invention include:
[0048] (1) By combining a multi-layer offloading architecture model with a multi-agent deep reinforcement learning optimization algorithm, the present invention can achieve efficient coordination of task offloading and resource allocation in complex marine environments, improve the speed and accuracy of task processing, and reduce system latency.
[0049] (2) The task offloading and resource scheduling method proposed in the present invention can realize adaptive scheduling decisions according to the dynamic changes of the sea surface stereo satellite network and different task requirements, and effectively improve resource utilization.
[0050] (3) The present invention reduces energy consumption in the computing and communication processes, reduces unnecessary resource consumption, effectively extends the service life of satellites and equipment, and reduces the operating costs of the system.
[0051] (4) Through the optimization of the MAPPO algorithm, the present invention can achieve efficient collaboration and decision-making in a multi-agent environment, has strong robustness, and can adapt to the complex changes in the marine environment.
[0052] Therefore, this paper establishes a multi-layer offloading system model for LEO satellites, HAPS, and users. With the weighted minimization of system processing latency, energy consumption, and cost as the optimization problem, it proposes a joint optimization algorithm for task offloading and resource allocation based on the MAPPO method. Experimental verification shows that, compared with existing methods, this model and method effectively solves the problem of coordinated task offloading for surface users, reducing latency, energy consumption, and cost, and optimizing overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 Schematic diagram of the three-dimensional ocean surface satellite network according to the present invention;
[0054] Figure 2 Schematic diagram of the process of a multi-layer offloading decision algorithm in a specific embodiment of the present invention;
[0055] Figure 3This is the structural diagram of the MAPPO unloading decision algorithm of the present invention. DETAILED DESCRIPTION
[0056] The following describes embodiments of the present invention in detail, and examples of the embodiments are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain the present invention, and should not be understood as limiting the present invention.
[0057] Within a single ocean area, the present invention constructs a multi-layered offloading system and its model for surface end users. This system is comprised of a space-based network, an air-based network, and an integrated ground-sea network. The space-based network includes LEO satellites equipped with airborne computing services, the air-based network includes a high-altitude platform system (HAPS) equipped with airborne computing services, and the integrated ground-sea network includes ground edge servers and surface end users. Each server in each layer of the network is considered an intelligent agent, each with different capabilities, resources, and goals, forming a heterogeneous intelligent agent system.
[0058] The tasks generated by the sea terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks have a certain amount of data and are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}, and the set of agents in the system is N, n l Represents the agents in layer l. The agents in each layer can monitor the communication, load, resource distribution, etc. within the layer and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observations.
[0059] In defining a joint optimization model that balances system latency, energy consumption, and cost, the size of the tasks assigned to each layer is used to calculate the transmission rate, transmission and computational latency, energy consumption, and cost. Agents at each layer are able to provide appropriate resource support and offloading strategies based on the task. Offloading models at different levels can collaborate to ensure efficient task offloading, ultimately minimizing the system's overall latency, energy consumption, and cost.
[0060] In conjunction with the above-mentioned communication system, a multi-layer unloading method to a three-dimensional ocean surface satellite network is provided, comprising:
[0061] Multi-agent collaboration enables intelligent decision-making and global optimization for task offloading in a multi-layer architecture. The MAPPO algorithm leverages the state-awareness of the multi-layer network structure to dynamically model the computing resources, communication bandwidth, and latency of the LEO satellites, HAPS, ground edge servers, and end-user devices in the architecture model, constructing a global state and local observations. Agents generate task offloading targets and resource allocation plans based on the current environmental state. By sharing global information, they collaboratively learn to reduce task latency, energy consumption, and cost. The MAPPO algorithm, an extension of the PPO algorithm, ensures training stability through policy updates using pruning, guides policy optimization using an advantage function, and designs a reward function that comprehensively considers task completion time, resource utilization, and energy consumption, achieving efficient offloading and resource scheduling in a dynamically changing environment. After multiple rounds of interaction and optimization, the agents learn the optimal policy, enabling the system to adapt to complex dynamic network conditions and achieve global optimization for task offloading in a multi-layer architecture. The introduction of the MAPPO algorithm significantly improves task processing efficiency and system stability, significantly optimizes resource utilization in multi-layer satellite networks, and provides an effective solution to address the high latency, uneven resource distribution, and energy consumption issues in maritime communications.
[0062] Specifically, see Figure 1 , Figure 1 It is a communication system established by the present invention, and also an application environment for multi-layer unloading of sea-surface stereo satellite networks. Among them, the space-based network includes several LEO satellites equipped with airborne computing services, the air-based network includes several HAPS equipped with airborne computing services, and the land-sea integrated network includes several ground edge servers and sea-surface terminal users. Based on the multi-layer unloading model, the computing power of sea-surface terminal users is the weakest, the computing resources are the least, but the computing cost is the lowest. LEO satellites and ground edge servers have strong computing power and more computing resources, but the server cost is higher. HAPS is equipped with certain computing and storage resources, but it is far less than that of ground edge servers, and the cost is also lower than that of ground edge servers. When a sea-surface terminal user generates a task, the terminal user passes the task-related information to the intelligent agent. The intelligent agent in the sea-surface network will decide whether the task can be processed directly or needs to be unloaded to other devices (LEO satellites, HAPS or ground edge servers) based on the current network conditions, load conditions and resource availability.
[0063] See also Figure 2 , Figure 2 The following is a flowchart of the multi-layer offloading decision algorithm based on MAPPO, which includes the following steps:
[0064] S1. The system builds a multi-layered heterogeneous architecture covering four main layers: surface end users, ground edge servers, HAPS, and LEO satellites. The core of this modeling process is to accurately define the functional characteristics of each layer and their synergistic relationships to adapt to different mission requirements and network environments. Specifically:
[0065] Layer 1 (surface terminal user layer): This layer primarily consists of terminal devices located on the sea surface, such as smart devices and ships. Layer 1 is responsible for handling simple tasks that can be directly calculated by terminal devices, such as data collection, basic processing, and user interaction.
[0066] Layer 2 (Ground Edge Server Layer): This layer is located on the ground and typically includes edge servers with high-performance computing resources. Compared to Layer 1, Layer 2 has more powerful computing resources and can handle tasks that require high computing power and low latency. It is suitable for offloading computing tasks from Layer 1, especially in scenarios with high timeliness requirements, such as real-time data processing and complex analysis.
[0067] Layer 3 (HPAS): This layer is deployed on high-altitude platforms above the ocean, such as balloons, airships, or high-altitude base stations. Layer 3 is suitable for tasks with computing resource and timeliness requirements between Layers 1 and 2. However, it provides highly flexible and low-cost task offloading, making it particularly suitable for applications in marine environments that require wide coverage and low timeliness requirements.
[0068] Layer 4 (LEO Satellite Layer): This layer consists of a constellation of satellites in low Earth orbit. Layer 4 is primarily designed to provide broad global communications and data transmission services. Compared to Layer 3, Layer 4 communication latency is higher, but its coverage is wider, making it particularly suitable for long-distance maritime communications.
[0069] S2. Various tasks generated by end users, such as ships and offshore platforms, generate navigation, weather, equipment status, and communication data. Converting this data into binary data through feature extraction improves data processing efficiency. Furthermore, binary data requires less storage space, enabling efficient use of storage resources. This facilitates subsequent processing and analysis, achieving consistent processing across the entire network. This includes constructing communication, latency, energy consumption, and cost models.
[0070] Communication model: The wireless transmission rate can be expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the transmission power and channel gain of the sea surface terminal user, respectively, and σ(t) 2 is the power of the additive white Gaussian noise.
[0071] Delay model: The computational delay and transmission delay models for each layer are: Among them, u l The CPU cycles required for each layer of agents to calculate 1 bit of task, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, f l is the maximum computing resource of each layer of agents, β l is the percentage of computing resources allocated to each layer of agents. Then the total delay T(t) of the system is the maximum value of each layer, that is,
[0072] Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is: Among them, the computing power of each layer of intelligent agent is: Where μ≥0 is the effective switching capacitance, the total energy consumption generated by the computing task can be expressed as: The transmission energy consumption model of task offloading to each layer is: Where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t. The total energy consumption generated by the transmission task can be expressed as: Therefore, the total energy consumption of the system is E(t) = E comp (t)+E tran (t).
[0073] Cost model: The model of the computational cost generated by the computational tasks in each layer is: in is the cost of processing each unit of data by each layer agent in each time slot t. The transmission cost model of task offloading to each layer is: The total cost of the system is C(t) = C comp (t)+C tran (t).
[0074] S3: Each server in each layer is considered an agent. The state space and action space of each agent are defined, including information such as task allocation, computing resources, and communication bandwidth. The action of each agent is the proportion of computing tasks in each time slot. This is defined as a partial Markov decision process. Specifically: In the definition of the partial Markov decision process quintuple,
[0075] State space: In each time slot t, the environment state can be expressed as: S(t) = {α l M(t),B(t),f l ,βl ,u l |l∈L}. The amount of task data per layer is α l M(t), the communication wireless bandwidth is B(t). The maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The CPU cycles required to compute a 1-bit task for each layer of the agent.
[0076] Action space: In each time slot t, the action of the agent can be expressed as: A(t) = {α l ,β l |l∈L}. Decisions are made based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents.
[0077] Reward function: The reward function is the key to measuring the quality of the agent's decision at each step. In each time slot t, the future rewards are usually weighted summed with a discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t = -(η1T(t)+η2E(t)+η3C(t)), where η1, η2, and η3 are the weighting factors for the total system latency, energy consumption, and cost, respectively. The total reward R(t) is the cumulative reward value obtained during the entire decision-making process.
[0078] The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem can be defined as:
[0079]
[0080] Among them, C1 represents the guaranteed computing delay of the task at each layer and the transmission time of each layer This should be within the maximum latency tolerable for the system. C2 ensures that the sum of the mission data allocated to each layer equals the total mission data volume. C3 represents the cost per unit of data processed by each layer of servers, from lowest to highest: maritime end users, ground edge servers, HAPS with onboard computing servers, and LEO satellites with onboard computing servers. C4 indicates that both the mission data volume and processing cycle are positive.
[0081] S4. After the task offloading decision is made, the agent performs resource scheduling through an optimization algorithm. This requires coordinating resources between multiple layers of models to ensure that computing and communication resources are properly allocated during the task offloading process. Specifically, after each agent performs task offloading and resource allocation actions, it updates its state based on environmental feedback. The MAPPO algorithm updates the agent's strategy through multiple rounds of training, enabling each agent to continuously optimize its task offloading and resource allocation strategies, gradually improving the overall performance of the system. Ultimately, the system outputs the task offloading decision and the corresponding performance indicators (stable cumulative rewards, optimized latency, energy consumption, and cost).
[0082] Furthermore, the process of optimizing task processing and resource allocation in a multi-level architecture based on the MAPPO algorithm specifically includes the following steps:
[0083] S41, the algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively.
[0084] S42. Initialize a resource-constrained environment and a memory buffer U, where U is used to store experience data generated during the training process.
[0085] The main loop of the S43 and MAPPO algorithms starts from each training cycle E, first initializing the environment state and agent action, and then iterating according to the maximum number of steps. Each step is carried out in a multi-level architecture, by traversing each layer (sea terminal user layer, ground edge server layer, HAPS server layer and LEO satellite layer), and performing the agent n in each layer. l Perform task assignment and action selection.
[0086] S44. In the specific operation of each layer, each agent first observes its current local state (i.e., partial observation), and then θ , select an optimal action a.
[0087] S45. Then, the MAPPO algorithm passes through the value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t To optimize long-term gains, the algorithm accumulates the total reward R over all steps using a discount factor γ.
[0088] S46. The environment updates the state S according to the agent's actions and provides new input for the next iteration.
[0089] S47. After each cycle, the algorithm stores the accumulated experience data (stable cumulative rewards, optimized delay, energy consumption and cost) into the memory buffer U.
[0090] S48. After completing all training cycles, the algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration. Through this process, the algorithm can efficiently handle task offloading problems in multi-level complex environments and significantly optimize the system's latency, energy consumption, and cost.
[0091] See also Figure 3 , Figure 3 This is an optimization method for multi-layer offloading of ocean-surface stereo satellite networks. Figure 3 The MAPPO algorithm framework shown in Figure 1 quantifies computing resources, communication bandwidth, latency, and task requirements in a multi-layer architecture through environmental modeling. Agents observe the current environmental state to generate optimal action strategies, including task offloading targets and resource allocation methods. The MAPPO algorithm leverages multi-agent collaboration to dynamically optimize task offloading and resource scheduling through an actor-critic architecture to minimize system latency, reduce energy consumption, and improve resource utilization. Each agent trains independently in a shared environment, and the PPO pruning update strategy ensures learning stability and convergence, thereby achieving global optimization in complex dynamic environments.
[0092] This invention achieves intelligent decision-making and global optimization for task offloading in a multi-layer architecture through multi-agent collaboration. The MAPPO algorithm leverages the state perception capabilities of the multi-layer network structure to dynamically model the computing resources, communication bandwidth, and latency of the LEO satellites, HAPS, ground edge servers, and end-user devices in the architecture model, constructing a global state and local observations. The agents generate task offloading targets and resource allocation plans based on the current environmental state. By sharing global information, they collaboratively learn to reduce task latency, energy consumption, and cost. The MAPPO algorithm, an extension of the PPO algorithm, ensures training stability through policy updates using a tailored strategy. It utilizes an advantage function to guide policy optimization and designs a reward function that comprehensively considers task completion time, resource utilization, and energy consumption, achieving efficient offloading and resource scheduling in a dynamically changing environment. After multiple rounds of interaction and optimization, the agents learn the optimal strategy, enabling the system to adapt to complex dynamic network conditions and achieve global optimization for task offloading in a multi-layer architecture. The introduction of the MAPPO algorithm significantly improves task processing efficiency and system stability, significantly optimizes resource utilization in multi-layer satellite networks, and provides an effective solution to address the high latency, uneven resource distribution, and energy consumption issues in maritime communications.
Claims
1. A task offloading and resource allocation system for a sea-surface stereo satellite network, characterized in that: The system is implemented based on a multi-layer offloading model of a three-dimensional satellite network facing the sea surface, which includes a space-based network, an air-based network and a ground-sea integrated network. The space-based network includes LEO satellites equipped with onboard computing services; The air-based network includes a high-altitude platform system (HAPS) equipped with airborne computing services, and the ground-sea integrated network includes ground edge servers and sea-surface terminal users; Each server in each layer of the network is regarded as an agent. The agents in each layer have different capabilities, resources and goals, forming a heterogeneous agent system. The heterogeneous agents include defining the state space and action space of each agent, including information such as task allocation, computing resource conditions, and communication bandwidth. The action of each agent is the proportion of computing tasks in each time slot, and it is defined as a partial Markov decision process. After performing task offloading and resource allocation actions, each agent updates its state based on environmental feedback. The agent's strategy is updated based on multiple rounds of training based on the MAPPO algorithm, so that each agent can continuously optimize the task offloading and resource allocation strategies, gradually improving the overall performance of the system. Finally, it includes outputting the task offloading decision and corresponding performance indicators. The performance indicators include stable cumulative rewards, optimized delay, energy consumption and cost. The strategy for updating the agent based on multiple rounds of training using the MAPPO algorithm is as follows: (1) The algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively; (2) Initialize a resource-constrained environment and a memory buffer U to store the experience data generated during the training process; (3) The main loop of the MAPPO algorithm starts from each training cycle E. First, the environment state and agent action are initialized, and then it is iterated according to the maximum number of steps. Each step is carried out in a multi-level architecture. By traversing each layer, the agent n in each layer is initialized. l Perform task assignment and action selection; (4) In the specific operation of each layer, each agent first observes its current local state and then θ , select an optimal action a; (5) MAPPO algorithm through value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t ; The MAPPO algorithm involves accumulating the total reward R of all steps using a discount factor γ to optimize long-term returns; (6) The environment updates the state S according to the agent’s actions and provides new input for the next iteration; (7) After each cycle, the MAPPO algorithm stores the accumulated experience data, including the stable cumulative reward, optimized delay, energy consumption, and cost, into the memory buffer U; (8) After completing all training cycles, the MAPPO algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration; The task offloading and resource allocation method of this system is: The tasks generated by the sea surface terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}. The set of agents in the system is N, n l Represents the intelligent agents in the lth layer. The intelligent agents in each layer can monitor the communication, load, and resource distribution within the layer, calculate the transmission rate, transmission and computing delay, energy consumption and cost based on the task size assigned to each layer, and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observation.
2. The task offloading and resource allocation system for the ocean surface stereo satellite network according to claim 1, characterized in that: The system's corresponding multi-layered heterogeneous architecture consists of at least four main layers, including surface end users, ground edge servers, HAPS, and LEO satellites. The intelligent agents in each layer can provide appropriate resource support and offloading strategies based on the task, and the offloading models at different levels can cooperate with each other. Various tasks generated by end users are converted into binary data through feature extraction to achieve consistent processing across the entire network, including the construction of communication models, latency models, energy consumption models, and cost models; Communication model: The wireless transmission rate is expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the unit data transmission power and channel gain of the sea surface terminal user in each time slot t, respectively, and σ(t) 2 is the power of additive white Gaussian noise; Delay model: The computational delay of each layer is expressed as The transmission delay models are Among them, u l The CPU cycles required for each layer of agents to calculate 1 bit of task, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, f l is the maximum computing resource of each layer of agents, β l is the percentage of computing resources allocated to each layer of agents; the total delay T(t) is the maximum value of each layer, that is, Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is Among them, the computing power of each layer of intelligent agent is Where μ≥0 is the effective switching capacitance, the total energy consumption generated by the computing task is expressed as: The transmission energy consumption model of task offloading to each layer is: Where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t; the total energy consumption generated by the transmission task is expressed as: Therefore, the total energy consumption of the system is E(t) = E comp (t)+E tran (t); Cost model: The model of the computational cost generated by the computational tasks in each layer is: in is the cost of processing each unit of data by each layer of agents in each time slot t; the transmission cost model of task offloading to each layer is: The cost of transmitting each unit of data in the link for each layer of agents in each time slot t; the total cost of the system is C(t) = C comp (t)+C tran (t).
3. The task offloading and resource allocation system for the ocean surface stereoscopic satellite network according to claim 2, characterized in that: In the multi-layer heterogeneous architecture described in step S1, the functional characteristics of each layer and their collaborative relationships are defined as follows: Sea surface terminal user layer: This layer includes terminal devices deployed on the sea surface, which are used for basic processing including data collection and user interaction; Ground edge server layer: This layer includes edge servers deployed on the ground, which are used to handle tasks that are more complex or require higher computing power than the surface terminal user layer, and also includes processing computing tasks offloaded from the surface terminal user layer; HPAS layer: This layer is deployed on high-altitude platforms on the sea surface, including balloons, airships, or high-altitude base stations. It is used to handle tasks offloaded from the sea surface terminal user layer and / or the ground edge server layer, and handles computing tasks with computing power and latency between the sea surface terminal user layer and the ground edge server layer. LEO satellite layer: This layer consists of a group of satellites in low Earth orbit to provide global communication and data transmission services.
4. The task offloading and resource allocation system for the ocean surface 3D satellite network according to claim 2, characterized in that: The partial Markov decision process described in step S3 is as follows: State space: In each time slot t, the environment state is expressed as: S(t) = {α l M(t),B(t),f l ,β l ,u l |l∈L}; where the amount of task data per layer is α l M(t), the communication wireless bandwidth is B(t), and the maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The number of CPU cycles required to compute a 1-bit task for each layer of the agent; Action space: In each time slot t, the action of the agent is expressed as: A(t) = {α l ,β l |l∈L}; Make decisions based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents; Reward function: In each time slot t, the future rewards are weighted summed using the discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t = -(η1T(t)+η2E(t)+η3C(t)), where η1, η2, and η3 are the weighting factors of the total system delay, energy consumption, and cost, respectively; the total reward R(t) is the cumulative reward value obtained during the entire decision-making process; The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem is defined as: Among them, the constraint C1 means to ensure the computing delay of the task at each layer and the transmission time of each layer It should be within the maximum delay range that the system can tolerate; Constraint C2 ensures that the sum of the task data allocated to each layer should be equal to the total task data volume; Constraint C3 represents the cost of processing each unit of data by the server at each layer, from low to high: marine terminal users, ground edge servers, HAPS with airborne computing servers, and LEO satellites with airborne computing servers; Constraint C4 indicates that both the task data volume and the processing cycle are positive.