Task unloading and resource allocation method and system for sea surface three-dimensional satellite network
By designing a multi-layer offloading architecture model in a sea surface three-dimensional satellite network and combining the MAPPO algorithm, the problems of inefficient communication and heavy computing burden in the marine network are solved, efficient task offloading and resource allocation are achieved, delay and energy consumption are reduced, and system performance is improved.
Patent Information
- Application Number
- CN202510068102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-16
AI Technical Summary
There are problems in marine networks with low communication efficiency and heavy computing burden, especially in dynamic marine environments. Traditional task offloading technology and multi-agent deep reinforcement learning algorithms have limitations in efficient coordination and resource optimization.
A multi-layer offloading architecture model for sea surface three-dimensional satellite network is proposed, combining multi-agent near-end strategy optimization (MAPPO) algorithm to optimize task offload decisions and resource allocation strategies in real time, improve system performance, and reduce latency, energy consumption and costs.
Through the combination of the multi-layer offloading architecture model and the MAPPO algorithm, efficient coordination of task offloading and resource allocation in the marine environment is achieved, task processing speed and accuracy are improved, system delay is reduced, and resource utilization and system stability are improved.
Smart Images

Figure CN119967486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite application technology in sea surface scenarios, and in particular to a technical solution for optimizing task offloading efficiency and resource scheduling based on a multi-layer offloading architecture and a multi-agent deep reinforcement learning algorithm (MARL), which is suitable for task offloading, computing resource allocation, delay optimization, energy consumption reduction and cost control in a sea surface stereo satellite network. Background Art
[0002] With the growing demand for marine communications, especially in offshore and marine applications, satellite networks have become a key solution to overcome the inherent limitations of traditional communication methods. However, marine users in remote areas often face significant challenges, including high latency, limited bandwidth, and constraints on energy and computing power, which hinder the effective execution of complex services. These challenges highlight the dual problems of communication inefficiency and computational burden in marine networks, necessitating innovative strategies for reliable service delivery under demanding conditions. Although task offloading techniques have shown great promise in terrestrial networks, their application in marine environments presents unique difficulties. The unique characteristics of marine networks, including the variability of communication links, the scarcity of computing resources, and the dynamic nature of marine weather, pose significant obstacles to the effective implementation of task offloading and deserve further research and development.
[0003] A satellite-drone-user multi-layer architecture has been proposed as a promising solution to enhance task offloading and resource allocation in marine networks. The architecture offers key advantages, including expanded communication coverage, reduced latency through drone relays, and support for distributed task processing. However, significant challenges hinder its practical deployment. The reliance on drones introduces serious weaknesses, such as limited flight time due to energy constraints, susceptibility to adverse weather conditions, and difficulty in maintaining stable connectivity with satellites and users. In addition, the architecture requires complex coordination across multiple network layers, which often leads to inefficiencies and increased latency in dynamic marine environments. Addressing continuous availability, meeting quality of service requirements, and adapting to environmental changes further exacerbate deployment difficulties. Among these challenges, the operational limitations of drones, a key component of the architecture, need to be thoroughly investigated.
[0004] Specifically, environmental factors, including distance, severe weather, and dynamic ocean conditions, often degrade the communication link between satellites and UAVs, resulting in interference, increased latency, and reduced reliability. In high-traffic scenarios, satellite bandwidth limitations exacerbate inefficient resource allocation and reduce service quality, especially for data-intensive missions such as video streaming or emergency transmissions. UAVs are constrained by limited battery capacity, payload limitations, and sensitivity to environmental conditions, and often experience connectivity interruptions during extended operations due to energy exhaustion, insufficient bandwidth, or mission overload. Such interruptions pose a significant challenge to maintaining stable and reliable communications, especially in dynamic ocean environments. Given the limitations of UAVs in meeting these demands, alternative solutions, including energy-efficient designs, hybrid communication systems, and advanced satellite technologies, deserve further exploration.
[0005] To address the inherent limitations of UAVs, existing research proposes the Sky-Ground-Sea Integrated Network (SAGSIN) as a robust and scalable solution for maritime communications. This multi-layer architecture leverages the collaborative operation of low Earth orbit (LEO) satellites and high altitude platform systems (HAPS) to improve communication relay, computational offloading, and resource optimization. LEO satellites are equipped with onboard computing servers to enhance computing power, while HAPS can provide longer operation time and wider coverage compared to UAVs, making them particularly suitable for wide-area maritime surface communications. The integration of HAPS and LEO satellites in SAGSIN not only significantly expands network coverage in remote ocean areas, but also alleviates the latency and reliability challenges associated with UAV-based architectures. In addition, the architecture facilitates efficient offloading of tasks by optimizing resource scheduling, making it a promising solution to address the complex requirements of marine communication networks.
[0006] However, despite these advantages, implementing SAGSIN, especially its LEO satellite-HAPS-user architecture, still faces significant challenges due to the complexity of managing complex communication links and allocating heterogeneous server resources across multiple layers. As a key component of SAGSIN, the architecture must adapt to a variety of offloaded task sources, including ships, buoys, and offshore platforms, which vary greatly in task types and computational requirements. These variations complicate resource allocation and require efficient conversion of tasks into unified binary data through feature extraction for consistent processing across the network. In addition, the heterogeneous computing capabilities of nodes at different network layers require advanced scheduling algorithms to optimize resource utilization and maintain system performance. These challenges highlight the need for further research to improve the operational efficiency of the SAGSIN framework in realistic ocean scenarios.
[0007] Based on the task offloading and resource optimization challenges in the SAGSIN architecture, deep reinforcement learning (DRL) has shown great potential in managing dynamic and uncertain conditions in multi-layer networks. However, traditional multi-agent reinforcement learning (MARL) algorithms, such as multi-agent deep Q-network (MADQN) and multi-agent dominant actor-critic (MAA2C), show limitations in scalability, adaptability, and coordination efficiency in highly dynamic marine environments. Specifically, MADQN suffers from instability and inefficiency when dealing with large state-action spaces, while MAA2C suffers from high computational complexity and suboptimal coordination under rapidly changing conditions. In contrast, multi-agent proximal policy optimization (MAPPO) exhibits superior sample efficiency, stability, and improved inter-agent coordination through centralized training and decentralized execution, making it particularly suitable for dynamic and complex environments. MAPPO enhances training stability in heterogeneous scenarios using clipping updates and implements a centralized training, decentralized execution (CTDE) framework to balance global information sharing with real-time agent autonomy. In addition, it effectively promotes cooperation and competition among agents, thereby achieving efficient task allocation and resource optimization in complex scenarios. These properties make MAPPO a powerful approach to addressing computational and communication challenges in marine applications.
[0008] To this end, the present invention proposes a multi-layer offloading architecture model for the ocean surface stereo satellite network. Combined with the MAPPO method, by optimizing resource scheduling and task offloading strategy, the task offloading efficiency in the ocean surface stereo satellite network is effectively improved, system latency is reduced, energy is saved and cost is optimized. Summary of the invention
[0009] The main purpose of the present invention is to provide a task offloading and resource allocation method and system for the ocean surface stereo satellite network. By designing a multi-layer offloading architecture model, combined with the MAPPO algorithm, under the collaboration of multiple intelligent agents, the task offloading decision and resource allocation strategy in the ocean surface stereo satellite network are optimized in real time, thereby effectively improving the system performance and minimizing the delay, energy consumption and cost.
[0010] Technical solution: A task offloading and resource allocation system for a sea-surface stereoscopic satellite network, which is implemented based on a multi-layer offloading model for a sea-surface stereoscopic satellite network, wherein the multi-layer offloading model includes a space-based network, an air-based network, and a land-sea integrated network;
[0011] The space-based network includes LEO satellites equipped with airborne computing services;
[0012] The air-based network includes a high-altitude platform system HAPS equipped with airborne computing services, and the ground-sea integrated network includes ground edge servers and sea surface terminal users;
[0013] The servers in each layer of the network are regarded as an agent, and the agents in each layer have different capabilities, resources, and goals, forming a heterogeneous agent system;
[0014] The task offloading and resource allocation method of this system is:
[0015] The tasks generated by the sea surface terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}. The set of agents in the system is N, n l Represents the agents in the lth layer. The agents in each layer can monitor the communication, load, and resource distribution within the layer, calculate the transmission rate, transmission and computing delays, energy consumption and cost using the task size assigned to each layer, and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observations.
[0016] Furthermore, the system adopts a multi-layer unloading method for the ocean surface stereo satellite network, which includes:
[0017] S1. Build a multi-layer heterogeneous architecture with at least four main layers, including sea surface end users, ground edge servers, HAPS and LEO satellites. The agents in each layer can provide appropriate resource support and offloading strategies according to the task, and the offloading models at different levels can cooperate with each other.
[0018] S2. Various tasks generated by end users are converted into binary data through feature extraction to achieve consistent processing of the entire network, including the construction of communication models, delay models, energy consumption models and cost models;
[0019] Communication model: The wireless transmission rate is expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the transmission power and channel gain of the sea surface terminal user, respectively, and σ(t) 2 is the power of additive white Gaussian noise;
[0020] Delay model: The computational delay of each layer is expressed as The transmission delay models are: Among them, u l The CPU cycles required for calculating 1 bit of task for each layer of agents, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, f l is the maximum computing resource of each layer of agents, βl is the percentage of computing resources allocated to each layer of agents; the total delay T(t) is the maximum value of each layer, that is,
[0021] Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is Among them, the computing power of each layer of agents is Where μ ≥ 0 is the effective switch capacitance, the total energy consumption generated by the computing task is expressed as: The transmission energy consumption model of task offloading to each layer is: Where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t; the total energy consumption generated by the transmission task is expressed as: Therefore, the total energy consumption of the system is E(t) = E comp (t)+E tran (t);
[0022] Cost model: The model of the computational cost generated by the computational tasks in each layer is: in is the cost of processing each unit of data by each layer of agents in each time slot t; the transmission cost model of task offloading to each layer is: is the cost of transmitting each unit of data in the link for each layer of agents in each time slot t; the total cost of the system is C(t) = C comp (t)+C tran (t);
[0023] S3. Each layer of servers is regarded as an intelligent agent, and the state space and action space of each intelligent agent are defined, including information such as task allocation, computing resources, and communication bandwidth. The action of each intelligent agent is the proportion of computing tasks in each time slot, and it is defined as a partial Markov decision process;
[0024] S4. After executing task offloading and resource allocation actions, each agent updates its state according to environmental feedback, and updates the agent's strategy based on multiple rounds of training based on the MAPPO algorithm, so that each agent can continuously optimize the task offloading and resource allocation strategies, and gradually improve the overall performance of the system. Finally, it includes outputting the task offloading decision and the corresponding performance indicators, and the performance indicators include stable cumulative rewards, optimized delays, energy consumption and costs.
[0025] Furthermore, in the multi-layer heterogeneous architecture described in step S1, the functional characteristics of each layer and their collaborative relationships are defined as follows:
[0026] Sea surface terminal user layer: This layer includes terminal devices deployed on the sea surface, which are used for basic processing including data collection and user interaction;
[0027] Ground edge server layer: This layer includes edge servers deployed on the ground, which are used to process tasks that are more complex or require higher computing power than the sea surface terminal user layer, and include computing tasks that are offloaded from the sea surface terminal user layer;
[0028] HPAS layer: This layer is deployed on high-altitude platforms on the sea surface, including balloons, airships, or high-altitude base stations. It is used to process tasks offloaded from the sea surface terminal user layer and / or the ground edge server layer, and to process computing tasks with computing power and latency between the sea surface terminal user layer and the ground edge server layer.
[0029] LEO Satellite Layer: This layer consists of a group of satellites in low Earth orbit to provide global communication and data transmission services.
[0030] In the above scheme, the partial Markov decision process described in step S3 is as follows:
[0031] State space: In each time slot t, the environment state is expressed as: S(t) = {α l M(t),B(t),f l ,β l ,u l |l∈L}; the amount of task data per layer is α l M(t), the communication wireless bandwidth is B(t), and the maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The CPU cycles required to compute 1 bit of tasks for each layer of agents;
[0032] Action space: In each time slot t, the action of the agent is expressed as: A(t) = {α l ,β l |l∈L}; Make decisions based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents;
[0033] Reward function: In each time slot t, the future rewards are weighted summed using the discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t =-(η1T(t)+η2E(t)+η3C(t)), η1, η2 and η3 are the weight factors of the total system delay, energy consumption and cost respectively; the total reward R(t) is the cumulative reward value obtained in the entire decision-making process;
[0034] The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem is defined as:
[0035]
[0036] Among them, constraint C1 means to ensure the computing delay of the task at each layer and the transmission time of each layer It should be within the maximum delay range that the system can tolerate; Constraint C2 ensures that the sum of the task data allocated to each layer should be equal to the total task data; Constraint C3 represents the cost of processing each unit of data by the server at each layer, from low to high: marine terminal users, ground edge servers, HAPS with airborne computing servers, and LEO satellites with airborne computing servers; Constraint C4 indicates that both the task data volume and the processing cycle are positive.
[0037] 5. The task offloading and resource allocation system for the sea surface stereo satellite network according to claim 2 is characterized in that the strategy of updating the agent based on multiple rounds of training of the MAPPO algorithm in step S4 is as follows:
[0038] S41, the algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively;
[0039] S42, initializing a resource-constrained environment and a memory buffer U for storing experience data generated during the training process;
[0040] S43, the main loop of the MAPPO algorithm starts from each training cycle E, first initializing the environment state and agent action, and then iterating according to the maximum number of steps; each step of the operation is carried out in a multi-level architecture, by traversing each layer, the agent n in each layer is initialized. l Perform task allocation and action selection;
[0041] S44. In the specific operation of each layer, each agent first observes its current local state (i.e., partial observation), and then θ , select an optimal action a;
[0042] S45, MAPPO algorithm through value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t ,
[0043] The MAPPO algorithm involves accumulating the total reward R of all steps using a discount factor γ to optimize long-term benefits;
[0044] S46, the environment updates the state S according to the agent's actions and provides new input for the next iteration;
[0045] S47. After each cycle, the MAPPO algorithm stores the accumulated experience data, including the stable accumulated rewards, optimized delays, energy consumption, and costs, into the memory buffer U;
[0046] S48. After completing all training cycles, the MAPPO algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration.
[0047] Beneficial effects: Compared with the prior art, the substantial features and significant improvements of the solution of the present invention include:
[0048] (1) The present invention combines a multi-layer offloading architecture model with a multi-agent deep reinforcement learning optimization algorithm to achieve efficient coordination of task offloading and resource allocation in a complex marine environment, thereby improving the speed and accuracy of task processing and reducing system latency.
[0049] (2) The task offloading and resource scheduling method proposed in the present invention can realize adaptive scheduling decisions according to the dynamic changes of the ocean surface stereo satellite network and different task requirements, thereby effectively improving resource utilization.
[0050] (3) The present invention reduces energy consumption in the computing and communication processes, reduces unnecessary resource consumption, effectively prolongs the service life of satellites and equipment, and reduces the operating costs of the system.
[0051] (4) Through the optimization of the MAPPO algorithm, the present invention can achieve efficient collaboration and decision-making in a multi-agent environment, has strong robustness, and can adapt to the complex changes in the marine environment.
[0052] Therefore, the present invention establishes a multi-layer offloading system model of LEO satellite-HAPS-user, takes the weighted minimization of the time delay, energy consumption and cost of system processing tasks as the optimization problem, and proposes a joint optimization algorithm for task offloading and resource allocation based on the MAPPO method. Experimental verification shows that compared with other existing methods, the model and method of the present invention can effectively solve the problem of collaborative task offloading for sea users, with less delay, energy consumption and cost, and optimizes the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A schematic diagram of the sea surface stereo satellite network according to the present invention;
[0054] Figure 2 It is a flowchart of a multi-layer offloading decision algorithm in a specific embodiment of the present invention;
[0055] Figure 3 This is a structural diagram of the MAPPO unloading decision algorithm of the present invention. DETAILED DESCRIPTION
[0056] Embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and cannot be understood as limitations of the present invention.
[0057] In a sea area, the present invention constructs a multi-layer unloading system and its model for sea surface terminal users, which consists of a space-based network, an air-based network and a land-sea integrated network. The space-based network has LEO satellites equipped with airborne computing services, the air-based network has a high-altitude platform system (HAPS) equipped with airborne computing services, and the land-sea integrated network has ground edge servers and sea surface terminal users. The servers in each layer of the network are regarded as an intelligent agent, and each layer of intelligent agents has different capabilities, resources and goals, forming a heterogeneous intelligent agent system.
[0058] The tasks generated by the sea surface terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks have a certain amount of data and are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}, and the set of agents in the system is N, n l Represents the agent in layer l. The agents in each layer can monitor the communication, load, resource distribution, etc. within the layer, and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observations.
[0059] In the process of defining a joint optimization model that takes into account system latency, energy consumption, and cost, the model is calculated using the size of the task assigned to each layer to calculate the transmission rate, transmission and computing latency, energy consumption, and cost. The agents in each layer can provide appropriate resource support and offloading strategies based on the task. Offloading models at different levels can cooperate with each other to ensure efficient offloading of tasks. Ultimately, the total latency, energy consumption, and cost of the system are minimized.
[0060] In combination with the above communication system, a multi-layer unloading method to a sea surface stereo satellite network is provided, comprising:
[0061] Intelligent decision-making and global optimization of task offloading under multi-layer architecture are achieved through multi-agent collaboration. The MAPPO algorithm uses the state perception capability in the multi-layer network structure to dynamically model the computing resources, communication bandwidth and delay of LEO satellites, HAPS, ground edge servers and end-user devices in the architecture model, and construct global states and local observations. The agent generates task offloading goals and resource allocation plans based on the current environmental state, and collaboratively learns by sharing global information to reduce the latency, energy consumption and cost of task calls. The MAPPO algorithm is based on the PPO extension, ensures training stability by updating the clipping strategy, guides strategy optimization by using the advantage function, and designs a reward function that comprehensively considers task completion time, resource utilization and energy consumption, achieving efficient offloading and resource scheduling in a dynamically changing environment. After multiple rounds of interaction and optimization, the agent learns the optimal strategy, enabling the system to adapt to complex dynamic network conditions and achieve global optimization of task offloading in a multi-layer architecture. The introduction of the MAPPO algorithm greatly improves the task processing efficiency and system stability, can significantly optimize the resource utilization of multi-layer satellite networks, and provides an effective solution to the problems of high latency, uneven resource distribution and energy consumption in sea surface communications.
[0062] For details, see Figure 1 , Figure 1 It is a communication system established by the present invention, and also an application environment for multi-layer unloading of a sea surface stereo satellite network. Among them, the space-based network includes several LEO satellites equipped with airborne computing services, the air-based network includes several HAPS equipped with airborne computing services, and the land-sea integrated network includes several ground edge servers and sea surface terminal users. Based on the multi-layer unloading model, the computing power of the sea surface terminal users is the weakest, the computing resources are the least, but the computing cost is the lowest. The computing power of LEO satellites and ground edge servers is strong, and they have more computing resources, but the server cost is high. HAPS is equipped with certain computing and storage resources, but it is far less than that of ground edge servers, and the cost is also lower than that of ground edge servers. When a sea surface terminal user generates a task, the terminal user passes the task-related information to the intelligent body. The intelligent body in the sea surface network will decide whether the task can be processed directly or needs to be unloaded to other devices (LEO satellites, HAPS or ground edge servers) based on the current network conditions, load conditions and resource availability.
[0063] See also Figure 2 , Figure 2 The flowchart of the multi-layer offloading decision algorithm based on MAPPO includes the following steps:
[0064] S1. The system builds a multi-layered heterogeneous architecture covering four main levels: sea surface end users, ground edge servers, HAPS and LEO satellites. The core of this modeling process is to accurately define the functional characteristics of each layer and their synergy to adapt to different mission requirements and network environments. Specifically:
[0065] Layer 1 (sea surface terminal user layer): This layer is mainly composed of terminal devices located on the sea surface, such as marine smart devices, ships, etc. Layer 1 is responsible for processing simple tasks that terminal devices can directly calculate, such as data collection, basic processing, user interaction, etc.
[0066] Layer 2 (ground edge server layer): This layer is located on the ground and usually includes edge servers with high-performance computing resources. Compared with layer 1, layer 2 has stronger computing resources and can handle tasks with high computing power and low latency requirements. It is suitable for offloading computing tasks from layer 1, especially in scenarios with high timeliness requirements, such as real-time data processing, complex analysis, etc.
[0067] Layer 3 (HPAS): This layer is deployed on high-altitude platforms above the sea, such as balloons, airships, or high-altitude base stations. Layer 3 is suitable for processing tasks that have requirements for computing resources and timeliness between layers 1 and 2, but it can provide highly flexible and low-cost task offloading, and is particularly suitable for application scenarios in marine environments that require large-scale coverage and have low timeliness requirements.
[0068] Layer 4 (LEO satellite layer): This layer consists of a group of satellites in low earth orbit. Layer 4 is mainly used to provide extensive global communication and data transmission services. Compared with layer 3, layer 4 has higher communication latency, but its coverage is wider and is particularly suitable for long-distance communications at sea.
[0069] S2. Various tasks generated by end users, ships, offshore work platforms, generate navigation, weather, equipment status and communication task data. Converting these data into binary data for processing through feature extraction can improve data processing efficiency. Secondly, binary data has a small storage space and can efficiently utilize storage resources. This facilitates subsequent processing and analysis to achieve consistent processing of the entire network, including the construction of communication models, delay models, energy consumption models and cost models;
[0070] Communication model: The wireless transmission rate can be expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the transmission power and channel gain of the sea surface terminal user, respectively, and σ(t) 2 is the power of the additive white Gaussian noise.
[0071] Delay model: The calculation delay and transmission delay models of each layer are: Among them, u l The CPU cycles required for calculating 1 bit of task for each layer of agents, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, f l is the maximum computing resource of each layer of agents, β l is the percentage of computing resources allocated to each layer of agents. Then the total delay T(t) of the system is the maximum value of each layer, that is,
[0072] Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is: Among them, the computing power of each layer of agents is: Where μ ≥ 0 is the effective switch capacitance, the total energy consumption generated by the computing task can be expressed as: The transmission energy consumption model of task offloading to each layer is: Where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t. The total energy consumption generated by the transmission task can be expressed as: Therefore, the total energy consumption of the system is E(t) = E comp (t)+E tran (t).
[0073] Cost model: The model of the computational cost generated by the computational tasks in each layer is: in is the cost of processing each unit of data by each layer of agents in each time slot t. The transmission cost model of task offloading to each layer is: is the cost of transmitting each unit of data in the link for each layer of agents in each time slot t. Then the total cost of the system is C(t) = C comp (t)+C tran (t).
[0074] S3: Each server in each layer is considered as an agent. Define the state space and action space of each agent, including task allocation, computing resources, communication bandwidth and other information. The action of each agent is the proportion of computing tasks in each time slot. And define it as a partial Markov decision process. Specifically: In the definition of the five-tuple of the partial Markov decision process,
[0075] State space: In each time slot t, the environment state can be expressed as: S(t) = {α l M(t), B(t), f l , β l ,ul |l∈L}. The amount of task data at each layer is α l M(t), the communication wireless bandwidth is B(t). The maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The CPU cycles required to compute a 1-bit task for each layer of the agent.
[0076] Action space: In each time slot t, the action of the agent can be expressed as: A(t) = {α l , β l |l∈L}. Decisions are made based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents.
[0077] Reward function: The reward function is the key to measuring the quality of the agent's decision at each step. In each time slot t, the future rewards are usually weighted summed with a discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t =-(η1T(t)+η2E(t)+η3C(t)), where η1, η2 and η3 are the weight factors of the total system delay, energy consumption and cost respectively. The total reward R(t) is the cumulative reward value obtained during the entire decision-making process.
[0078] The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem can be defined as:
[0079]
[0080]
[0081] Among them, C1 represents the guaranteed computing delay of the task at each layer and the transmission time of each layer It should be within the maximum delay range that the system can tolerate. C2 ensures that the sum of the task data allocated to each layer should be equal to the total task data. C3 represents the cost of processing each unit of data by each layer of server, from low to high: marine end users, ground edge servers, HAPS with airborne computing servers, LEO satellites with airborne computing servers. C4 indicates that both the task data volume and the processing cycle are positive.
[0082] S4. After the task offloading decision is made, the agent performs resource scheduling through the optimization algorithm. It is necessary to coordinate resources between multi-layer models to ensure that the computing and communication resources are reasonably allocated during the task offloading process. Specifically, after each agent performs task offloading and resource allocation actions, it updates its state based on environmental feedback. The MAPPO algorithm updates the agent's strategy through multiple rounds of training, allowing each agent to continuously optimize the task offloading and resource allocation strategies and gradually improve the overall performance of the system. Ultimately, the system will output the task offloading decision and the corresponding performance indicators (stable cumulative rewards, optimized delays, energy consumption, and costs).
[0083] Furthermore, the process of optimizing task processing and resource allocation under the multi-level architecture based on the MAPPO algorithm specifically includes the following steps:
[0084] S41, the algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively.
[0085] S42. Initialize a resource-constrained environment and a memory buffer U, where U is used to store experience data generated during the training process.
[0086] S43, the main loop of the MAPPO algorithm starts from each training cycle E, first initializing the environment state and agent action, and then iterating according to the maximum number of steps. Each step of the operation is carried out in a multi-level architecture, by traversing each layer (sea surface terminal user layer, ground edge server layer, HAPS server layer and LEO satellite layer), the agent n in each layer is initialized. l Perform task assignment and action selection.
[0087] S44. In the specific operation of each layer, each agent first observes its current local state (i.e., partial observation), and then θ , select an optimal action a.
[0088] S45. Then, the MAPPO algorithm passes through the value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t To optimize the long-term reward, the algorithm accumulates the total reward R of all steps using a discount factor γ.
[0089] S46. The environment updates the state S according to the agent's actions and provides new input for the next iteration.
[0090] S47. After each cycle, the algorithm stores the accumulated experience data (stable cumulative rewards, optimized delays, energy consumption, and costs) into the memory buffer U.
[0091] S48. After completing all training cycles, the algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration. Through this process, the algorithm can efficiently handle task offloading problems in multi-level complex environments and significantly optimize the system's latency, energy consumption, and cost.
[0092] See also Figure 3 , Figure 3 This is an optimization method for multi-layer offloading of ocean surface stereo satellite networks. Figure 3 The MAPPO algorithm framework shown in the figure quantifies the computing resources, communication bandwidth, latency, and task requirements in the multi-layer architecture through environmental modeling. The agent generates the optimal action strategy by observing the current environmental state, including task offloading goals and resource allocation methods. The MAPPO algorithm uses multi-agent collaboration to dynamically optimize task offloading and resource scheduling through the Actor-Critic architecture to minimize system latency, reduce energy consumption, and improve resource utilization. Each agent is trained independently in a shared environment, and the PPO pruning update strategy is used to ensure the stability and convergence of learning, thereby achieving global optimization in a complex dynamic environment.
[0093] The present invention realizes intelligent decision-making and global optimization of task offloading under a multi-layer architecture through multi-agent collaboration. The MAPPO algorithm uses the state perception capability in the multi-layer network structure to dynamically model the computing resources, communication bandwidth and delay of LEO satellites, HAPS, ground edge servers and end-user devices in the architecture model, and construct global states and local observations. The agent generates task offloading targets and resource allocation plans according to the current environmental state, and collaboratively learns by sharing global information to reduce the delay, energy consumption and cost of task calls. The MAPPO algorithm is based on the PPO extension, ensures training stability by updating the clipping strategy, guides strategy optimization by using the advantage function, and designs a reward function that comprehensively considers task completion time, resource utilization and energy consumption, so as to achieve efficient offloading and resource scheduling in a dynamically changing environment. After multiple rounds of interaction and optimization, the agent learns the optimal strategy, enabling the system to adapt to complex dynamic network conditions and achieve global optimization of task offloading in a multi-level architecture. The introduction of the MAPPO algorithm greatly improves the task processing efficiency and system stability, can significantly optimize the resource utilization of multi-layer satellite networks, and provides an effective solution to solve the problems of high latency, uneven resource distribution and energy consumption in sea surface communications.
Claims
1. A task offloading and resource allocation system for a sea surface stereo satellite network, characterized in that: The system is implemented based on a multi-layer offloading model of a three-dimensional satellite network facing the sea surface, and the multi-layer offloading model includes a space-based network, an air-based network and a land-sea integrated network; The space-based network includes LEO satellites equipped with airborne computing services; The air-based network includes a high-altitude platform system HAPS equipped with airborne computing services, and the ground-sea integrated network includes ground edge servers and sea surface terminal users; The servers in each layer of the network are regarded as an agent, and the agents in each layer have different capabilities, resources, and goals, forming a heterogeneous agent system; The task offloading and resource allocation method of this system is: The tasks generated by the sea surface terminal users in each time slot t∈T={1,2,…,|T|} are recorded as D(t). The tasks are split and unloaded to the servers of other layers in the model. The set of layers is defined as l∈L={1,2,…,|L|}. The set of agents in the system is N, n l Represents the agents in the lth layer. The agents in each layer can monitor the communication, load, and resource distribution within the layer, calculate the transmission rate, transmission and computing delays, energy consumption and cost using the task size assigned to each layer, and participate in decision-making to optimize their task offloading and resource allocation decisions through partial observations.
2. The task offloading and resource allocation system for the ocean surface stereoscopic satellite network according to claim 1 is characterized in that: The system adopts a multi-layer unloading method for the ocean surface stereo satellite network, which includes: S1. Build a multi-layer heterogeneous architecture with at least four main layers, including sea surface end users, ground edge servers, HAPS and LEO satellites. The agents in each layer can provide appropriate resource support and offloading strategies according to the task, and the offloading models at different levels can cooperate with each other. S2. Various tasks generated by end users are converted into binary data through feature extraction to achieve consistent processing of the entire network, including the construction of communication models, delay models, energy consumption models and cost models; Communication model: The wireless transmission rate is expressed as: Where B(t) is the communication wireless bandwidth, w(t) and h(t) are the transmission power and channel gain of the sea surface terminal user, respectively, and σ(t) 2 is the power of additive white Gaussian noise; Delay model: The computational delay of each layer is expressed as l∈L, the transmission delay models are l∈L; where u l The CPU cycles required for calculating 1 bit of task for each layer of agents, M(t) is the amount of data generated for task D(t) in each time slot t, α l is the ratio of the amount of data in each layer to the total amount of data, and the sum is 1, that is, l∈L;f l is the maximum computational resource of each layer of agents, β l is the percentage of computing resources allocated to each layer of agents; the total delay T(t) is the maximum value of each layer, that is, Energy consumption model: The model of computing energy consumption generated by computing tasks in each layer is l∈L, where the computational power of each layer of agents is Where μ ≥ 0 is the effective switch capacitance, the total energy consumption generated by the computing task is expressed as: The transmission energy consumption model of task offloading to each layer is: l∈L, where w(t) is the unit data transmission power of the sea surface terminal user in each time slot t; the total energy consumption generated by the transmission task is expressed as: l∈L; therefore, the total energy consumption of the system is E(t)=E comp (t)+E tran (t); Cost model: The model of the computational cost generated by the computational tasks in each layer is: l∈L, where is the cost of processing each unit of data by each layer of agents in each time slot t; the transmission cost model of task offloading to each layer is: l∈L, is the cost of transmitting each unit of data in the link for each layer of agents in each time slot t; the total cost of the system is C(t) = C comp (t)+C tran (t); S3. Each layer of servers is regarded as an intelligent agent, and the state space and action space of each intelligent agent are defined, including information such as task allocation, computing resources, and communication bandwidth. The action of each intelligent agent is the proportion of computing tasks in each time slot, and it is defined as a partial Markov decision process; S4. After executing task offloading and resource allocation actions, each agent updates its state according to environmental feedback, and updates the agent's strategy based on multiple rounds of training based on the MAPPO algorithm, so that each agent can continuously optimize the task offloading and resource allocation strategies, and gradually improve the overall performance of the system. Finally, it includes outputting the task offloading decision and the corresponding performance indicators, and the performance indicators include stable cumulative rewards, optimized delays, energy consumption and costs.
3. The task offloading and resource allocation system for the ocean surface stereoscopic satellite network according to claim 2 is characterized in that: In the multi-layer heterogeneous architecture described in step S1, the functional characteristics of each layer and their collaborative relationships are defined as follows: Sea surface terminal user layer: This layer includes terminal devices deployed on the sea surface, which are used for basic processing including data collection and user interaction; Ground edge server layer: This layer includes edge servers deployed on the ground, which are used to process tasks that are more complex or require higher computing power than the sea surface terminal user layer, and include computing tasks that are offloaded from the sea surface terminal user layer; HPAS layer: This layer is deployed on high-altitude platforms on the sea surface, including balloons, airships, or high-altitude base stations. It is used to process tasks offloaded from the sea surface terminal user layer and / or the ground edge server layer, and to process computing tasks with computing power and latency between the sea surface terminal user layer and the ground edge server layer. LEO Satellite Layer: This layer consists of a group of satellites in low Earth orbit to provide global communication and data transmission services.
4. The task offloading and resource allocation system for the ocean surface stereoscopic satellite network according to claim 2 is characterized in that: The partial Markov decision process described in step S3 is as follows: State space: In each time slot t, the environment state is expressed as: S(t) = {α l M(t),B(t),f l ,β l ,u l |l∈L}; the amount of task data in each layer is α l M(t), the communication wireless bandwidth is B(t), and the maximum computing resource of each agent in each layer is f l , the percentage of computing resources allocated to each layer of agents is β l ,u l The CPU cycles required to compute 1 bit of tasks for each layer of agents; Action space: In each time slot t, the action of the agent is expressed as: A(t) = {α l ,β l |l∈L}; Make decisions based on the current state, the amount of task data assigned to each layer of agents, and the percentage of computing resources allocated to each layer of agents; Reward function: In each time slot t, the future rewards are weighted summed using the discount factor γ, and the single-step reward r t Reflects the immediate reward given by the environment after the agent performs action a, r t =-(η1T(t)+η2E(t)+η3C(t)), η1, η2 and η3 are the weight factors of the total system delay, energy consumption and cost respectively; the total reward R(t) is the cumulative reward value obtained in the entire decision-making process; The goal is to minimize the weighted total delay, energy consumption and cost. The joint optimization problem is defined as: Among them, constraint C1 means to ensure the computing delay of the task at each layer and the transmission time of each layer It should be within the maximum delay range that the system can tolerate; Constraint C2 ensures that the sum of the task data allocated to each layer should be equal to the total task data; Constraint C3 represents the cost of processing each unit of data by the server at each layer, from low to high: marine terminal users, ground edge servers, HAPS with airborne computing servers, and LEO satellites with airborne computing servers; Constraint C4 indicates that both the task data volume and the processing cycle are positive.
5. The task offloading and resource allocation system for the ocean surface stereoscopic satellite network according to claim 2 is characterized in that: The strategy for updating the agent based on multiple rounds of training of the MAPPO algorithm in step S4 is as follows: S41, the algorithm first initializes the policy network π θ and value network They are used for action selection and state value evaluation of the agent respectively; S42, initializing a resource-constrained environment and a memory buffer U for storing experience data generated during the training process; S43, the main loop of the MAPPO algorithm starts from each training cycle E, first initializing the environment state and agent action, and then iterating according to the maximum number of steps; each step of the operation is carried out in a multi-level architecture, by traversing each layer, the agent n in each layer is initialized. l Perform task allocation and action selection; S44. In the specific operation of each layer, each agent first observes its current local state and then θ , select an optimal action a; S45, MAPPO algorithm through value network Evaluate the value of the current state and calculate the single-step reward r through environmental feedback t ; The MAPPO algorithm involves accumulating the total reward R of all steps using a discount factor γ to optimize long-term benefits; S46, the environment updates the state S according to the agent's actions and provides new input for the next iteration; S47. After each cycle, the MAPPO algorithm stores the accumulated experience data, including the stable accumulated rewards, optimized delays, energy consumption, and costs, into the memory buffer U; S48. After completing all training cycles, the MAPPO algorithm outputs the final optimized policy network π θ , achieving global optimal performance under multi-agent collaboration.
Citation Information
Patent Citations
Self-adaptive task unloading method considering space-time load in satellite edge calculation
CN117608812A
Cited By
Underwater edge privacy protection task unloading method based on differential federated learning
CN120434626A
Satellite edge computing task unloading method and system
CN120729406A