A method for UAV trajectory optimization and task offloading for integrated air-space-ground networks

By optimizing UAV trajectories and task offloading in an integrated air-space-ground network, and utilizing temporal prediction and multi-agent deep reinforcement learning algorithms, the problems of communication link disconnection and low task offloading efficiency are solved, achieving high-efficiency services with low latency and low energy consumption.

CN119907048BActive Publication Date: 2025-11-14NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510113388.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-11-14
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In an integrated air-space-ground network, existing technologies struggle to effectively address the communication link disruptions caused by network node movement, leading to network service interruptions and increased latency. Furthermore, the lack of efficient task offloading solutions results in decreased system performance.

Method used

We employ time-series prediction algorithms and multi-agent deep reinforcement learning algorithms to optimize UAV trajectories and task offloading. We establish network, communication, and energy consumption models and optimize UAV flight trajectories and task offloading strategies through collaborative cooperation between UAVs and low-Earth orbit satellites, thereby reducing latency and energy consumption.

Benefits of technology

It enables efficient collaboration between UAVs and low-orbit satellites, optimizes system performance, meets users' needs for low latency and low energy consumption, and improves the quality of network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119907048B_ABST
    Figure CN119907048B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of integrated air-space-ground network technology, and discloses a method for UAV trajectory optimization and task offloading in integrated air-space-ground networks. It aims to optimize task offloading efficiency in integrated air-space-ground networks, focusing on energy consumption and latency optimization. The proposed method comprehensively considers various situations in integrated air-space-ground networks, incorporating decoupled communication methods to reduce the impact of service interruptions. An optimization model is constructed to minimize the weighted sum of latency and energy consumption. To improve the collaboration of heterogeneous agents in integrated air-space-ground networks, a hierarchical reinforcement learning framework is used to propose a two-level Qmix algorithm based on hierarchical learning to enhance the collaboration between low-Earth orbit satellites and UAVs. Simulations on the PyCharm platform in multiple scenarios verify that this invention has significant advantages in average response latency, energy consumption, and QoS indicators, effectively improving network performance and meeting user needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated air-space-ground network technology, and in particular to a method for optimizing UAV trajectory and offloading tasks for integrated air-space-ground networks. Background Technology

[0002] With the rapid development of internet technology, the amount of data processed by smart terminal devices is constantly increasing. Due to limitations in hardware technology, these devices cannot provide sufficient computing resources for emerging computationally intensive and latency-sensitive tasks. Since 2008, cloud computing has been applied in this field. Cloud computing centers, relying on their powerful computing capabilities, have overcome the resource limitations of smart devices, providing diverse and efficient applications and computing services. However, while traditional cloud computing can provide ample resources for end users, it cannot meet the latency requirements of rapid application response. Therefore, Mobile Edge Computing (MEC) has emerged as a new technology. MEC is a technology that moves computing tasks from the cloud center to the edge devices, achieving tight integration between the computing platform and the devices. It provides powerful computing and storage capabilities to terminal devices while minimizing energy consumption and latency, thereby enabling users to obtain higher Quality of Service (QoS).

[0003] Traditional MEC (Multi-access Edge Computing) technology is difficult to deploy in complex terrains such as remote mountainous areas and oceans, making it impossible to provide internet services to end users. Furthermore, network infrastructure is vulnerable to damage from natural disasters such as earthquakes and floods, leading to service interruptions. Therefore, a new technology is urgently needed to overcome these limitations. With major technology companies vigorously developing low-Earth orbit (LEO) satellite technologies in recent years, as of May 2024, the number of satellites worldwide exceeded 9,000, with LEO satellites accounting for over 90%. Space-Air-Ground Integrated Network (SAGIN) has become a hot research topic for next-generation mobile communication networks. Integrating various mobile communication devices less affected by environmental constraints, SAGIN is expected to overcome the aforementioned technical difficulties and alleviate the limitations of existing network communication technologies. The SAGIN architecture mainly consists of three parts: a space-based network composed of LEO satellites, an air-based network composed of unmanned aerial vehicles (UAVs), and a ground-based network composed of ground users and traditional ground communication equipment. Of these three components, SAGIN is based on a ground-based network, combined with the wide coverage of satellite-based networks and the flexible deployment advantages of UAVs, achieving seamless coverage across the air, space, and ground domains through deep integration of multi-layered networks. Therefore, integrating MEC technology with SAGIN leverages SAGIN's advantages of flexible deployment, low cost, wide coverage, and multi-layered networks to provide services to terminal devices.

[0004] Current research on the application of MEC technology in SAGIN has yielded many valuable results. However, a series of key issues still need to be addressed: First, during data transmission, the relative movement of network nodes can cause communication link breaks, leading to service interruptions and a significant decrease in user service quality. Current research lacks attention to data transmission after link breaks, resulting in high latency. Second, due to the complexity and mobility of the SAGIN network structure, existing algorithms struggle to quickly find an efficient task offloading solution. Furthermore, the collaboration of various heterogeneous devices in SAGIN is often overlooked, leading to overall system performance degradation. Solving these problems requires in-depth research and innovation to improve system performance and efficiency to meet the ever-growing demands of mobile applications. Summary of the Invention

[0005] To address the aforementioned issues in SAGIN, this invention proposes a UAV trajectory optimization and task offloading method for an integrated air-space-ground network. The core of this method is to utilize the nodes at each layer of SAGIN to provide efficient task offloading services to ground users. The UAV is responsible for collecting user data and deciding whether to perform calculations internally or upload them to the LEO (Local Area Network) for computation. The results from both methods are then returned to the ground users. This reduces task processing latency and energy consumption.

[0006] The technical solution of the present invention is as follows: a method for UAV trajectory optimization and task offloading for an integrated air-space-ground network, which establishes a network model, a node communication model, a task offloading energy consumption model, and a latency model;

[0007] The optimization problem is to minimize the energy consumption and average latency of completing the task. The optimization problem is solved by using a time-series prediction algorithm and a multi-agent deep reinforcement learning algorithm to obtain the optimal trajectory optimization and task unloading strategy.

[0008] The network model is divided into three layers: the ground layer, the air layer, and the space layer.

[0009] The ground floor includes ground users;

[0010] The air layer is mainly composed of unmanned aerial vehicles (UAVs); UAVs establish communication links with ground users, receive tasks uploaded by ground users, and calculate the tasks themselves or transmit them to upper-level satellites for calculation based on decisions; at the same time, there are communication links between UAVs and with satellites to realize information transmission and interaction.

[0011] The space layer includes multiple low-Earth orbit (LEO) satellites in the upper atmosphere; a time-series prediction model based on iTransformer is deployed in it; the LEO satellites are regarded as high-level intelligent agents, which make high-level decisions at regular intervals based on the observed state information and the user distribution predicted by the time-series prediction model based on iTransformer, to guide the low-level strategies of the UAV.

[0012] The network model includes N ground users, M drones, and L low-Earth orbit satellites. The set of ground users is represented as... Its three-dimensional coordinates are represented by l n (t)=(x n (t),y n (t),h n Let (t) represent the coordinates of ground users, which are randomly assigned at the beginning of each round and then randomly moved at the beginning of each time slot. The set of aerial drones is represented as... Its three-dimensional coordinates are represented by l m (t)=(x m (t),y m (t),h m(t) represents a satellite that moves at a speed of V in each time slot and is responsible for collecting data from ground users. The set of low-Earth orbit satellites above it is represented by l∈L={1,2,...,L}, and its three-dimensional coordinates are represented as l l (t)=(x l (t),y l (t),h l (t)), LEO moves along a fixed orbit; ground users can only establish communication with LEO satellites when they enter the communication window.

[0013] In the node communication model, the communication link between the UAV and the ground user is either a line-of-sight (LoS) link or a non-line-of-sight (NLoS) link; the path loss of LoS is... The path loss corresponding to NLoS is Where β LoS and β NLoS These are two different attenuation factors that distinguish between LosS and NloS links, α mn (t) is the power gain, and the average path loss for this communication link is defined as:

[0014]

[0015] Among them, P LoS (t) represents the probability that the UAV and the user establish a LoS channel in a specific network environment, and the probability that they establish an NLoS channel is P. NLoS (t)=1-P LoS (t);

[0016] The communication link between the drone and the ground user is decoupled, separating their uplink and downlink, so that each ground user can establish uplink and downlink communication with different drones. A binary variable is defined. To indicate whether UAVm is associated with user n, it is represented as:

[0017]

[0018] When (·) = D, it indicates that the communication link is a downlink; when (·) = U, it indicates that the link is an uplink. In each time slot, each user is only allowed to establish an uplink or downlink with one UAV. At the same time, due to the limitations of UAV hardware, a UAV can only establish a connection with a maximum of Q ground users in each time slot.

[0019] Considering channel interference during decoupling, communication links are established between UAVs and between UAVs and satellites using line-of-sight transmission. Therefore, the information transmission rates between UAVs and between UAVs and satellites are defined as follows:

[0020]

[0021] Among them B back For backhaul link bandwidth, σ 2 For noise power, p o (t) represents the transmission power;

[0022] The information transmission rate between the drone and the ground user is expressed as:

[0023]

[0024] Let B be the link connection status between UAV m and user n at time t, and let B be the channel bandwidth. This is the signal-to-noise ratio plus interference. For uplink and downlink, the corresponding signal-to-noise ratios are:

[0025]

[0026] Where, p n (t) represents the transmit power of user n, p m (t) represents the transmit power of the drone m. This indicates the interference caused by other users' uplink transmissions to drone m. This indicates the interference caused by downlink transmissions from other drones to drone m. Indicates the self-interference of drone m, This indicates the interference caused to user n by the downlink transmissions of other drones. This indicates the interference caused to user n by the uplink transmissions of other users. This represents the self-interference of user n;

[0027] During the initialization of the task offloading energy consumption model, each ground user is randomly assigned a certain amount of data for computation. The ground user needs to upload the task and transmit it to the UAV via the uplink. The UAV then decides whether to upload the task to the LEO satellite for computation or to perform the computation within itself. Throughout the entire task computation cycle, the energy consumed by the ground user and the UAV is mainly considered, specifically in four aspects:

[0028] (1) Energy consumption for ground user data transmission: Users need to transmit all their data to the UAV, and the data volume is I. n The power of user-uploaded data is defined as p. n (t), therefore the transmission energy consumption is defined as:

[0029]

[0030] in Indicates the uplink information transmission rate;

[0031] (2) Data transmission power consumption of UAVs: When the UAVs that establish uplink and downlink links with the same user are different, the transmission power of the uplink UAV transmitting data to the corresponding downlink UAV, or forwarding data to the LEO satellite, is defined as p. m (t), energy consumption is:

[0032]

[0033] The energy consumption calculated by the UAV and returned to the ground user upon completion of the task calculation is as follows:

[0034]

[0035] I m This indicates the amount of data that drone m needs to transmit. Indicates the information transmission rate of the downlink. This indicates the information transmission rate between drones or between a drone and a satellite.

[0036] (3) Energy consumption of UAV computing tasks: Set the amount of data that the UAV needs to compute as TA. m The calculation speed is The calculated power is P c The energy consumption of a computing task is defined as:

[0037]

[0038] (4) UAV flight energy consumption: The flight energy consumption of UAV is relatively large compared to hovering energy consumption, so the flight energy consumption of UAV is the main consideration. The flight power of UAV is defined as P f Flight energy consumption is defined as:

[0039]

[0040] The latency model mainly includes two parts: task computation latency and information transmission latency. The task computation latency involves UAV computation latency and satellite computation latency, which are defined as follows: Information transmission latency is divided into uplink latency, downlink latency, and transmission latency between UAVs and between UAVs and satellites, defined as:

[0041]

[0042] Among them, I n Indicates the amount of tasks uploaded by ground users, I m Indicates the amount of data transmitted by the drone, I' m Indicates the amount of data transmitted between drones. This indicates the amount of data transmitted by the drone to the low-Earth orbit satellite;

[0043] Therefore, the cycle delay for each task to complete is defined as:

[0044]

[0045] The time-series prediction algorithm divides the ground area into multiple small blocks, using each block as a unit. It simulates the number of ground users within each unit using a real dataset, treating this as historical time-series data. This data is then used to make predictions through the iTransformer time-series prediction model. The iTransformer-based time-series prediction model is a Transformer-based time-series prediction architecture that does not change the Transformer's network architecture but transforms the roles of the attention mechanism and the feedforward network. Different variables are considered separately, with each variable encoded as an independent token. The attention mechanism models the correlation between different variables, and the feedforward network models the temporal correlation of the variables to obtain a sequence time-series representation. Historical time-series data is input, and the algorithm outputs the number of users in each unit at future time points.

[0046] The optimization problem is specifically:

[0047] By optimizing the UAV's flight trajectory and offloading decision variables, a computational offloading constrained optimization problem is obtained, aiming to minimize the weighted sum of the total system task energy consumption and average latency, while ensuring constraints on communication resources, computational resources, LEO satellite coverage, and the maximum tolerable latency of the task. The optimization problem is as follows:

[0048]

[0049] The constraints are:

[0050]

[0051] Among them, constraint C1 means that the power of the UAV at any time cannot exceed its maximum power; constraint C2 means the UAV battery capacity limit; constraint C3 means the maximum coverage range limit of the UAV; constraint C4 means the task calculation tolerance delay; constraint C5 means that each user can only establish an uplink or downlink connection with one UAV; constraint C6 means the UAV flight speed limit; and constraint C7 means the instantaneous power consumption limit of the user and the UAV.

[0052] The defined optimization problem is modeled as a partially observable Markov decision process (POMDP). The MDP is simplified and transformed into a model-free process. The state space, action space, and reward function are defined as follows.

[0053] 1) State Space: The state space includes all network factors that influence the agent's decisions, represented as... This includes the location coordinates of UAVs, LEOs, and ground users, the workload, and energy consumption;

[0054] 2) Action Space: This includes the actions of UAV flight and the actions of establishing connections with ground users. The action space in time slot t is represented as follows:

[0055] 3) Reward Function: The goal is to minimize the objective function, which is defined as the reward function:

[0056] The multi-agent deep reinforcement learning algorithm is a two-level Qmix algorithm based on hierarchical learning; it consists of two parts: a high-level decision network and a low-level decision network.

[0057] Low-Earth orbit satellites are regarded as high-level intelligent agents in a high-level decision-making network, and drones are regarded as low-level intelligent agents in a low-level decision-making network.

[0058] The high-level decision network and the low-level decision network respectively include an agent decision network RNN ​​and a mixing network for fitting the Q values ​​of multiple agents;

[0059] The RNN network is cyclically input to the current observation o and the agent's action a from the previous time step. t-1 The mixing network takes the Q-values ​​of multiple agents and the state information of the global environment as input, and outputs the global Q-value Q. tot ;

[0060] The mixing network incorporates a multi-head attention-based encoder to reduce the dimensionality of the global environment's state information. It then extracts effective spatiotemporal features using the multi-head attention mechanism. Simultaneously, the results predicted by the temporal prediction algorithm are used as part of the reward function of the high-level decision network, guiding the agent to identify high-reward regions in the early stages of policy exploration and avoiding the impact of sparse environmental rewards. The loss function of the mixing network is defined as:

[0061]

[0062] The mixing network uses the DQN network update method, where θ represents the current network parameters and b represents the number of samples; where R is the reward function, θ - The target network is represented; the high-level decision network and the low-level decision network jointly learn the same task and share a single target network to update the network parameters.

[0063] A high-level decision-making network is deployed on a low-Earth orbit (LEO) satellite, while a low-level decision-making network is deployed within the drone. The LEO satellite, acting as the high-level agent, makes a high-level decision every K time interval based on observed state information and user distribution predicted by an iTransformer-based time-series prediction model, guiding the drone's low-level strategy for the next k steps. At each time interval, the drone determines its action based on its own observations and the reward function provided by the high-level decision-making network on the LEO satellite. The reward value obtained by the high-level decision-making network is evenly distributed to the low-level decision-making network in the next k time slots. The reward function of the low-level decision-making network is expressed as:

[0064]

[0065] Where R H The reward function representing the high-level decision-making network is defined as:

[0066]

[0067] in, This indicates the number of users covered by the drone. This represents the difference between the prediction result of the time series prediction model based on iTransformer and the actual data, where j and k represent the weights of the two.

[0068] The beneficial effects of this invention are as follows: In summary, this invention provides a comprehensive and accurate method for handling the flight trajectory and task offloading decision-making problems of unmanned aerial vehicles (UAVs) in integrated air-space-ground networks. By optimizing the flight trajectory and offloading strategy of UAVs and continuously learning and adapting to the environment, this invention enables UAVs to dynamically adjust their flight trajectory and offloading strategy based on factors such as the distribution of ground users, changes in task requirements, and network status. This method not only meets the computing needs of users in specific areas but also, through the designed algorithm, achieves a high degree of synergy between the offloading strategies of UAVs and satellites, significantly improving the overall network performance and better meeting users' needs for high energy efficiency and low latency. Attached Figure Description

[0069] Figure 1 Display the overall network model structure diagram;

[0070] Figure 2 Show the network structure based on the iTransformer time series prediction model;

[0071] Figure 3 This paper presents a design method for a two-level Qmix algorithm based on hierarchical learning.

[0072] Figure 4 This paper presents a two-level Qmix algorithm network structure based on hierarchical learning.

[0073] Figure 5 It displays the changes in the algorithm reward value under specified parameters, where HA_Qmix is ​​the method of this invention;

[0074] Figure 6 Displays energy consumption information for multiple ground users;

[0075] Figure 7 Displays latency information for multiple ground users. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0077] Incorporating a decoupled communication model between UAVs and ground users into the invention scenario, during task offloading, ground users can select different UAVs to establish uplink and downlink connections. During a task computation cycle, if the link is broken during uplink transmission, the UAV needs to forward data to a new uplink UAV, while the ground user resumes the transmission of any unsuccessfully transmitted data. After completing the task, the uplink UAV needs to transmit the computation results to the downlink UAV, which then delivers them to the user. This approach effectively overcomes the limitations of full-duplex coupled communication, reducing the long waiting time for users due to link breaks. Furthermore, the research problem is modeled as an optimization problem, aiming to minimize the weighted sum of the average latency and energy consumption for the user to complete the computation task. Establishing uplink and downlink connections between the UAV and the ground user is treated as two actions in the agent's action space, jointly considering the association between the UAV and the user and their flight trajectories to ensure the continuity of the task offloading service and high-quality data delivery.

[0078] To address the challenges of task offloading in SAGIN, this invention proposes a two-level Qmix algorithm based on hierarchical learning, using the HAVEN framework—a value decomposition framework for multi-agent fully cooperative problems—comprising a high-level decision network and a low-level decision network. The low-Earth orbit (LEO) satellite, acting as the high-level agent, makes a high-level decision every K time interval based on observed state information and user distribution predicted by the prediction module. This decision guides the UAV's low-level strategy for the next k steps. At each time step, the UAV determines its action based on its own observations and the reward function provided by the high-level strategy. Simultaneously, a time-series prediction model based on iTransformer predicts user distribution to guide algorithm training. By designing the reward function in the two-layer hierarchical structure, a dual coordination mechanism is introduced between layers (LEO satellite and UAV) and between agents (UAVs), ensuring that agents consider the actions and goals of other agents when executing tasks, thus improving agent collaboration.

[0079] Step 1: Establish a system model

[0080] Step 1.1: Network Model

[0081] This invention considers the scenario where ground users have a certain number of computationally intensive or latency-sensitive tasks that are difficult to complete independently. It utilizes an integrated air-space-ground network to provide computing services to ground users, leveraging the good maneuverability of UAVs and the high coverage of low-Earth orbit satellites to provide continuous and high-quality services, meeting the latency and energy consumption needs of ground users' computing tasks. After ground users upload their tasks to the corresponding UAVs, the UAVs, acting as the offloading decision-makers, decide whether the task is computed within the UAV itself or delivered to upper-level low-Earth orbit satellites for computation. The overall SAGIN structure consists of three layers:

[0082] Ground layer: This layer includes ground users who have computationally intensive or latency-sensitive tasks that they cannot perform on their own and need to access computing services through an integrated air-space-ground network. For example, equipment in a smart factory on the ground may need to process large amounts of production data in real time, but its own computing resources are limited. In this case, the task can be uploaded to other nodes in the network for processing.

[0083] The air layer primarily consists of unmanned aerial vehicles (UAVs). UAVs possess excellent maneuverability and play a crucial role in mission unloading. They can establish communication links with ground users, receive missions uploaded by ground users, and, based on decisions, perform mission calculations themselves or transmit them to upper-level satellites for calculation. Simultaneously, communication links exist between UAVs and between UAVs and satellites to enable information transmission and interaction.

[0084] The space layer comprises multiple low-Earth orbit (LEO) satellites. LEO satellites offer high coverage, enabling them to observe the state of a large area and possess relatively strong computing power. Within the overall network architecture, LEO satellites are considered high-level agents. At regular intervals, they make high-level decisions based on observed state information and user distribution predictions from the forecasting module, guiding the low-level strategies of UAVs. They play a crucial role in optimizing UAV flight trajectories and task offloading strategies, ensuring continuous, high-quality service for ground users and meeting their latency and energy consumption requirements for computing tasks.

[0085] Different computing nodes possess varying computing and communication capabilities, enabling more efficient task processing through collaboration. The network model considered in this invention includes N ground users, M unmanned aerial vehicles (UAVs), and L low-Earth orbit (LEO) satellites. The set of ground users is represented as... Its three-dimensional coordinates are represented by l n (t)=(x n (t),y n (t),h nThe coordinates of ground users are randomly assigned at the beginning of each time slot and then move according to a stochastic model. The set of aerial drones is represented as (t). Its three-dimensional coordinates are represented by l m (t)=(x m (t),y m (t),h m (t) represents a satellite that moves at a speed of V in each time slot and is responsible for collecting data from ground users. The set of low-Earth orbit satellites above it is represented by l∈L={1,2,...,L}, and its three-dimensional coordinates are represented as l l (t)=(x l (t),y l (t),h l (t)) LEO satellites move along fixed orbits. Due to their high speed and limited coverage, ground users do not maintain constant connectivity with LEO satellites. Communication can only be established when a LEO satellite enters a communication window.

[0086] Step 1.2: Communication Model

[0087] The communication link between the drone and the ground user is generally considered to be a line-of-sight (LoS) link or a non-line-of-sight (NLoS) link, which is closely related to the network environment. The path loss for LoS is... The path loss corresponding to NLoS is Where β LoS and β NLoS These are two different attenuation factors that distinguish between LosS and NloS links, α mn (t) is the power gain, therefore the average path loss for this communication link is defined as:

[0088]

[0089] Where P LoS (t) represents the probability that the UAV and the user establish a Loss channel in a specific network environment. Correspondingly, the probability that they establish an NLoS channel is P. NLoS (t)=1-P LoS (t).

[0090] Due to the limitations of full-duplex coupled channels on task offloading scenarios, this invention proposes to decouple the communication link between the UAV and the ground user, separating their uplink and downlink, so that each ground user can establish uplink and downlink communication with different UAVs, defining a binary variable. To indicate whether UAVm is associated with user n, it can be represented as:

[0091]

[0092] When (·) = D, it indicates that the communication link is a downlink; when (·) = U, it indicates that the link is an uplink. It is worth noting that in each time slot, each user is only allowed to establish an uplink or downlink with one UAV. At the same time, due to the limitations of UAV hardware, a UAV can only establish a connection with a maximum of Q users in each time slot.

[0093] Unlike conventional full-duplex coupled communication methods, the decoupled communication method considered in this invention may have some additional channel interference.

[0094] (1) Uplink: Interference originates from interference from other uplinks, interference from other downlinks of the UAV, and self-interference of the channel, which can be expressed by the formula:

[0095]

[0096] (2) Downlink: Similar to the uplink, the interference experienced by the user in the downlink can be represented as:

[0097]

[0098] Therefore, the information transmission rate between the drone and the user can be expressed as:

[0099]

[0100] in This is the signal-to-noise ratio plus interference. For uplink and downlink, the corresponding signal-to-noise ratios are:

[0101]

[0102]

[0103] Since communication between drones typically occurs in free and open space, communication links between drones and between drones and satellites will be established using line-of-sight transmission. Therefore, the information transmission rate is defined as follows:

[0104]

[0105] Step 1.3: Task Unloading Model

[0106] The main models for task unloading include the task unloading energy consumption model and the latency model.

[0107] Step 1.3.1: Energy Consumption Model

[0108] During initialization, each ground user is randomly assigned a computational task with a certain amount of data, denoted as I.n Ground users need to upload the task information via uplink to the UAV, which then decides whether to upload the task to the LEO satellite for computation or perform the computation itself. Throughout the entire task computation cycle, this invention primarily considers the energy consumed by ground users and the UAV, which can be divided into four aspects:

[0109] (1) Energy consumption for ground user data transmission: Users need to transmit all their data to the UAV, and the data volume is I. n The power of user-uploaded data is defined as p. n (t), therefore the transmission energy consumption can be defined as:

[0110]

[0111] (2) Data transmission power consumption of UAVs: When different UAVs establish uplink and downlink links with the same user, the uplink UAV needs to transmit data to the corresponding downlink UAV, or forward data to the LEO satellite, and its transmission power is defined as p m (t), energy consumption is:

[0112]

[0113] If the calculation is performed by this UAV, the energy consumption required to return the result to the ground user after the task is completed is:

[0114]

[0115] (3) Energy consumption of UAV computing tasks: Set the amount of data that the UAV needs to compute as TA. m The calculation speed is The calculated power is P c The energy consumption of a computing task can be defined as:

[0116]

[0117] (4) UAV flight energy consumption: The flight energy consumption of UAV is relatively large compared to hovering energy consumption. Therefore, this invention mainly considers the flight energy consumption of UAV. The flight power of UAV is defined as P. f Flight energy consumption can be defined as:

[0118]

[0119] Step 1.3.2: Delay Model

[0120] The latency component mainly comprises two parts: task computation latency and information transmission latency. Task computation latency involves both UAV computation and satellite computation latency, which are defined as follows: Information transmission latency is divided into uplink latency, downlink latency, and transmission latency between UAVs and between UAVs and satellites, defined as:

[0121]

[0122] Therefore, the cycle delay for each task to complete is defined as:

[0123]

[0124] Step 2: Problem Definition

[0125] For each mathematical model, by optimizing the UAV's flight trajectory and unloading decision variables, a computational unloading constraint optimization problem is obtained, aiming to minimize the weighted sum of the total system task energy consumption and average latency, while ensuring constraints on communication resources, computational resources, LEO satellite coverage, and the maximum tolerable latency of the task. The optimization problem is given as follows:

[0126]

[0127] The constraints are:

[0128]

[0129]

[0130] Among them, constraint C1 means that the power of the UAV at any time cannot exceed its maximum power; constraint C2 means the UAV battery capacity limit; constraint C3 means the maximum coverage range limit of the UAV; constraint C4 means the task calculation tolerance delay; constraint C5 means that each user can only establish an uplink or downlink connection with one UAV; constraint C6 means the UAV flight speed limit; and constraint C7 means the instantaneous power consumption limit of the user and the UAV.

[0131] Step 3: Time Series Prediction Model Based on iTransformer

[0132] The ground area is divided into multiple small blocks, and the number of users is simulated using a real dataset. This data is then used as time-series data, and a time-series prediction model based on iTransformer is employed for prediction. The results are used as one of the parameters of the reward function in subsequent algorithms to accelerate their convergence.

[0133] This iTransformer-based time series prediction model is a novel Transformer-based time series prediction architecture. It does not change the Transformer network architecture, but rather transforms the roles of the attention mechanism and the feedforward network. iTransformer considers different variables separately, encoding each variable as an independent token. It uses the attention mechanism to model the correlation between different variables, while utilizing the feedforward network to model the temporal correlation of the variables, thus obtaining a better sequence temporal representation. The iTransformer-based time series prediction model includes an embedding module, a multi-head attention mechanism, a normalization layer, a feedforward neural network, and a mapping module.

[0134] (1) For historical window data The goal is to obtain pre-data.

[0135] (2) First, obtain the sequence encoding of the univariate: using a simple MLP layer, Mapped to

[0136] (3) Analyze different variables using the multi-head attention mechanism Correlation between them, such as Figure 2 As shown, Q, K, and V are all derived from variable H. :n Starting with the multivariate correlation map, which represents the calculated attention coefficients indicating the magnitude of the correlation between different variables, this module integrates the correlations between different variables to update the variable embeddings.

[0137] (4) Normalization layer: In order to ensure the independence of variables and avoid the influence of sampling misalignment or time delay on the cross-correlation of modeling variables, each variable is standardized separately.

[0138]

[0139] (5) Feedforward Neural Network: Finally, the feedforward neural network is used to further model the temporal effects within each variable to obtain an efficient temporal representation.

[0140] The process of the time series prediction model algorithm based on iTransformer is as follows:

[0141]

[0142] Step 4: Design of a two-level Qmix algorithm based on hierarchical learning

[0143] The defined optimization problem is modeled as a partially observable Markov decision process (POMDP). Most MARL algorithms for POMDP currently follow the centralized training distributed execution (CTDE) framework, where each agent can utilize all available information during training, but can only make decisions based on local observations.

[0144] Typically, a POMDP is described as a 4-tuple, including the state space, state transition probabilities, action space, and reward function. In this context, since the agent's state is high-dimensional and continuous, and it is difficult to accurately obtain the state transition probabilities, the MDP is simplified and transformed into a model-free process. To apply deep reinforcement learning (DRL) to trajectory optimization and task unloading problems, this invention models the problem as an MDP. In this model, the state space, action space, and reward function are defined as follows.

[0145] State space: The state space includes all network factors that influence the agent's decisions, represented as... This includes the location coordinates of UAVs, LEOs, and ground users, the workload, and energy consumption.

[0146] Action Space: The action space mainly consists of two parts: the actions of UAV flight and the actions of establishing connections with ground users. Therefore, the action space in time slot t can be represented as follows:

[0147] Reward function: The objective function is minimized, therefore it is defined as the reward function:

[0148] As the complexity of problems and scenarios increases, situations where multiple heterogeneous intelligent agents cooperate to complete the same task frequently occur. Existing research algorithms lack cooperation relationships between agent layers and only consider cooperation relationships within layers. Inspired by the human nervous system, when completing a complex task, close cooperation between specific areas of the brain and lower-level organs is indispensable. The brain, as a high-level strategy, needs to guide the lower-level strategies of each agent. Therefore, when the human body is regarded as a multi-agent system, coordination between layers and within layers is crucial for solving tasks requiring complete cooperation.

[0149] Based on the above, in the SAGIN scenario, low-orbit satellites are considered high-level agents due to their large coverage area, ability to observe the global state, and strong computing power. UAVs, on the other hand, have limited observation range and relatively weak computing power, but the specific strategy implementation needs to be completed by them, so they can be considered low-level agents. Based on the HAVEN value decomposition framework for multi-agent fully cooperative problems using hierarchical reinforcement learning, a two-level hierarchical Qmix algorithm is proposed, which includes a high-level decision network and a low-level decision network.

[0150] The network structures of the high-level and low-level decision networks are similar to the traditional value decomposition algorithm Qmix. The agent's Actor network uses an RNN network, taking the current observation and the hidden state from the previous time step as the current input to extract the hidden information of the time sequence. The mix network is equivalent to the Critic network, simultaneously receiving the Q-value of the Agent RNN Network and the current global state st, and outputting the behavioral utility value Q of the joint action u of all agents in the current state. tot .

[0151] By designing a time-series prediction model based on iTransformer to predict user distribution at future moments, the agent can identify high-reward areas early in policy exploration, avoiding the impact of sparse environmental rewards. A high-level decision network and the iTransformer-based time-series prediction model are deployed on a low-Earth orbit (LEO) satellite, while the lower-level decision network is deployed in a drone. The LEO satellite, acting as the high-level agent, makes a high-level decision every K time interval based on observed state information and the user distribution predicted by the prediction module, guiding the drone's low-level strategy for the next k steps. At each moment, the drone determines its action based on its own observation information and the reward function provided by the high-level strategy. By designing the reward function in the two-layer hierarchical structure, a dual coordination mechanism is introduced between layers (satellite and drone) and between agents (drones), ensuring that the agent considers the actions and goals of other agents when executing tasks, thus improving agent collaboration.

[0152] Meanwhile, when fitting the global q-value in the mix network, pre-encoding the global state using a Transformer encoder allows for the extraction of features containing spatial correlations between different regions' states at the current time. This also reduces the dimensionality, leading to faster algorithm convergence. The overall training process of the algorithm is shown below:

[0153]

[0154]

[0155] Step 5: Performance Verification

[0156] The simulation implementation of this method is based on the PyCharm platform. Python is used to simulate the scenario model. The experiment randomly deployed 5 UAVs, 20 ground users, and 3 LEOs in a 1000m x 1000m square area, assuming the UAVs hovered at a fixed altitude of 50 meters. The task data size for ground users was 8MB, and the maximum tolerable latency for ground users was 1000ms. The altitude of the ground above the LEO satellites was 784km. The channel power gain at a 1-meter reference distance was -80dB, the path loss was 8, the available channel bandwidth for UAVs was 300MHz, the available channel bandwidth for LEOs was 40GHz, the background noise for ground user communication was 40GHz, the computing power of UAVs was assumed to be 4G / s, the computing power of LEOs was assumed to be 10G / s, the transmission power of UAVs was 1W, and the transmission power of LEOs was 3W. A comparison algorithm was also set up during the simulation verification of this method. To evaluate three performance metrics—average response latency, average response power consumption, and Quality of Service (QoS)—comparative experiments were conducted on the PyCharm platform for DQN, MADDPG, and Qmix. Figure 5 As shown, this method outperforms MADDPG and Qmix algorithms in terms of reward value during training, while DQN performs the worst. Figure 6 This indicates that under the above parameter conditions, 10 users were randomly selected from ground users to analyze the task computation energy consumption of different algorithms. It was found that the task energy consumption of this method is the lowest, Qmix and MADDPG performed relatively poorly, and DQN had the highest energy consumption. Figure 7 This means that under the above parameter conditions, 10 users were randomly selected from ground users to analyze the task computation latency of different algorithms. This method can meet the specified maximum tolerable latency of the task, while some tasks of DQN exceed the specified maximum tolerable latency of the task.

Claims

1. A method for UAV trajectory optimization and task offloading for an integrated air-space-ground network, characterized in that, Establish network models, node communication models, task offloading energy consumption models, and latency models; The optimization problem is to minimize the energy consumption and average latency of completing the task. The optimization problem is solved by using a time-series prediction algorithm and a multi-agent deep reinforcement learning algorithm to obtain the optimal trajectory optimization and task unloading strategy. During the initialization of the task offloading energy consumption model, each ground user is randomly assigned a certain amount of data for computation. The ground user needs to upload the task and transmit it to the UAV via the uplink. The UAV then decides whether to upload the task to the LEO satellite for computation or to perform the computation within itself. Throughout the entire task computation cycle, the energy consumed by the ground user and the UAV is mainly considered, specifically in four aspects: (1) Energy consumption for ground user data transmission: Users need to transmit all their data to the UAV, and the data volume is [data volume not specified]. The power of user-uploaded data is defined as Therefore, transmission energy consumption is defined as: ; in Indicates the uplink information transmission rate; (2) Data transmission power consumption of UAVs: When the UAVs that establish uplink and downlink links with the same user are different, when the uplink UAV transmits data to the corresponding downlink UAV or forwards data to the LEO satellite, its transmission power is defined as Energy consumption is: ; The energy consumption calculated by the UAV and returned to the ground user upon completion of the task calculation is as follows: ; This indicates the amount of data that drone m needs to transmit. Indicates the information transmission rate of the downlink. This indicates the information transmission rate between drones or between a drone and a satellite. (3) Energy consumption of UAV computing tasks: The amount of data that the UAV needs to compute is set to be... The calculation speed is The calculated power is The energy consumption of a computing task is defined as: ; (4) UAV flight energy consumption: The flight energy consumption of UAV is relatively large compared to hovering energy consumption. Therefore, the flight energy consumption of UAV is mainly considered. The flight power of UAV is defined as Flight energy consumption is defined as: ; The latency model mainly includes two parts: task computation latency and information transmission latency. The task computation latency involves UAV computation latency and satellite computation latency, which are defined as follows: and Information transmission latency is divided into uplink latency, downlink latency, and transmission latency between drones and between drones and satellites, defined as follows: ; ; ; ; in, Indicates the amount of tasks uploaded by ground users. Indicates the amount of data transmitted by the drone. Indicates the amount of data transmitted between drones. This indicates the amount of data transmitted by the drone to the low-Earth orbit satellite; Therefore, the cycle delay for each task to complete is defined as: ; The optimization problem is specifically: By optimizing the UAV's flight trajectory and offloading decision variables, a computational offloading constrained optimization problem is obtained, aiming to minimize the weighted sum of the total system task energy consumption and average latency, while ensuring constraints on communication resources, computational resources, LEO satellite coverage, and the maximum tolerable latency of the task. The optimization problem is as follows: ; The constraints are: C1: ; C2: ; C3: ; C4: ; C5: ; C6: ; C7: ; Among them, constraint C1 means that the power of the UAV at any time cannot exceed its maximum power; constraint C2 means the UAV battery capacity limit; constraint C3 means the maximum coverage range limit of the UAV; constraint C4 means the task calculation tolerance delay; constraint C5 means that each user can only establish an uplink or downlink connection with one UAV; constraint C6 means the UAV flight speed limit; and constraint C7 means the instantaneous power consumption limit of the user and the UAV. The defined optimization problem is modeled as a partially observable Markov decision process (POMDP). The MDP is simplified and transformed into a model-free process. The state space, action space, and reward function are defined as follows. 1) State Space: The state space includes all network factors that influence the agent's decisions, represented as... This includes the location coordinates of UAVs, LEOs, and ground users, the workload, and energy consumption; 2) Action Space: This includes the actions of UAV flight and the actions of establishing connections with ground users. The action space in time slot t is represented as follows: ; 3) Reward Function: The goal is to minimize the objective function, which is defined as the reward function: .

2. The method for UAV trajectory optimization and task offloading for an integrated air-space-ground network according to claim 1, characterized in that, The network model is divided into three layers: the ground layer, the air layer, and the space layer. The ground floor includes ground users; The air layer is mainly composed of unmanned aerial vehicles (UAVs); UAVs establish communication links with ground users, receive tasks uploaded by ground users, and calculate the tasks themselves or transmit them to upper-level satellites for calculation based on decisions; at the same time, there are communication links between UAVs and with satellites to realize information transmission and interaction. The space layer includes multiple low-Earth orbit (LEO) satellites above it; Deploy iTransformer-based time series forecasting models within it; The low-orbit satellite is regarded as a high-level intelligent agent. It makes high-level decisions at regular intervals based on the observed state information and the user distribution predicted by the iTransformer time series prediction model, which guides the low-level strategy of the UAV. The network model includes N ground users, M drones, and L low-Earth orbit satellites. The set of ground users is represented as... Its three-dimensional coordinates are used Let represent the situation where the coordinates of ground users are randomly assigned at the beginning of each round, and then randomly moved at the beginning of each time slot. The set of aerial drones is represented as . Its three-dimensional coordinates are used This indicates that it moves at a speed of V in each time slot, responsible for collecting data from ground users, and the low-Earth orbit satellites above it are used for... To represent, its three-dimensional coordinates are as follows LEO moves along a fixed orbit; ground users can only establish communication with LEO satellites when they enter a communication window.

3. The method for UAV trajectory optimization and task offloading for an integrated air-space-ground network according to claim 1, characterized in that, In the node communication model, the communication link between the UAV and the ground user is either a line-of-sight (LoS) link or a non-line-of-sight (NLoS) link; the path loss of LoS is... The corresponding path loss for NLoS is ,in and These are two different attenuation factors that distinguish between LosS and NloS links. It is the power gain, and the average path loss for this communication link is defined as: ; in, Let be the probability that a UAV and a user establish a Loss channel in a specific network environment, and let be the probability that they establish an NLoS channel. ; The communication link between the drone and the ground user is decoupled, separating their uplink and downlink, so that each ground user can establish uplink and downlink communication with different drones. A binary variable is defined. To indicate whether UAVm is associated with user n, it is represented as: ; when This indicates that the communication link is a downlink. The time indicates that the link is an uplink. In each time slot, each user is only allowed to establish an uplink or downlink with one UAV. At the same time, due to the limitations of UAV hardware, a UAV can only establish a maximum of Q ground users in each time slot. Considering channel interference during decoupling, communication links are established between UAVs and between UAVs and satellites using line-of-sight transmission. Therefore, the information transmission rates between UAVs and between UAVs and satellites are defined as follows: ; in For backhaul link bandwidth, For noise power, This refers to the transmission power. The information transmission rate between the drone and the ground user is expressed as: ; The link connection status between drone m and user n at time t. For channel bandwidth, This is the signal-to-noise ratio plus interference. For uplink and downlink, the corresponding signal-to-noise ratios are: ; ; ; ; in, Indicates the transmit power of user n, Indicates the transmit power of drone m, This indicates the interference caused by other users' uplink transmissions to drone m. This indicates the interference caused by downlink transmissions from other drones to drone m. Indicates the self-interference of drone m. This indicates the interference caused to user n by the downlink transmissions of other drones. This indicates the interference caused to user n by the uplink transmissions of other users. This represents the self-interference of user n.

4. The method for UAV trajectory optimization and task offloading for an integrated air-space-ground network according to claim 1, characterized in that, The time series prediction algorithm divides the ground area into multiple small blocks, takes each small block as a unit, uses a real dataset to simulate the number of ground users in the unit, uses it as historical time series data, and makes predictions through the iTransformer time series prediction model. The iTransformer-based time series prediction model is a Transformer-based time series prediction architecture that does not change the Transformer network architecture but transforms the roles of the attention mechanism and the feedforward network. It considers different variables separately, with each variable encoded as an independent token. It uses the attention mechanism to model the correlation between different variables and the feedforward network to model the temporal correlation of variables, thereby obtaining the temporal representation of the sequence. It takes historical time series data as input and outputs the number of users in each unit at future time.

5. The method for UAV trajectory optimization and task offloading for an integrated air-space-ground network according to claim 1, characterized in that, The multi-agent deep reinforcement learning algorithm is a two-level Qmix algorithm based on hierarchical learning; it consists of two parts: a high-level decision network and a low-level decision network. Low-Earth orbit satellites are regarded as high-level intelligent agents in a high-level decision-making network, and drones are regarded as low-level intelligent agents in a low-level decision-making network. The high-level decision network and the low-level decision network respectively include an agent decision network RNN ​​and a mixing network for fitting the Q values ​​of multiple agents; The RNN network repeatedly inputs the current observation. and the actions of the agent in the previous moment. Output the Q value; The mixing network takes the Q-values ​​of multiple agents and the state information of the global environment as inputs, and outputs the global Q-value. ; The mixing network incorporates a multi-head attention-based encoder to reduce the dimensionality of the global environment's state information. It then extracts effective spatiotemporal features using the multi-head attention mechanism. Simultaneously, the results predicted by the temporal prediction algorithm are used as part of the reward function of the high-level decision network, guiding the agent to identify high-reward regions in the early stages of policy exploration and avoiding the impact of sparse environmental rewards. The loss function of the mixing network is defined as: ; The mixing network uses the DQN network update method. Indicates the current network parameters. Indicates the number of samples; where , For the reward function, The target network is represented; the high-level decision network and the low-level decision network jointly learn the same task and share a single target network to update the network parameters. A high-level decision-making network is deployed on a low-Earth orbit (LEO) satellite, while a low-level decision-making network is deployed within the drone. The LEO satellite, acting as the high-level agent, makes a high-level decision every K time interval based on observed state information and user distribution predicted by an iTransformer-based time-series prediction model, guiding the drone's low-level strategy for the next k steps. At each time interval, the drone determines its action based on its own observations and the reward function provided by the high-level decision-making network on the LEO satellite. The reward value obtained by the high-level decision-making network is evenly distributed to the low-level decision-making network in the next k time slots. The reward function of the low-level decision-making network is expressed as: ; in The reward function representing the high-level decision-making network is defined as: ; in, This indicates the number of users covered by the drone. This represents the difference between the prediction results of the time series forecasting model based on iTransformer and the actual data. , This indicates the weights of the two.

Citation Information

Patent Citations

  • Space-air-ground integrated mobile edge computing unloading and resource allocation optimization method

    CN118524446A

  • Flow sensing lightweight layered unloading framework for adaptive slice enabled space-air-ground integrated network

    CN118784599A