Safe communication method for task offloading of space-air-ground integrated network
By optimizing UAV mode selection, user association, and power allocation through hierarchical reinforcement learning, conflict search, and successive convex approximation algorithms, the challenges of resource allocation and task scheduling in integrated air-space-ground networks are solved, achieving low-latency and high-security communication.
Patent Information
- Application Number
- CN202511018040.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-18
AI Technical Summary
In an integrated air-space-ground network, how can we optimize resource allocation and mission scheduling in collaboration between UAVs and low-Earth orbit satellites, prevent mission information from being eavesdropped on, and control the directionality of interference signals to balance privacy protection and communication performance?
A hierarchical reinforcement learning algorithm is used to optimize drone mode selection and user association, a conflict search algorithm is used to plan drone trajectories, and a successive convex approximation algorithm is used to optimize power allocation. A system model is constructed to minimize the average system latency and eavesdropping rate.
It effectively reduced task delays and suppressed the transmission rate at the eavesdropper's location, improving communication security and resource utilization efficiency.
Smart Images

Figure CN120979513A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of communication, in particular to a secure communication method for space-air-ground integrated network task offloading, aiming at optimizing the utilization rate of computing resources and ensuring communication security, and designs a hierarchical reinforcement learning algorithm which can simultaneously optimize the unmanned aerial vehicle mode selection and user association decision. BACKGROUND
[0002] In recent years, satellite and unmanned aerial vehicle technologies have received extensive attention as key complements to ground networks. Benefiting from their flexible deployment and efficient management, unmanned aerial vehicles are increasingly widely used in military and civilian fields, providing on-demand communication and computing resource support for various scenarios. However, due to the broadcast nature of wireless signals, users have the risk of being intercepted when offloading tasks, and some physical layer security techniques have been developed to address this challenge, such as unmanned aerial vehicle cooperative relaying, friendly unmanned aerial vehicle jamming, etc. Unmanned aerial vehicles and low-orbit satellites complement each other, combining wide coverage with flexibility and controllability. Due to the high mobility and controllability of unmanned aerial vehicles, they can dynamically adjust their positions to meet network demands and cooperate with friendly unmanned aerial vehicles to perform jamming and computing offloading tasks. This cooperative mechanism effectively alleviates the eavesdropping threat during offloading, thereby providing users with low-latency, wide-coverage, and enhanced privacy computing services. However, there are still two key challenges to achieving these goals. First, both satellites and unmanned aerial vehicles exhibit high mobility, resulting in a highly dynamic network that poses significant challenges to resource allocation, task scheduling, and maintaining stable connections, necessitating the development of adaptive network management and optimization mechanisms. Second, while friendly unmanned aerial vehicles can transmit jamming signals to help users resist potential eavesdroppers, these signals may inadvertently interfere with legitimate users, thereby degrading their communication quality. Therefore, how to control the directionality and scheduling of jamming signals to balance privacy protection and communication performance is a key issue. SUMMARY
[0003] In view of some deficiencies of the prior art, the application proposes an air-space-ground network security task offloading method, which uses hierarchical reinforcement learning, conflict search, and successive convex approximation to optimize unmanned aerial vehicle mode selection and user association, unmanned aerial vehicle trajectory and power allocation, respectively.
[0004] In view of this, the technical scheme adopted by the application is: a secure communication method for space-air-ground integrated network task offloading, comprising the following steps:
[0005] A system model is constructed, a movement model, a communication model and a time delay model are determined, and an optimization problem aiming at minimizing system average time delay and eavesdropping rate is constructed; the system model comprises low-orbit satellites, a plurality of unmanned aerial vehicles, a plurality of ground users and eavesdroppers, the computing task of the ground user is handed over to the unmanned aerial vehicle or the satellite for processing, and the unmanned aerial vehicle transmits an interference signal to prevent the user task information from being eavesdropped by the eavesdropper;
[0006] The layered reinforcement learning algorithm is used to solve the unmanned aerial vehicle mode selection and user association sub-problems in the optimization problem;
[0007] The conflict search algorithm is used to solve the unmanned aerial vehicle trajectory planning sub-problem in the optimization problem;
[0008] The successive convex approximation algorithm is used to solve the power allocation sub-problem in the optimization problem, and the entire optimization problem is solved by using the alternating optimization algorithm according to the solutions of the three sub-problems.
[0009] The optimization problem aiming at minimizing system average time delay and eavesdropping rate is represented as:
[0010]
[0011] The target of the present application is to minimize the task time delay and the transmission rate of the user at the eavesdropper by jointly optimizing the power control Subchannel allocation Unmanned aerial vehicle trajectory And unmanned aerial vehicle mode selection I. m [n] represents the number of tasks generated by user m in time slot n, t m,i [n] represents the total delay of task i of user m, R m,e [n] represents the transmission rate of user m at eavesdropper e, δ u [n] represents a binary variable of the working mode of the unmanned aerial vehicle in time slot n, the unmanned aerial vehicle can select to transmit an interference signal or act as an edge server, if the unmanned aerial vehicle selects to transmit an interference signal δ u [n] = 1, otherwise δ u [n] = 0, represents whether user m communicates with the unmanned aerial vehicle through subchannel k, represents whether user m communicates with satellite l through subchannel o, represents the transmission power of the user in time slot n, represents the transmission power of the unmanned aerial vehicle in time slot n, respectively represent the maximum transmission power of the user and the unmanned aerial vehicle. λ m represents the task arrival rate of user m, b U represents the service rate of the unmanned aerial vehicle, F l [n] represents the number of tasks that can be processed by satellite l in time slot n, qu [n], q i [n] represents the coordinates of the drone u, i in time slot n, D min represents the minimum distance between drones, D max represents the maximum flight distance of the drone in a time slot, t m,i [n] represents the total delay of task i of user m, represents the maximum tolerable delay of the task.
[0012] Constraint C1 is the drone mode selection variable, the drone can only work in one mode in a period, constraint C2 and C3 ensure that the user can only access one subchannel at any time, constraint C4 ensures that the drone cannot provide services to the user in the interference mode, constraint C5 is the transmission power limit of the user and the drone, constraint C6 ensures the stability of the drone processing queue, constraint C7 is the task quantity limit that the satellite processing queue can handle, constraint C8 and C9 are the drone trajectory constraints, and constraint C10 is the maximum delay constraint.
[0013] The application also provides a secure communication system for space-air-ground integrated network task offloading, which can perform the above-mentioned secure communication method for space-air-ground integrated network task offloading, comprising: a low-orbit satellite, a plurality of drones, a plurality of ground users and an eavesdropper, the computing task of the ground user is handed over to the drone or the satellite for processing, and the drone transmits an interference signal to prevent the user task information from being eavesdropped by the eavesdropper.
[0014] The beneficial effects of the application include:
[0015] The application considers that the service time of the satellite is limited, and the computing task processing process of the satellite is modeled as a finite capacity queuing system to ensure efficient use of computing resources. When the drone interferes with the eavesdropper, it inevitably affects the legitimate user, so the application aims to minimize the task delay and suppress the transmission rate of the user at the eavesdropper. The application divides the formulated optimization problem into three subproblems. First, the hierarchical reinforcement learning algorithm processes the user subcarrier allocation and the drone mode selection. Second, the application introduces a conflict search-based algorithm to approximate the optimal drone trajectory. Finally, the application designs a successive convex approximation-based algorithm to optimize the power allocation. Experimental results prove the effectiveness of the application in terms of average task delay and secrecy rate. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 System model for secure computing offloading in space-air-ground networks
[0017] Figure 2 Geometry between the earth and the satellite
[0018] Figure 3This is a schematic diagram of a hierarchical reinforcement learning algorithm;
[0019] Figure 4 The HRLCS algorithm designed for this invention compares the interference signal power performance of the other three algorithms under different user conditions.
[0020] Figure 5 The algorithm of this invention provides the system average latency performance under different user conditions;
[0021] Figure 6 The algorithm of the invention is used to measure the average eavesdropping rate performance of the system under different user conditions. Detailed Implementation
[0022] To make the technical solution and advantages of the present invention clearer, the counting scheme of the present invention will be described in detail below.
[0023] This invention provides a secure communication scheme for offloading tasks in an integrated space-ground network, the method comprising:
[0024] Step 1: Construct a system model, determine the mobility model, communication model, and time delay model, and plan a multi-objective optimization problem.
[0025] This invention constructs a such Figure 1 The system model shown assumes that, in the considered scenario, the computational tasks of ground users are handled by drones or satellites. Simultaneously, there may be eavesdroppers on the ground, and the drones will emit jamming signals to prevent the eavesdropping of user task information. Consider a drone flight cycle with a duration of T, which is discretized into N time slots of equal duration, with a time slot length τ = T / N. The set of time slots is represented by... Representation. M ground users use a set. In time slot n, the coordinates of user m are denoted as q. m [n] = (x m [n],y m [n], 0). U drones are used in a set. It is stated that the drone operates at a constant altitude H U Flight, the coordinates of the UAV u in time slot n are denoted as q u [n] = (x u [n],y u [n],H U L satellites are used in a set The coordinates of satellite l in time slot n are denoted as q. l [n] = (x l [n],y l [n],H l ), H lis the orbit height of satellite l. The e eavesdroppers use the set ε = {1,..., e,..., E}, and the coordinates of the eavesdroppers are q e [n] = (x e [n], y e [n], 0). Assume that the computation task generated by each ground user follows a Poisson distribution with arrival rate λ m . For task i of user m, it can be represented by the tuple , where c m,i is the number of CPU cycles required for task i, s m,i is the size of task i, is the maximum tolerable delay of task i, and I m is the number of tasks generated by user m in time slot n.
[0026] The movement of users and eavesdroppers follows a random walk model, and the position of UAVs can be considered as static within a time slot if the time slot length is small enough. Define the flight speed of UAVs as V U , then the maximum flight distance of UAVs within a time slot is D max = V U τ, and the movement constraint of UAVs can be represented as:
[0027]
[0028] To prevent collisions between UAVs, define the minimum distance between UAVs as D min , which is limited by:
[0029]
[0030] Due to the high-speed movement of low-orbit satellites, the service time of a single low-orbit satellite for ground users is limited, as shown in Figure 2 , where θ G is a predefined minimum elevation angle of the satellite coverage area, and when the elevation angle between the satellite and the ground user is lower than this predefined threshold, the satellite will no longer provide service to the ground user. In the formula, D R is the radius of the Earth, and θ C is the satellite coverage angle, which can be represented as θ C = arccos(D R cos θ G / (D R + H l ))- θ G , and then the arc length of satellite l within the coverage time can be calculated as Given the orbit speed V l of satellite l, the longest service time of satellite l can be calculated as
[0031] The present application adopts a probabilistic line-of-sight channel model to determine the large-scale attenuation of the link between the UAV and the ground users. The geometric outage probability between the UAV and the ground users depends on the environment-related parameters and the elevation angle of the UAV. The line-of-sight probability of user m to UAV u at time slot n is denoted as where a, b are environment parameters, Therefore, the expected channel power gain is:
[0032]
[0033] where is the regularized LoS probability, κ represents the regularization factor considering the attenuation effect of NLoS channel with κ < 1. β o denotes the channel gain at the reference distance equal to 1 m, is the path loss exponent. The channel from the UAV to the eavesdropper is similar to the above.
[0034] Since the satellite is high above the ground, the channel between the satellite and the ground users is dominated by the LOS, thus the channel between the satellite and the ground users adopts the free space path loss model, given the satellite orbit height H l , then the channel from user m to satellite l at time slot n is
[0035]
[0036] Since the UAVs are flying in the air, the channel between the UAVs is also dominated by the LOS, thus similar to the above, the channel between the UAVs is is the distance between the two UAVs. The channel from user m to eavesdropper e is where d m,e [n] is the distance between the eavesdropper and the user, ξ e is the Rayleigh fading coefficient of the unit mean exponential distribution.
[0037] Assume that the total bandwidth of the UAV is divided into K completely orthogonal sub-channels, and the sub-channel bandwidth is B U , let denote the set of sub-channels. The binary sub-channel allocation variable denotes whether user m communicates with the UAV through sub-channel k, if user m communicates with the UAV through sub-channel k otherwise δ u [n] is a binary variable indicating the working mode of the UAV at time slot n, the UAV can choose to send an interference signal or act as an edge server, if the UAV chooses to send an interference signal δ u [n] = 1, otherwise δ u [n] = 0. In time slot n, the signal-to-noise ratio of user m to UAV u on sub-channel k can be represented as:
[0038]
[0039] respectively, is the transmit power of user m and UAV u at time slot n, σ 2 is the channel noise, is the subchannel allocation variable of user m, each subchannel has bandwidth B U is the transmission rate of user m to UAV u
[0040] Assume that the total bandwidth of satellite l is divided into O fully orthogonal subchannels, each subchannel has bandwidth B L Let denote the set of subchannels. The binary subchannel allocation variable denotes whether user m communicates with satellite l through subchannel o, if user m communicates with satellite l through subchannel o otherwise Since the antenna direction of UAV u transmitting interference signal is opposite to the direction of satellite l and the body of UAV u blocks the interference signal, the impact of interference signal transmitted by UAV u on satellite l is ignored. Then the signal-to-noise ratio of user m to satellite l in subchannel o can be expressed as:
[0041]
[0042] is the transmission rate of user m to satellite l Assume that user m can only occupy one subchannel at time slot n, i.e. is the transmission rate of user m The transmission rate of user m needs to meet the minimum transmission rate requirement i.e. R m [n] > R, m.
[0043] User m can be associated with UAV or satellite, if user m is associated with UAV u, then the signal-to-noise ratio of user m to eavesdropper e in subchannel k is:
[0044]
[0045] is the transmission rate of user m
[0046] If user m is associated with satellite l, then the signal-to-noise ratio of user m to eavesdropper e in subchannel o is:
[0047]
[0048] is the transmission rate of user m In summary, the transmission rate of user m at eavesdropper e is
[0049] The time spent by user m to transmit task i to the offloading terminal in time slot n is The time spent by user m to transmit task i to the offloading terminal in time slot n is The time spent by user m to transmit task i to the offloading terminal in time slot n is b U , b L are the service rates of the UAV and the satellite, respectively.
[0050] The waiting time of task i of user m in the satellite l queue can be derived from queuing theory. In addition, to ensure that the satellite does not lose connection in a time slot, when the remaining service time of the satellite is less than the length of a time slot, the satellite is no longer available as an offloading terminal.
[0051] The computing task generation of each ground user follows a Poisson distribution with arrival rate m According to the merging property of the Poisson process, the UAV computing task arrival rate is The processing queue of the UAV in time slot n can be expressed as a queuing model of M / M / 1, and the waiting time of task i of user m in the UAV u queue is which can be derived from queuing theory. In summary, the total delay of task i of user m can be calculated by the following formula:
[0052]
[0053] The optimization goal of the present application is to minimize the system average delay and eavesdropping rate, and the problem is described as follows:
[0054]
[0055] Constraint C1 is the mode selection variable of the UAV, which can only work in one mode in a time period. Constraints C2 and C3 ensure that the user can only access one subchannel at any time, constraint C4 ensures that the UAV cannot provide services to the user in the interference mode. Constraint C5 is the transmission power limit of the user and the UAV. Constraint C6 ensures the stability of the UAV processing queue, and constraint C7 is the limit of the number of tasks that the satellite processing queue can handle. Constraints C8 and C9 are the trajectory constraints of the UAV. Constraint C10 is the maximum delay constraint.
[0056] Step 2: Solve the UAV mode selection and user association sub-problems in step 1) by using a hierarchical reinforcement learning algorithm.
[0057] The constraints C1, C2, C3 and C4 in the optimization problem in step 1) are all constraints that limit the selection of the UAV mode and the user association, and it can be seen that the selection of the UAV mode and the user association are highly coupled, and if they are optimized separately, the user association needs to be fixed in advance. However, the UAV in the interference mode cannot serve the user, and the user association depends on the mode selection, so the two must be optimized jointly to achieve optimal performance. Therefore, the layered reinforcement learning algorithm is adopted to optimize the selection of the UAV mode and the user association, and the layers interact with each other to obtain a better solution.
[0058] Since the variables a and d are both binary and coupled with each other, the problem P1 is still a typical mixed integer nonlinear programming problem, which is difficult to solve effectively by traditional algorithms. Therefore, a layered reinforcement learning framework is proposed to divide the UAV mode selection and user association decision into two levels, thereby reducing the overall computational complexity. Specifically, the upper controller is responsible for the selection of the UAV mode, and the lower controller performs the user association decision according to the mode determined by the upper controller. As shown in Figure 3 During the training process, the upper controller outputs the selection of the UAV mode every n time slots, and the lower controller performs the user association selection according to the output result of the upper controller.
[0059] In the proposed architecture, the first type of agent is deployed at the high level to control the selection of the UAV mode, and the second type of agent is deployed at the low level to perform the specific user association task. The environment state of the high-level agent includes the positions of the UAVs, the eavesdroppers and the number of eavesdroppers in the nth time slot, denoted as The corresponding action space represents the selection strategy of the operation mode of each UAV in each time slot, denoted as
[0060] The reward function of the high-level controller aims to guide it to select the appropriate operation mode to minimize the eavesdropping rate and reduce the interference to the legal communication. Therefore, the reward function is defined as the inverse of the sum of the average communication rate of the user and the eavesdropping rate, as follows:
[0061]
[0062] The low-level agent corresponds to all users. Its environment state includes the positions of the users, the UAVs and the satellites, and the current selection strategy of the UAV mode, defined as Considering that the low-level agent can only obtain local information related to itself, its observation space is defined as the channel state between each user and the accessible nodes (including the UAVs, the satellites and the potential eavesdroppers), that is, O m,n ={h m,u [n],h m,l [n],h m,e [n]}.
[0063] The action space of the low-level agent refers to the task offloading target selected by user m at time slot n, denoted as Since the impact of the task offloading strategy on the eavesdropping rate is relatively small, the reward function of the low-level agent is mainly designed based on the task offloading delay, which is as follows:
[0064]
[0065] The high-level architecture contains three types of neural networks: policy network, critic network, and target critic network. Among them, the policy network maps the environment state to the probability distribution of each available action, and samples the action according to the distribution. The target critic network receives the environment state and the corresponding action probability generated by the policy network as input, and outputs two Q values for evaluating the value of the action. By introducing two Q values, the overestimation problem often occurring in the training process of traditional single Q network is effectively alleviated. The Q network can be trained through Bellman residual, and its loss function is which is expressed as follows:
[0066]
[0067] wherein can be obtained by soft Bellman iteration, is the parameter of the target Q network, and θ is the parameter of the current Q network, s n , a n represent the corresponding state and action.
[0068] In the low-level framework, each agent is equipped with a duel network and a corresponding target duel module. The Q value in the duel architecture is divided into two parts: a state value function related to the state, and an advantage function representing the advantage of each action. This structure enables the network to better distinguish the importance of the state itself and the revenue brought by each optional action in that state.
[0069] Compared with the traditional Q network, this design provides more accurate action value estimation and can update the Q values of all actions simultaneously, not just the Q value of the selected action. Therefore, it significantly improves the learning efficiency and enhances the stability of policy convergence. By modeling the state value and action advantage separately, the duel architecture can better understand the relevance of the state and make more refined Q value predictions. The duel network parameters of each agent are optimized by minimizing the following loss function:
[0070]
[0071] denotes the reward of user m at time slot n+1, and γ is the reward discount factor, Target network Q-value and network Q-value, respectively.
[0072] Step 3: solve the UAV trajectory planning sub-problem in step 1) by using a conflict search algorithm.
[0073] The UAV trajectory planning sub-problem is a non-convex optimization problem, and the present application discretizes the direction and speed of the UAV, and then uses a conflict search algorithm to solve the UAV trajectory.
[0074] The UAV charging scheduling and trajectory design sub-problem in step 3) is still difficult to solve, mainly because of the strong coupling relationship between the UAV scheduling variables and the UAV position. This step first divides all the charging scheduling and trajectory design tasks of the UAVs into U sub-tasks, and then uses the Markov optimization method to transform the problem. The constraints are relaxed, and the reward of each UAV specific task is updated. At the same time, the entropy regularization discount difference is used to establish a multi-task deep reinforcement learning-based optimization problem.
[0075] We discretize the flight direction and speed level of the UAV. In the main loop, the algorithm iterates through all the UAVs and their feasible direction-speed combinations one by one. In each iteration, the utility value Γ related to the current combination is calculated and compared with the currently known optimal utility value (denoted as bestΓ). If the current utility is better, update bestΓ and its corresponding flight direction and speed. After completing a complete round of traversal, if the selection of all UAVs does not violate any constraints, the algorithm terminates. Otherwise, for each UAV that causes a conflict, the conflicting direction-speed pair will be recorded and added to the constraint set to prevent the selection of infeasible strategies in subsequent iterations.
[0076] Step 4: solve the power allocation sub-problem in step 1) by using a successive convex approximation algorithm, and solve the entire optimization problem by using an alternating optimization algorithm based on the solutions of the three sub-problems.
[0077] A first-order Taylor expansion is used to approximate the original problem to obtain an approximate convex optimization problem, and then a convex optimization solving tool is used to solve the approximate problem to obtain an approximate solution of the original problem. The three sub-problems are integrated, and an alternating optimization algorithm is used to obtain the solution of the entire problem.
[0078] The following is the pseudo code of the algorithm designed by the present application:
[0079]
[0080]
[0081]
[0082] As Figure 4As shown in the experimental results, it is illustrated that the algorithm designed in the application is superior to the other three algorithms in terms of interference to the eavesdropper.
[0083] As shown in the experimental results, it is illustrated that the algorithm designed in the application is superior to the other three algorithms in terms of interference to the eavesdropper. Figure 5 As shown in the experimental results, it is illustrated that the algorithm designed in the application is superior to the other three algorithms in terms of interference to the eavesdropper.
[0084] As shown in the experimental results, it is illustrated that the algorithm designed in the application is superior to the other three algorithms in terms of interference to the eavesdropper. Figure 6 As shown in the experimental results, it is illustrated that the algorithm designed in the application is superior to the other three algorithms in terms of interference to the eavesdropper.
[0085] The above-mentioned is the specific embodiment of the application and the technical principle used, if the change is made according to the conception of the application, the function generated still does not exceed the spirit covered by the specification and drawings, still should belong to the protection scope of the application.
Claims
1. A secure communication method for offloading tasks in an integrated air-space-ground network, characterized in that, Includes the following steps: A system model is constructed, and the mobility model, communication model, and time delay model are determined. An optimization problem is constructed with the goal of minimizing the average system latency and the eavesdropping rate. The system model includes a low-orbit satellite, several UAVs, several ground users, and eavesdroppers. The computing tasks of the ground users are handled by UAVs or satellites. The UAVs emit jamming signals to prevent the user's task information from being eavesdropped on by the eavesdroppers. The hierarchical reinforcement learning algorithm is used to solve the drone mode selection and user association sub-problems in the optimization problem. The UAV trajectory planning subproblem in the optimization problem is solved using a conflict search algorithm. The power allocation subproblem in the optimization problem is solved using a successive convex approximation algorithm, and the entire optimization problem is solved using an alternating optimization algorithm based on the solutions to the three subproblems.
2. The secure communication method for task offloading in an integrated air-space-ground network according to claim 1, characterized in that: In the aforementioned movement model: the movement of the user and the eavesdropper follows a random walk model, and the drone's flight speed is defined as V. U Then the maximum flight distance of the UAV within one time slot is D. max =V U τ, where τ represents the time slot length, then the UAV motion constraint is expressed as: q u [n] represents the coordinates of UAV u in time slot n. Represents a set of time slots; Define the minimum distance between drones as D min Its limitations are: Indicates a collection of drones; Due to the high-speed motion of low-Earth orbit (LEO) satellites, the service time for a single LEO satellite to ground users is limited. This service is only available when the elevation angle between the satellite and the ground user falls below a predetermined threshold θ. G At that time, the satellite will no longer provide services to ground users, and the satellite coverage angle θ C Represented as θ C =arccos(D R cosθ G / (D R +H l ))-θ G D R H is the Earth's radius. l The orbital altitude of satellite l; The arc length of satellite l during its coverage time is l = 2(D) R +H l )θ C Given the orbital velocity V of satellite l l Calculate the longest service time T of satellite l. l max =l / V l .
3. The secure communication method for offloading tasks in an integrated air-space-ground network according to claim 1, characterized in that: The communication model expresses the line-of-sight probability between user m and drone u at time slot n as follows: In the formula, a and b are environmental parameters. H U q represents the drone's flight altitude. m [n] represents the coordinates of user m in time slot n; therefore, the expected channel power gain is: In the formula It is the regularized Loss probability, κ represents the regularization factor, and β o This represents the channel gain at a reference distance of 1m. It is the path loss exponent; the channel from the drone to the eavesdropper is similar to the one above; Given satellite orbital altitude H l Then the channel from user m to satellite l in time slot n is: q l [n] represents the coordinates of satellite l in time slot n; Channel between drones It is the distance between the two drones, and the channel from user m to eavesdropper e is... Where d m,e [n] is the distance between the eavesdropper and the user, ξ e Rayleigh fading coefficient for a unit-mean exponential distribution; Assume the total bandwidth of the UAV is divided into K completely orthogonal sub-channels, and the bandwidth of each sub-channel is B. U ,make Represents the set of sub-channels, a binary sub-channel allocation variable. This indicates whether user m communicates with the drone through sub-channel k. If user m communicates with the drone through sub-channel k... otherwise δ u [n] is a binary variable indicating the UAV's operating mode in time slot n. If the UAV chooses to transmit interference signal δ... u [n] = 1, otherwise δ u When [n] = 0, the signal-to-noise ratio (SNR) from user m to drone u in subchannel k within time slot n can be expressed as: The transmit power of the user and the drone in time slot n are respectively, σ 2 It's channel noise. This represents the user's sub-channel allocation variable, with each sub-channel having a bandwidth of B. U The transmission rate from user m to drone u is Assume the total bandwidth of the satellite is divided into O completely orthogonal sub-channels, and the bandwidth of each sub-channel is B. L ,make Represents the set of sub-channels, a binary sub-channel allocation variable. This indicates whether user m communicates with satellite l through subchannel o. If user m communicates with satellite l through subchannel o... otherwise The signal-to-noise ratio (SNR) from user m to satellite l in subchannel o is expressed as: The transmission rate from user m to satellite l is Assume that user m can only occupy one subchannel in time slot n, i.e. Then the transmission rate of user m is User m's transmission rate needs to meet the minimum transmission rate requirement. Right now User m may be associated with a drone or a satellite. If user m is associated with drone u, then the signal-to-noise ratio from user m to eavesdropper e in subchannel k is: The transmission rate is If user m is associated with satellite l, then the signal-to-noise ratio from user m to eavesdropper e in subchannel o is: The transmission rate is In summary, the transmission rate of user m at the eavesdropper e is:
4. The secure communication method for offloading tasks in an integrated air-space-ground network according to claim 3, characterized in that: The latency model is as follows: in time slot n, the time spent by user m to transmit task i to the offloading terminal is... If user m is served by a drone, then the time is calculated. If the service is provided by satellite, then b U b L These are the service rates for drones and satellites, respectively. The total latency of user m's task i is calculated by the following formula: α m,k,u [n], α m,o,l [n] represents the sub-channel allocation variable for user m at the UAV or satellite, respectively. This represents the waiting time of user m's task i in the drone u queue. This represents the waiting time of user m's task i in the satellite l queue.
5. A secure communication method for offloading tasks in an integrated air-space-ground network according to claim 1, characterized in that: The optimization problem aimed at minimizing the system's average latency and eavesdropping rate is expressed as: Power control Sub-channel allocation drone trajectory and drone mode selection I m [n] represents the number of tasks generated by user m in time slot n, t m,i [n] represents the total latency of user m's task i, R m,e [n] represents the transmission rate of user m at the eavesdropper e, δ u [n] represents a binary variable representing the UAV's operating mode in time slot n. The UAV can choose to send jamming signals or act as an edge server. If the UAV chooses to send jamming signals δ u [n] = 1, otherwise δ u [n] = 0, This indicates whether user m communicates with the drone through subchannel k. This indicates whether user m communicates with satellite l through subchannel o. This represents the user's transmit power in time slot n. This represents the transmit power of the UAV in time slot n. λ represents the maximum transmit power of the user and the drone, respectively. m b represents the task arrival rate for user m. U F represents the service rate of the drone. l [n] represents the number of tasks that satellite l can handle in time slot n, q u [n]、q i [n] represents the coordinates of UAVs u and i in time slot n, respectively, and D min D represents the minimum distance between drones. max t represents the maximum flight distance of a drone within a time slot. m,i [n] represents the total delay of user m's task i. Indicates the maximum tolerable delay for the task; Constraint C1 is the UAV mode selection variable, which means that the UAV can only operate in one mode at a time. Constraints C2 and C3 ensure that the user can only access one sub-channel at any time. Constraint C4 ensures that the UAV cannot provide services to the user in interference mode. Constraint C5 is the transmission power limit for the user and the UAV. Constraint C6 ensures the stability of the UAV processing queue. Constraint C7 is the limit on the number of tasks that the satellite processing queue can handle. Constraints C8 and C9 are UAV trajectory constraints. Constraint C10 is the maximum delay constraint.
6. A secure communication method for offloading tasks in an integrated air-space-ground network according to claim 1 or 5, characterized in that: The hierarchical reinforcement learning algorithm divides drone mode selection and user association decision into two levels. The upper-level controller is responsible for drone mode selection, while the lower-level controller executes user association decision based on the mode determined by the upper level. The upper-level environmental state includes the drone's location, the eavesdropper's location, and the number of eavesdroppers in the nth time slot, denoted as . Its corresponding action space represents the selection strategy for each UAV operation mode in each time slot, denoted as: reward function Defined as the negative of the sum of the user's average communication rate and the eavesdropping rate, as shown below: R m [n] represents the transmission rate of user m; The lower-level environment state includes the positions of the user, drone, and satellite, as well as the current drone mode selection strategy, defined as follows: The observation space is defined as the channel state between each user and accessible nodes (including drones, satellites, and potential eavesdroppers), i.e., O m,n ={h m,u [n],h m,l [n],h m,e [n]}; Action space refers to the task unloading target selected by user m in time slot n, denoted as reward function Based on the task unloading delay design, its form is as follows:
7. A secure communication method for offloading tasks in an integrated air-space-ground network according to claim 6, characterized in that: The upper-level architecture comprises three types of neural networks: a policy network, a critic network, and a target critic network. The policy network maps the environment state to a probability distribution of available actions and samples actions based on this distribution. The target critic network receives the environment state and the corresponding action probabilities generated by the policy network as input and outputs two Q-values to evaluate the value of the actions. By introducing two Q-values, the overestimation problem that often occurs during the training process of traditional single-Q networks is effectively alleviated. The Q-network can be trained using Bellman residuals, and its loss function... The expression is as follows: in, It can be obtained through soft Bellman iteration. θ represents the parameters of the target Q-network, while θ represents the parameters of the current Q-network. n a n This represents the state and action corresponding to time slot n; In the lower-level framework, each agent is equipped with a duel network and a corresponding target duel module. The Q-value in the duel network is divided into two parts: a state-value function related to the state, and an advantage function representing the advantage of each action. The duel network parameters are optimized by minimizing the following loss function: This represents the reward for user m in time slot n+1, where γ is the reward discount factor. These are the target network Q-value and the network Q-value, respectively.
8. A secure communication method for offloading tasks in an integrated air-space-ground network according to claim 1, characterized in that: In the conflict search algorithm, the flight direction and speed level of the UAV are discretized. In the main loop, all UAVs and their feasible direction-speed combinations are traversed one by one. In each iteration, the utility value Γ associated with the current combination is calculated and compared with the currently known best utility value (denoted as bestΓ). If the current utility is better, bestΓ and its corresponding flight direction and speed are updated. After completing a full traversal, if the selection of all UAVs does not violate any constraints, the algorithm terminates. Otherwise, for each UAV that causes a conflict, its conflicting direction-speed pair will be recorded and added to the constraint set.
9. A secure communication system for offloading tasks in an integrated air-space-ground network, characterized in that: A secure communication method for offloading tasks in an integrated air-space-ground network, as described in any one of claims 1-8, includes: a low-orbit satellite, several drones, several ground users, and an eavesdropper. The computing tasks of the ground users are offloaded to the drones or satellites for processing. The drones transmit jamming signals to prevent the user's task information from being eavesdropped on by the eavesdropper.