Assisted edge calculation method for space-air-ground integrated network

Through the integrated air-space and earth-based network-assisted edge computing method, task offload decisions and drone resource allocation are optimized, and the problem of low efficiency of mobile edge computing services in remote areas is solved, and high response speed and low energy consumption is achieved.

CN120151948AActive Publication Date: 2025-06-13JILIN UNIVERSITY

Patent Information

Application Number
CN202510624221.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The prior art has difficulty providing efficient mobile edge computing services in remote or difficult-to-reach areas, especially in dynamic environments of highly mobility applications such as drones or connected vehicles, where increased latency and service disruptions are often present.

Method used

The integrated air-space and earth network assisted edge computing method is adopted to optimize the unloading decisions of user tasks, the unloading ratio, the computing resource allocation and flight trajectory of the drone, and use reinforcement learning algorithms and alliances to form a game algorithm to improve the task response speed and reduce energy consumption.

Benefits of technology

It significantly improves task response speed, reduces energy consumption, balances network load, reduces task delay, and improves resource utilization efficiency. It is suitable for high mobility application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151948A_ABST
    Figure CN120151948A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of wireless network communication, and discloses a space-air-ground integrated network auxiliary edge calculation method, which comprises the following steps: determining the number and position of unmanned aerial vehicles, the number and position of satellites, and the number and position of users needing communication in a space-air-ground integrated network; the calculation task amount and the calculation density generated by the user needing to communicate are determined; constructing an objective function, and determining an optimal task unloading proportion, an optimal task unloading strategy, an optimal unmanned aerial vehicle resource allocation strategy and an optimal unmanned aerial vehicle trajectory according to the objective function; the user communicates with the edge server for task unloading according to the optimal task unloading proportion and the optimal task unloading strategy; the unmanned aerial vehicle allocates computing resources to the computing task according to the optimal unmanned aerial vehicle resource allocation strategy, and performs position adjustment according to the optimal unmanned aerial vehicle trajectory; wherein the edge server is an unmanned aerial vehicle or a satellite; and step 4, the edge server returns a calculation result to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless network communication, and particularly relates to a space-air-ground integrated network assisted edge computing method. Background Art

[0002] The exponential growth of data-driven services such as real-time monitoring, intelligent transportation, and immersive entertainment urgently demands unprecedented support in terms of computing resources and low latency. These services usually require instant processing and quick response, while traditional user devices often cannot meet these demanding requirements due to their limited computing power and energy constraints. To address this challenge, mobile edge computing has emerged rapidly as an effective solution. By offloading computationally intensive tasks to edge servers located near user devices, mobile edge computing can significantly reduce latency and save energy, thus ensuring that emerging applications can provide seamless user experiences and efficient operations. However, the ground mobile edge computing infrastructure is often restricted by limited coverage, especially in remote or rural areas with less or unreliable network access. In addition, due to the fixed location of edge servers, these systems also have difficulty meeting the requirements of highly mobile applications (such as drones or connected vehicles) in dynamic and real-time environments, which may lead to increased latency and service interruptions.

[0003] To overcome these challenges, the space-air-ground integrated network has emerged as an innovative solution to enhance edge computing capabilities. The space-air-ground integrated network realizes a wider and more flexible coverage by comprehensively utilizing satellite, aerial, and ground communication resources. Compared with traditional ground communication networks, the space-air-ground integrated network can provide a more extensive and continuous connection, especially in areas with insufficient ground network coverage. By deeply integrating satellite and aerial platforms with the ground network, the space-air-ground integrated network can provide enhanced connectivity for mobile edge computing, significantly reducing the latency of highly mobile applications while improving the service availability in remote or hard-to-reach areas. This hybrid architecture not only expands the coverage of mobile edge computing but also optimizes the utilization of computing and communication resources, becoming an effective support platform for future intelligent applications, especially for those real-time application scenarios with extremely strict performance requirements, such as autonomous driving, drone control, and emergency rescue, providing strong technical support.

[0004] In the process of space-air-ground integrated network-assisted edge computing, due to limited communication and computing resources, unreasonable resource allocation strategies are difficult to meet the computing needs of all users in the scenario. In addition, considering the privacy and security of communication, task information is not shared and communicated between different users, and only necessary location information is shared between different drones. These constraints make it difficult to achieve satisfactory quality of service. In addition, current work on space-air-ground integrated network-assisted edge computing only focuses on drones serving static ground users, without considering the mobility of the served users (such as pedestrians, robots, or vehicles), which poses challenges to the real-time response ability of the communication system. Summary of the Invention

[0005] The purpose of the present invention is to provide a space-air-ground integrated network-assisted edge computing method, which can effectively improve the task response speed, reduce energy consumption, balance the network load, reduce task delay, and reduce energy consumption by optimizing the offloading decision of user tasks, the offloading ratio, the computing resource allocation of drones, and the flight trajectory.

[0006] The technical solution provided by the present invention is as follows: A space-air-ground integrated network-assisted edge computing method includes the following steps: Step 1: Determine the number and positions of drones, the number and positions of satellites, and the number and positions of users to communicate in the space-air-ground integrated network; and determine the amount of computing tasks and computing density generated by the users to communicate. Step 2: Construct an objective function, and determine the optimal task offloading ratio, the optimal task offloading strategy, the optimal drone resource allocation strategy, and the optimal drone trajectory according to the objective function. Among them, the objective function is: ;

[0007] In the formula, represents the task offloading ratio decision, represents the task offloading strategy, represents the drone resource allocation strategy, represents the drone trajectory decision; represents the time slot, represents the user, represents the total system time, represents the set of users to communicate; represents the user at the execution delay of the time slot, represents the user at the energy consumption of the time slot; represents the weight coefficient of the delay, Denote the weight coefficient of energy consumption; Among them, the task offloading strategy is: offloading the computing task to the UAV or satellite; the UAV trajectory decision is the moving direction and speed of the UAV; Step 3: The user communicates with the edge server according to the optimal task offloading ratio and the optimal task offloading strategy for task offloading; the UAV allocates computing resources for the computing task according to the optimal UAV resource allocation strategy, and adjusts its position according to the optimal UAV trajectory; Among them, the edge server is the UAV or satellite; Step 4: The edge server returns the calculation result to the user.

[0008] Preferably, the user at The execution delay of the time slot is:

[0009] In the formula, Denote the user at The completion delay of the local execution part of the computing task generated in the time slot, Denote the user at The completion delay of the part of the computing task generated in the time slot offloaded to the edge server part, , respectively represent the UAV set and the satellite set; Denote the The user set that offloads the computing task generated in the time slot to the edge server , Denote the user at The offloading decision in the time slot.

[0010] Preferably, when the user offloads the computing task to the UAV:

[0011] In the formula, Denote the user at The completion delay of the part of the computing task generated in the time slot offloaded to the UAV part, Denote the offloading ratio of the computing task, Denote the number of bits of the computing task, Denote the user and the UAV at The data transmission rate in the time slot, Denote the computing density of the task, Denote the drone Allocate to the user Computing resources.

[0012] Preferably, when the user offloads the computing task to the satellite:

[0013] Wherein, Denote the user At The completion delay of the computing task generated in the time slot offloaded to the satellite Part of the completion delay, Denote the offloading ratio of the computing task, Denote the number of bits of the computing task, Denote the user And the satellite At The data transmission rate in the time slot, Denote The time slot satellite And The number of network hops passed during the forwarding between them, Denote the satellite with the fewest network hops between the user and the satellite Of the visible ground base stations, Denote the inter-satellite link transmission rate, Denote the transmission rate between the satellite and the ground base station, Denote The time slot user And the satellite The distance between them, Denote The time slot satellite And The distance between them, Denote the satellite And the ground base station The distance between them, Denote the speed of light.

[0014] Preferably, the energy consumption of the user At The energy consumption in the time slot is:

[0015] Wherein, Denote the local computing energy consumption of the computing task, Denote the user At The transmission energy consumption of the computing task generated in the time slot offloaded to the edge server Of the transmission energy consumption.

[0016] Preferably, when the user offloads the computing task to the UAV:

[0017] where represents the transmission energy consumption of the computing task generated by the user in time slot being offloaded to the UAV , and represents the transmit power of the user in time slot when transmitting to the UAV .

[0018] Preferably, when the user offloads the computing task to the satellite:

[0019] where represents the transmission energy consumption of the computing task generated by the user in time slot being offloaded to the satellite , and represents the transmit power of the user in time slot when transmitting to the satellite .

[0020] Preferably, the data transmission rate between the user and the satellite in time slot is calculated by the following formula:

[0021] where , ; where represents the bandwidth between the user and the satellite in time slot, represents the instantaneous channel gain between the user and the satellite in time slot, represents the path loss between the user and the satellite represents the carrier frequency of the band, represents the distance between the satellite and the user ;

[0022] Preferably, in the second step, the method for determining the optimal task offloading ratio, the optimal task offloading strategy, the optimal UAV resource allocation strategy, and the optimal UAV trajectory is as follows: Construct a first reinforcement learning network model and a second reinforcement learning network model; Integrate the user's location, the user's task volume, the task computing density, and the user's energy consumption limit into a first state space, input it into the first reinforcement learning network model, and obtain a first action; Integrate the user's location, the UAV's location, and the UAV's own energy consumption limit into a second state space, input it into the second reinforcement learning network model, and obtain a second action; Among them, the first action is the task offloading ratio strategy, and the second action is the UAV trajectory strategy; Integrate the first action and the second action into a joint action, and use the coalition formation game algorithm based on the joint action to obtain the task offloading strategy and the UAV resource allocation strategy under the current action; Respectively determine the reward values obtained by the user and the UAV, and obtain the updated first state space and second state space, and store the first state space, the second state space, the joint action, the reward value, and the updated first state space and second state space as an experience tuple in the buffer pool; Use the experience tuples in the buffer pool to perform iterative optimization training on the first reinforcement learning network model and the second reinforcement learning network model until the set number of iterations is reached, and obtain the optimal first reinforcement learning network model and the optimal second reinforcement learning network model; Obtain the first state space and the second state space of the current time slot, and input them into the optimal first reinforcement learning network model and the optimal second reinforcement learning network model respectively to obtain the optimal task offloading ratio strategy and the optimal UAV trajectory strategy; Based on the optimal task offloading ratio strategy and the optimal UAV trajectory strategy, the user set performs a coalition formation game to obtain the optimal task offloading strategy and the optimal UAV resource allocation strategy.

[0023] Preferably, the user reward function is:

[0024] The UAV reward function is:

[0025] Among them, , ; In the formula, represents the reward value of the user in the time slot, represents the UAV in the time slot The reward value represents the total system cost of the time slot represents the sum of the timeout penalties of all users in the time slot represents the individual reward of the UAV represents the user in the time slot cost and are the weighted factors of the total system cost and the individual user cost respectively

[0026] The beneficial effects of the present invention are as follows The air-space-ground integrated network-assisted edge computing method provided by the present invention can effectively improve the task response speed, reduce energy consumption, and balance the network load by optimizing the offloading decision and offloading ratio of user tasks, so as to achieve the goals of minimizing the total task delay and energy consumption of users; in addition, by optimizing the computing resource allocation and flight trajectory of UAVs, it can not only improve the resource utilization efficiency, but also shorten the communication distance between UAVs and users, thereby further reducing the task delay and energy consumption; therefore, the mobile edge computing method assisted by the air-space-ground integrated network can significantly improve the user experience and reduce the system cost

[0027] The present invention designs the optimal task offloading strategy and optimal offloading ratio of users, the optimal computing resource allocation and trajectory of UAVs by using a multi-agent policy gradient algorithm combined with coalition formation game. It adopts a distributed architecture, and each agent makes decisions only based on its own state and limited information sharing, without relying on the decisions of other agents and without the need for a central processor for processing Description of the Drawings

[0028] Figure 1 is a schematic diagram of the communication scenario of the air-space-ground integrated network-assisted edge computing described in the present invention

[0029] Figure 2 is a flowchart of the air-space-ground integrated network-assisted edge computing method described in the present invention Detailed Embodiment

[0030] The following further elaborates on the present invention with reference to the drawings, so that those skilled in the art can implement it according to the description in the specification

[0031] Such as Figure 1As shown in the figure, the space-air-ground integrated network comprehensively utilizes satellite, aerial, and ground communication resources to achieve a wider and more flexible coverage area and improve the service availability in remote or hard-to-reach areas. However, due to the long satellite communication path and high latency, and the limited coverage area and resources of drones, it is a major challenge to meet the intensive computing tasks and latency-sensitive requirements of users.

[0032] The present invention provides a space-air-ground integrated network-assisted mobile edge computing method. Each drone is equipped with an MEC server, which can provide computing offloading services for users within its coverage area. The satellite network is connected to the core network through inter-satellite links and can forward the tasks uploaded by users to the cloud computing center of the core network for processing. Compared with the traditional MEC framework with base stations as edge servers, this architecture has the characteristics of being unrestricted by terrain and having flexible deployment. Compared with the single-layer drone network architecture, it introduces satellite nodes as a supplement, which can provide a wider coverage area and more stable connection for users within the coverage area. Users can choose to communicate with drones or satellites according to the designed optimal task offloading strategy and optimal offloading ratio. At the same time, the drone allocates computing resources for users according to the optimal resource allocation strategy and adjusts the optimal position.

[0033] As Figure 2 shown in the figure, the present invention provides a space-air-ground integrated network-assisted edge computing method, which establishes a multi-objective joint optimization model for reducing the user task execution latency and reducing the user energy consumption. Then, an improved multi-agent reinforcement learning algorithm is used to design the offloading ratio, association decision, resource allocation decision, and optimal position of drone movement for user tasks according to the limited observation information at each moment. The multi-agent reinforcement learning algorithm decouples the discrete continuous hybrid action space and reduces the space dimension by introducing an association decision method based on coalition formation game and a resource allocation method based on convex optimization, improving the convergence speed and effectiveness of the algorithm. The specific implementation process is as follows: I. Design two optimization functions according to the long-term latency and energy consumption requirements, and construct a multi-objective optimization problem according to the multi-objective optimization theory.

[0034] First, the overall objective function is designed as follows: ; Among them, ; ; ; ; When the user chooses to offload to the drone : ; ; ; ; ; ; When the user selects to unload to the satellite : ; ; ; ; ; The first target represents the total completion time delay of the user task, where represents the user in The completion time delay of the part of the task generated in the time slot and selected to be executed locally, represents the user in The completion time delay of the part of the task generated in the time slot and selected to be unloaded to the edge server ; represents the set of users who unload the computational tasks generated in the time slot to the edge server ; represents the user's in The unloading decision in the time slot, , , respectively represent the sets of all users, drones, and visible satellites, The user in The unloading ratio of the computational task in the time slot, represents the computational density of the task, represents the number of bits of the task, represents the local computing power of the user ; represents the user and the drone in The data transmission rate in the time slot, represents the computing power allocated by the drone to the user ; represents the user and the drone The bandwidth between, represents the user The transmission power between the user and the UAV in the time slot, where represents the background noise power, represents the user and the UAV in the instantaneous channel gain in the time slot, represents the user and the UAV the path loss between them, and represent the additional path losses of the LoS and NLoS links respectively, represents the user and the UAV the distance between them, represents the carrier frequency of the band, represents the speed of light, represents the probability of the Los link, and represent the environment-related variables respectively, represents the user and the UAV the elevation angle between them, represents the user and the satellite the data transmission rate in the time slot , represents the inter-satellite link transmission rate represents the transmission rate between the satellite and the ground station, represents the user and the satellite the bandwidth between them, represents the number of network hops the satellite forwards through, represents the user and the satellite in the transmission power at the moment, where represents the background noise power, represents the user and the satellite in the instantaneous channel gain in the time slot, represents the user and the satellite the path loss between them, represents the user and the satellite the distance between them, represents the carrier frequency of the Ka band, Denotes rainfall attenuation, determined based on a Weibull random process.

[0035] The second objective Denotes the total energy consumption of the user equipment. Among them, Denotes the local computing energy consumption of the task, Denotes the energy consumption efficiency of the user equipment when processing tasks, determined by the CPU chip architecture, Denotes the user At The transmission energy consumption when the task generated in the time slot is offloaded to the drone When Denotes the user in the time slot To the drone The transmission power when transmitting, Denotes the user At The transmission energy consumption when the task generated in the time slot is offloaded to the satellite When Denotes the user in the time slot To the satellite The transmission power when transmitting.

[0036] II. By using the weight coefficient method, the original multi-objective optimization problem is transformed into a single-objective optimization problem, which simplifies the complexity, unifies the evaluation criteria, and facilitates the trade-off between different objectives.

[0037] The total objective function is designed as follows: ;

[0038] , , , , , , Among them, Denotes the task offloading ratio decision, Denotes the task offloading strategy (association strategy), Denotes the drone resource allocation strategy, Denotes the drone trajectory decision, , , Denotes the sets of all users, drones, and visible satellites respectively, Denotes the total system time, Denotes the selection to offload to the drone The set of users, indicating the user at the offloading ratio of the task generated in the time slot, indicating the user at the offloading decision in the time slot, indicating the user at the execution delay in the time slot, indicating the user the maximum tolerable delay, indicating the UAV at the computing resources allocated to the user in the time slot , indicating the UAV the available maximum computing resources, and respectively indicate the UAV and the user at the energy consumption in the time slot, and indicating the UAV and the user the maximum energy limit.

[0039] Based on the above objective function, an optimal policy network is trained using an improved multi-agent reinforcement learning algorithm. The user policy network and the UAV policy network can respectively output the best offloading ratio of the user and the best position of the UAV movement according to the environmental state observed at each moment, and execute a coalition formation game according to the actions taken to find the optimal association decision and resource allocation decision to minimize the above system total cost.

[0040] In terms of architecture, the present invention takes into account the privacy and security of communication. Task information is not shared and communicated between different users, and only necessary location information is shared between different UAVs.

[0041] The present invention considers the mobility of ground users and a more comprehensive communication model. A probabilistic line-of-sight (LoS) channel model is used to model UAV communication and satellite communication. Considering that user satellite communication is sensitive to environmental factors such as rainfall, the path loss between the user and the satellite comprehensively considers rainfall attenuation and path loss fading.

[0042] III. Network initialization. Initialize the experience buffer pool and the actor network and critic network of the UAV and user agents. To improve the stability of the training process, the actor network and critic network are respectively composed of an online network and and the corresponding target network and which consists of, where and respectively represent the network parameters of the online networks of the actor and the critic, and respectively represent the network parameters of the target networks of the actor and the critic. denotes the set of two types of agents, i.e., users and UAVs. These two types of agents adopt similar network architectures and initialize the network parameters in a randomly initialized manner.

[0043] IV. Train the network parameters based on the architecture of centralized training and distributed execution. First, the agents make actions according to the observations at the current moment, execute the joint actions, and make the association decision and resource allocation decision for the current time slot according to the coalition formation game algorithm. The environment gives the rewards for the current time slot according to the reward functions of each agent respectively, stores the joint observations, joint actions, joint rewards and the observations at the next moment in the experience buffer pool as a quadruple, and updates the network parameters according to the update rules. The specific process is as follows: (1) Agent obtains the observations of the current time slot from the environment where for user it includes the positions of all users and UAVs , its own task volume , the computing density and its own energy consumption limit , for UAV , it includes the position information of all users and UAVs and its own energy consumption limit . The observation information only includes its own visible information and the limited information publicly disclosed by other agents. It is a subset of the global state , and the global state includes the position information of all users and UAVs, the energy limit, and the size and computing density of the tasks generated by all users. The observation of agent is input into the policy network to obtain the action , where the user agent makes the offloading ratio decision , and the UAV corresponds to the trajectory decision , corresponding to the direction and moving speed of the UAV.

[0044] (2) According to the joint action execute the coalition formation game algorithm to obtain the best association decision and the best resource allocation strategy under the current action.

[0045] (3) The environment gives rewards to the actions of the agents, where the user and the drone have the following reward functions respectively: ,

[0046] ,

[0047] where, and are the weighting factors of the total system reward and the individual reward respectively, represents the total system cost at the time slot, represents the sum of the timeout penalties of all users, which can be expressed as where is a binary variable indicating whether the user's task times out, represents the timeout penalty for this task, represents the user's individual cost, represents the individual reward of the drone, which can be calculated as The individual reward of the drone includes the sum of the individual costs of all users served in this time slot and the possible out-of-bounds penalty and collision penalty of itself. Among them, and are binary variables indicating whether the drone goes out of bounds or collides, represents the out-of-bounds penalty of the drone, represents the collision penalty of the drone.

[0048] (4) The joint observation , the action , the reward and the observation at the next moment are stored as an experience tuple in the buffer pool .

[0049] (5) In the network update stage, randomly sample a batch of experiences of size from the buffer pool for network update.

[0050] Specifically, in the training stage, the drone and the user act as agents interacting with the environment. First, the observation information of the agent is fed into the policy to obtain the action at the current time step. Then, by executing the joint action and using the user association and resource allocation obtained through the CFG method, the environment provides the joint reward and the next joint observation value . Then, the experience tuple is stored in the replay buffer . During the network update phase, a batch of experiences of size is randomly selected from the replay buffer . The online critic network is updated according to the sampled experiences to minimize the temporal difference error. After updating the critic, the actor network is updated every steps to improve the policy based on the critic's feedback. Finally, stability is ensured by gradually updating the target networks and .

[0051] Among them, the specific process of the coalition formation game is as follows: During the coalition formation process, users are regarded as players, and edge servers are regarded as coalitions, aiming to minimize the task cost. These players independently choose drones or satellites for task offloading, and players who choose the same edge server are regarded as joining the same coalition. Among them, represents the set of edge servers.

[0052] To measure the performance difference between two coalitions, we first introduce the concept of preference-based coalitions.

[0053] Definition (preference coalition): Given two coalitions and in the CFG, if the former has a higher coalition utility, that is, performs better than . Therefore, the preference coalition can be described as: , where represents the coalition utility, that is, the total cost of drone serving users. represents the set of users who choose to offload to the edge server .

[0054] Under the guidance of the preference relationship, each user will follow the following rules to join or leave the coalition, including the coalition switching rule and the insertion rule.

[0055] Switching rule: Suppose two users , , when the following conditions are met, the two users are more inclined to switch their coalitions, that is: , Insertion rule: Suppose user , user prefers to join the coalition from ​ : , where represents the users included in the new coalition formed after the coalition with the offloading decision to the server removes user and introduces user . represents the users included in the new coalition formed after the coalition with the offloading decision to the server removes user and introduces user . represents the offloading decision of user . represents the set of offloading decisions of other users except user . represents the utility obtained from the offloading decision made by user . and represent the utilities of the original two coalitions of users and before the exchange is made. The above two rules ensure that when an exchange or insertion occurs, the total utility of the two new coalitions is greater than that of the original coalition. The key idea of the CFG algorithm is to iteratively update the association strategies of the participants using the above two rules until a Nash equilibrium is reached. represents the users included in the new coalition formed after the coalition with the offloading decision to the server removes user . represents the users included in the new coalition formed after the coalition with the offloading decision to the server introduces user . The above two rules ensure that when an exchange or insertion occurs, the total utility of the two new coalitions is greater than that of the original coalition. The key idea of the CFG algorithm is to iteratively update the association strategies of the participants using the above two rules until a Nash equilibrium is reached.

[0056] The process of the algorithm is as follows: (1) Randomly select a user , determine the coalition it currently belongs to, denoted as , and randomly select a coalition from the set of coalitions it can join.

[0057] (2) Randomly select a user from who can exchange coalitions with it.

[0058] (3) Determine the insertion rule. Assume that user joins the new coalition . Obtain the sub-optimal resource allocation expression under the current coalition state through the Lagrangian dual method and the KKT conditions, Calculate the utility and determine whether the overall utility becomes better after joining the new coalition. If it becomes better, perform the exchange.

[0059] (4) Otherwise, determine the exchange rule. Assume the user and exchange coalitions with each other. Allocate resources through the sub-optimal resource allocation expression in the current coalition state, and calculate the overall utility.

[0060] (5) Determine whether the total utility of the new coalition becomes better after the two users exchange coalitions. If it becomes better, perform the exchange.

[0061] (6) Iterate the first six steps cyclically until the Nash equilibrium is reached.

[0062] V. Online review network Update according to the sampled experience to minimize the temporal difference error. After updating the critic, the actor network is updated every d steps to improve the policy according to the feedback of the critic. Finally, ensure stability by gradually updating the target network and The specific update process is as follows: (1) actor: Update the actor network of each agent to maximize the expected return. Calculate the policy gradient using the deterministic policy gradient theorem.

[0063]

[0064] where, is the gradient vector of the function with respect to the critic network parameters , represents the policy objective function, usually expressed as , representing the expected value of the discounted cumulative reward of the agent over the entire time span , where represents the discount factor. represents the Q-value function of the agent calculated by the online sub-network of the agent network with parameters , representing the expected cumulative reward after performing the action in the global state , represents the gradient of the Q-value function with respect to the action, measuring the rate of change of the Q-value with respect to the action under the current policy. is the output of the actor network with parameters of the agent . Then, these gradients are backpropagated to the online sub-network of the actor network to update , , Among them, represents the learning rate of the online actor network.

[0065] (2) critic: The critic network can be updated through the temporal difference error. The mean squared error loss of the Critic network can be calculated as ; Where is the TD target value, used to calculate the value error, which can be expressed as: , Among them, and are the outputs of the target subnets of the actor and critic networks of the agent respectively, and represents the discount factor. Then, the gradient of

[0066] Therefore, the critic network can be updated as follows , Among them, is the learning rate of the critic network, and then by making the two target networks of the agent slowly track the online network that is learning to update their weights , , Among them, represents the update rate of the target network.

[0067] Whether the iteration termination condition reaches the maximum number of iterations; if the iteration termination condition is met, the current network parameters are saved as the final policy network in the distributed execution phase, otherwise the iteration update process continues.

[0068] In the execution phase, each user and drone respectively obtain the optimal offloading ratio or optimal position with the current time slot observation as the input according to their own actor network parameters.

[0069] According to the current actions, the user set performs a coalition formation game to obtain the optimal association and optimal resource allocation.

[0070] Complete the edge computing tasks of this time slot according to the obtained solution, and the drone moves to the optimal position.

[0071] The above process is repeated in each time slot until the communication ends.

[0072] Due to the dynamics and heterogeneity of the network, as well as the limited information sharing in the scenario, it is difficult for traditional methods to fully model our scenario. In addition, traditional algorithms have limitations in dealing with trajectory planning and similar tasks because they are often ineffective in optimizing the cumulative objectives over multiple time slots and focus more on short-term results.

[0073] Traditional MADRL algorithms are difficult to effectively model and solve the discrete-continuous hybrid action space, and these algorithms are usually optimized for discrete or continuous domains. At the same time, with the addition of multiple optimization variables such as UAV trajectory planning, offloading ratio, and computing resource allocation, the action and state spaces grow exponentially, significantly increasing the computational burden and reducing the efficiency of the algorithm to find the optimal solution. In addition, the number of tasks offloaded to the UAV varies with time slots, which leads to fluctuations in the action dimension, complicating the learning dynamics of the MADRL algorithm and significantly increasing the difficulty of solving.

[0074] Based on the above challenges, we incorporate the coalition formation game into the MADRL framework to solve the UD association and resource allocation sub-problems under given actions. By doing so, we transform the problem into a partially observable Markov decision process (POMDP) with a smaller action space, thus reducing the complexity associated with the mixed discrete-continuous action space and combining the advantages of the two methods, with better convergence and performance.

[0075] And we regard UAVs and users as two types of heterogeneous agents, and design their respective reward functions according to their own utility expectations, with better fairness performance and being able to reasonably balance the overall utility and their own utility.

[0076] The present invention provides a method for high-reliability and low-latency computing offloading service based on space-air-ground integrated network assisted edge computing. The UAV cluster can move to the optimal position in real time according to the real-time position of mobile users and allocate reasonable computing resources to the users who choose to offload. Ground users can reasonably select the offloading UAVs and the offloading ratio, and obtain good latency and energy consumption performance under the condition of only providing necessary location and task information, ensuring information security and service quality. In addition, by designing reasonable weight factors, a certain trade-off can be achieved between minimizing the execution latency and minimizing the energy consumption. Therefore, the purpose of realizing high-reliability and low-latency computing offloading service can be achieved through the method of space-air-ground integrated network assisted edge computing. The present invention combines multi-agent reinforcement learning algorithms with traditional methods such as game theory and convex optimization to solve multi-objective optimization problems, and strengthens the stability and convergence speed of training by designing target networks, experience replay mechanisms, etc. At the same time, by designing reward functions for the two types of agents specifically, the trade-off between overall utility and self-utility is achieved, ensuring fairness.

[0077] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A space-ground integrated network-assisted edge computing method, characterized in that: The steps include: Step 1: Determine the number and location of drones, the number and location of satellites, and the number and location of users who need to communicate in the air-ground integrated network; and determine the amount of computing tasks and computing density generated by users who need to communicate; Step 2: construct an objective function, and determine the optimal task offloading ratio, the optimal task offloading strategy, the optimal UAV resource allocation strategy and the optimal UAV trajectory according to the objective function; Wherein, the objective function is: ; In the formula, represents the task offloading ratio decision, Indicates the task offloading strategy, represents the UAV resource allocation strategy, represents the UAV trajectory decision; Indicates time slot, Represents the user, Indicates the total system time, Represents the set of users who need to communicate; Indicates user exist The execution delay of the time slot, Indicates user exist Energy consumption of time slots; represents the weight coefficient of delay, Represents the weight coefficient of energy consumption; The task offloading strategy is: offloading the computing task to the drone or satellite; the drone trajectory decision is the moving direction and speed of the drone; Step 3: The user communicates with the edge server to offload tasks according to the optimal task offloading ratio and the optimal task offloading strategy; the drone allocates computing resources to the computing tasks according to the optimal drone resource allocation strategy, and adjusts its position according to the optimal drone trajectory; Wherein, the edge server is a drone or a satellite; Step 4: The edge server returns the calculation result to the user.

2. The air-ground integrated network-assisted edge computing method according to claim 1, characterized in that: user exist The execution delay of a time slot is: ; In the formula, Indicates user exist The completion delay of the local execution part of the computing task generated by the time slot, Indicates user exist The computational tasks generated by time slots are offloaded to edge servers Partial completion delay, , denote the set of drones and the set of satellites respectively; Indicates that The computational tasks generated by time slots are offloaded to edge servers A collection of users, Indicates user exist Unloading decision for time slots.

3. The air-ground integrated network-assisted edge computing method according to claim 2, characterized in that: When users offload computing tasks to unmanned devices: ; In the formula, Indicates user exist The computational tasks generated by the time slot are offloaded to the drone Partial completion delay, Indicates the offloading ratio of computing tasks, represents the number of bits of the computational task, Indicates user With drones exist The data transmission rate of the time slot, represents the computational density of the task, Indicates drone Assign to user computing resources.

4. The air-ground-space integrated network-assisted edge computing method according to claim 3, characterized in that: When users offload computing tasks to satellites: ; In the formula, Indicates user exist The computational tasks generated by the time slot are offloaded to the satellite Partial completion delay, Indicates the offloading ratio of computing tasks, represents the number of bits of the computational task, Indicates user With satellite exist The data transmission rate of the time slot, express Time slot satellite and The number of network hops forwarded between Indicates and satellite Satellites with the fewest network hops between visible ground base stations, represents the intersatellite link transmission rate, Indicates the transmission rate between the satellite and the ground base station. express Time slot user With satellite The distance between express Time slot satellite and The distance between Indicates satellite With ground base station The distance between Represents the speed of light.

5. The air-ground integrated network-assisted edge computing method according to any one of claims 2 to 4, characterized in that: user exist The energy consumption of a time slot is: ; In the formula, represents the local computing energy consumption of the computing task, Indicates user exist The computational tasks generated by time slots are offloaded to edge servers transmission energy consumption.

6. The air-ground-space integrated network-assisted edge computing method according to claim 5, characterized in that: When users offload computing tasks to drones: ; In the formula, Indicates user exist The computational tasks generated by the time slot are offloaded to the drone The transmission energy consumption, Indicates user exist Time slot to drone The transmit power during transmission.

7. The air-ground-space integrated network-assisted edge computing method according to claim 6, characterized in that: When users offload computing tasks to satellite machines: ; In the formula, Indicates user exist The computational tasks generated by the time slot are offloaded to the satellite The transmission energy consumption, Indicates that the user is Time slot to satellite The transmit power during transmission.

8. The air-ground-space integrated network-assisted edge computing method according to claim 7, characterized in that: user With satellite exist The data transmission rate of a time slot is calculated by the following formula: ; in, , ; In the formula, express Time slot user With satellite The bandwidth between Indicates user With satellite exist The instantaneous channel gain of the time slot, Indicates user With satellite The path loss between express The carrier frequency of the band, Indicates satellite With users The distance between Represents rainfall attenuation.

9. The air-ground-space integrated network-assisted edge computing method according to claim 8, characterized in that: In step 2, the method for determining the optimal task offloading ratio, the optimal task offloading strategy, the optimal UAV resource allocation strategy and the optimal UAV trajectory is: Constructing a first reinforcement learning network model and a second reinforcement learning network model; Integrate the user's position, the user's task volume, the task computing density, and the user's energy consumption limit into a first state space, input the first reinforcement learning network model, and obtain a first action; Integrate the user's position, the drone's position, and the drone's own energy consumption limit into a second state space, input it into the second reinforcement learning network model, and obtain a second action; Among them, the first action is the task offloading ratio strategy, and the second action is the drone trajectory strategy; Iteratively optimize and train the first reinforcement learning network model and the second reinforcement learning network model until a set number of iterations is reached to obtain an optimal first reinforcement learning network model and an optimal second reinforcement learning network model; Obtain the first state space and the second state space of the current time slot and input the optimal first reinforcement learning network model and the optimal second reinforcement learning network model respectively to obtain the optimal task offloading ratio strategy and the optimal UAV trajectory strategy; Based on the optimal task offloading ratio strategy and the optimal drone trajectory strategy, the user set performs an alliance formation game to obtain the optimal task offloading strategy and the optimal drone resource allocation strategy.

10. The air-ground-space integrated network-assisted edge computing method according to claim 9, characterized in that: In the first reinforcement learning network model, the user reward function is: ; In the second reinforcement learning network model, the drone reward function is: ; in, , ; In the formula, express Time slot user The reward value, express Time slot drone The reward value, represent Total system cost of the time slot, represent The sum of the timeout penalties of all users in the time slot, Represents individual drone rewards, represent Time slot user The cost, and are the total system cost and user individual cost weighting factors respectively.

Citation Information

Patent Citations

  • Calculation unloading method and system based on space-air-ground remote Internet of Things

    CN110868455A

  • Time delay minimization calculation task unloading method and system in space-air-ground integrated network

    CN113346944A

  • Task unloading method and system based on game theory under air-space-ground integrated network

    CN116192228A

  • Task unloading method in space-air-ground network based on multi-target depth Q network

    CN116431240A

  • Multi-unmanned aerial vehicle assisted mobile edge calculation method based on robustness

    CN116528301A

Cited By

  • Double-time-scale optimization method for assisting task unloading and caching in cooperation of multiple unmanned aerial vehicles

    CN121541943A

  • Multi-uav cooperative auxiliary task unloading and cache double-time-scale optimization method

    CN121541943B