Unmanned aerial vehicle networking calculation unloading method based on task age

By using a drone networking computational offloading method, and combining task age and multi-agent deep reinforcement learning to optimize drone trajectory and task offloading, the problem of insufficient adaptability of traditional mobile edge computing in dynamic user demand scenarios is solved, thereby improving task freshness and optimizing system efficiency.

CN121126450APending Publication Date: 2025-12-12NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511354710.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Traditional mobile edge computing is not adaptable to dynamic user demand scenarios, especially in disaster recovery, crowded places or intermittent connection environments, and cannot effectively quantify the freshness of information, leading to task delays and wrong decisions.

Method used

A task-age-based UAV networking computational offloading method is adopted. By establishing an objective function based on the dynamic correspondence between UAVs and users, transmission power, CPU frequency, and energy consumption limit, the service allocation and task offloading strategies are iteratively solved. The UAV trajectory and task offloading are optimized by combining multi-agent deep reinforcement learning and convex optimization theory.

Benefits of technology

It improved the freshness of task completion, reduced erroneous decisions caused by outdated data, enhanced data timeliness and system efficiency, and optimized the energy and resource utilization of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126450A_ABST
    Figure CN121126450A_ABST
Patent Text Reader

Abstract

The invention discloses a task age-based unmanned aerial vehicle networking calculation unloading method and device. The method comprises the steps of obtaining a calculation task generated by a user; generating task freshness of the user according to the calculation task; taking a one-to-one corresponding service relationship between the unmanned aerial vehicle and the user, a transmitting power peak value of the unmanned aerial vehicle, a transmitting power peak value of the user, a CPU frequency upper limit of the user and a total energy consumption upper limit of the centralized control architecture as constraints, and establishing a target function with minimum task freshness of all users in the centralized control architecture; iteratively solving the objective function; according to the method, the task freshness of the user is generated according to the calculation task information generated by the user, and then the target function is generated through the task freshness, so that the operation of a centralized control framework can be realized through the task freshness, the data timeliness is improved, and wrong decisions caused by expiration results are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Internet of Things (IoT) communication technology, and in particular relates to a method for unloading unmanned aerial vehicle (UAV) networking calculations based on task age. Background Technology

[0002] Traditional mobile edge computing (MEC) deploys computing resources at the edge of a fixed network, significantly reducing latency and energy consumption compared to centralized cloud computing by offloading tasks to nearby servers. However, its static deployment limits its adaptability to dynamic user demands, especially in scenarios requiring rapid scaling, such as disaster recovery, congested venues (e.g., stadiums), or environments with a large number of intermittent connections. These challenges are exacerbated by the increasing prevalence of compute- and communication-intensive applications where terrestrial MEC may be temporarily unavailable or inefficient.

[0003] Unmanned Aerial Vehicle (UAV)-assisted mobile edge computing offers a transformative solution by leveraging the maneuverability, flexible deployment, and high-probability line-of-sight (LOS) connectivity of UAVs to overcome these limitations. UAVs can dynamically reposition themselves and hover near ground users (GUs), reducing propagation distance and improving channel conditions for task offloading. This mobility-driven approach not only ensures reliable communication even when ground infrastructure is compromised but also enhances resource accessibility for edge computing. Furthermore, UAVs can act as airborne relays, efficiently forwarding tasks to remote MEC servers while mitigating additional latency through optimized relay strategies, including transmission power control and relay selection.

[0004] Large-scale Internet of Things (IoT) and low-latency applications rely on timely delivery of up-to-date information. The timeliness of data is crucial when it reaches IoT devices, as outdated data can be meaningless and lead to flawed decisions. However, traditional throughput and latency metrics cannot adequately quantify the freshness of information. Summary of the Invention

[0005] The purpose of this invention is to provide a method for calculating unloading in UAV networking based on task age, so as to effectively improve the freshness of task completion.

[0006] This invention adopts the following technical solution: a method for unmanned aerial vehicle (UAV) network computational offloading based on task age, applied in a centralized control architecture, which includes several users, several UAVs, and an MEC server; the method includes the following steps: Obtain user-generated computing tasks; computing tasks can be computed locally by the user, offloaded to the MEC server via a drone for computation, or computed jointly by the user and the MEC server; The user's task freshness is generated based on the computational task; the task freshness increases with the increase of time slots. The objective function is established with the constraints of the one-to-one service relationship between drones and users, the peak transmission power of drones, the peak transmission power of users, the upper limit of CPU frequency of users, and the upper limit of total energy consumption of centralized control architecture, and the minimum task freshness of all users in centralized control architecture. The objective function is solved iteratively to obtain the service allocation indication matrix, task offloading ratio matrix, user computing power matrix, transmission power matrix, and UAV global trajectory vector matrix for each time slot.

[0007] Another technical solution of the present invention: a UAV networking calculation offloading device based on task age, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement any of the above methods.

[0008] The beneficial effects of this invention are: this invention generates the user's task freshness based on the user-generated computing task information, and then generates the objective function through the task freshness. Thus, the operation of a centralized control architecture can be achieved through task freshness, improving the timeliness of data and avoiding erroneous decisions caused by expired results. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the system architecture of the UAV-assisted mobile edge computing system according to an embodiment of the present invention; Figure 2 This is a diagram illustrating the evolution of AoT according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the MADDPG-SCA method according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating the iterative algorithm of an embodiment of the present invention; Figure 5 This is a UAV trajectory diagram of the experimental method in an embodiment of the present invention; Figure 6 The drone trajectory diagram is a comparison method of the present invention embodiment; Figure 7 This is a comparison chart of the average PAOT and unloading ratio of the experimental method and its comparative method in the embodiments of the present invention; Figure 8 This is a comparison chart of the comparative method of the present invention and the comparison of the energy consumption and travel distance of the drone using the comparative method. Detailed Implementation

[0010] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0011] This invention is based on the novel concept of Information Age (AoI), defined as the time elapsed from its initial generation to its reception by the device. However, since the concept of AoI cannot accurately describe the freshness of information when a specific task needs to be received, a new performance metric called Task Age (AoT) is proposed, which further considers the freshness of information for a given task execution. Specifically, AoT measures the execution time from the generation of the latest task at the source to its completion.

[0012] Therefore, this invention discloses a UAV networking computation offloading method based on task age, applied in a centralized control architecture, which includes several users, several UAVs, and an MEC server. The method includes the following steps: obtaining computation tasks generated by users; computation tasks are computed locally by users, offloaded to the MEC server by UAVs, or computed jointly by users and the MEC server; generating task freshness for users based on computation tasks; task freshness increases with the increase of time slots; establishing an objective function with the minimum task freshness of all users in the centralized control architecture, constrained by the one-to-one service relationship between UAVs and users, the peak transmit power of UAVs, the peak transmit power of users, the upper limit of CPU frequency of users, and the upper limit of total energy consumption of the centralized control architecture; iteratively solving the objective function to obtain the service allocation indication matrix, task offloading ratio matrix, user computing power matrix, transmission power matrix, and global trajectory vector matrix of UAVs for each time slot.

[0013] This invention generates a user's task freshness based on the user-generated computing task information, and then uses the task freshness to generate an objective function. This allows for the operation of a centralized control architecture through task freshness, improving data timeliness and avoiding erroneous decisions caused by outdated results.

[0014] Flying Ad Hoc Networks (FANETs), with their unique multi-hop communication mechanism combined with UAV device-to-device (D2D) relay technology, can effectively overcome the limitations of ground base stations in complex scenarios. In extreme conditions such as military confrontations, sudden natural disasters, and uninhabited desert areas, ground infrastructure often suffers from discontinuous coverage or even complete inability to be deployed. UAVs can dynamically construct aerial communication links through multi-hop collaboration. When the communication range of a single UAV is limited, the multi-hop mechanism allows data to be relayed between UAV nodes, forming a temporary network with wide-area coverage, thereby filling the coverage gaps of ground base stations.

[0015] In the integrated application of UAV networking and task processing, the value of the multi-hop mechanism is further highlighted: On the one hand, UAVs can act as mobile terminals, "offloading" the sensor data or computationally intensive tasks they acquire to the cloud / edge for efficient processing via FANET's multi-hop links (i.e., the "UAV-to-the-cloud" mode). In this way, multi-hop transmission can flexibly bypass communication barriers, ensuring that task data is stably delivered to processing nodes in a dynamic topology. On the other hand, UAVs can act as multi-hop relay nodes, assisting ground users in completing task offloading: the ground user's data is first transmitted to the nearest UAV, and then relayed to the cloud / edge through multi-hop collaboration between UAVs.

[0016] This invention designs a UAV networking protocol and introduces a potential function-based UAV networking control method to address the network coordination challenges posed by multiple heterogeneous link layers. In UAV swarm networking design, the routing coordination mechanism is a core element in ensuring network performance. This mechanism encompasses two main functional modules: route discovery and maintenance. Route discovery relies on link-layer broadcast route request (RREQ) messages to perform topology probing; on-demand routing mode establishes low-latency connections for bursty data transmission; and proactive routing adapts to low-speed, steady-state transmission scenarios through periodic maintenance.

[0017] The network layer employs a standardized IPv6 / IPv4 protocol stack to ensure compatibility and relies on ICMP / IGMP protocols for route adaptation management. Addressing the needs for computing power sharing and task offloading in multi-hop communication networks, this invention proposes a method based on link connectivity... The key optimization metric—which integrates multi-dimensional parameters such as signal-to-noise ratio, energy consumption, and outage probability—directly characterizes topology performance through its network-wide sum function, Fit. While traditional centralized optimization can be solved using shortest path algorithms, it faces the challenge of excessive global synchronization overhead in distributed scenarios.

[0018] To this end, this invention innovatively constructs an on-demand topology optimization strategy based on game theory: triggering local updates through a pre-deletion connection mechanism, combined with link importance weights quantized by feature vectors (…). The connectivity gain function drives the distributed decision-making to converge to Nash equilibrium, ultimately achieving near-global optimal network topology control.

[0019] In network protocol design, route coordination comprises two parts: route discovery and maintenance. For route discovery, both on-demand routing and proactive routing require the widespread distribution of Route Request (RREQ) messages at the link layer for topology discovery. The difference lies in that on-demand routing establishes routes only before data transmission is required, resulting in a longer route establishment time suitable for bursty transmissions of large amounts of data; while proactive routing involves periodic route discovery and maintenance, making it suitable for low-rate periodic transmissions.

[0020] The network layer packetization protocol uses standard IPv6 / IPv4 protocols to ensure good compatibility with existing systems, while network management uses ICMP (Internet Control Message Protocol) and IGMP (Internet Group Management Protocol). This is used to coordinate node routing adaptation issues, such as destination unreachability, source suppression, timeouts, changes in routing protocols, and routing parameter reporting. For uplink and downlink protocol adaptation, well-defined service primitives (interface APIs) are used to separate the upper and lower layer protocols.

[0021] Whether it's sharing computing power among drone clusters or unloading tasks, both require transmission through a multi-hop communication network formed by the drones.

[0022] Connectivity of drone n to drone n' It can be obtained by fusing a series of indicators such as signal-to-noise ratio (SNR), energy consumption, and interruption probability. That is: (1) Among them, based on The fusion function is determined based on the characteristics of the task; the larger the better. The fusion method can be... Linear weighted summation or other nonlinear fusion methods will not be discussed in this embodiment. The signal-to-noise ratio and interference ratio are respectively. For interruption probability, For nodes Energy consumption. Furthermore, the overall network topology Fit can be represented by the connectivity and function of all links, i.e.: (2) Typically, in centralized environments, maximizing the connectivity between any two nodes can be modeled as a shortest path problem based on the sum of connectivity, and solved using algorithms such as Dijkstra's algorithm and Floyd's algorithm. However, in distributed environments, optimizing the connectivity of all connections requires obtaining the signal-to-noise ratio, actual link interruption probability, energy consumption, etc., of all nodes in the entire network, resulting in significant global synchronization overhead. Therefore, this invention proposes a UAV networking control method based on a potential function, which achieves near-global topology control performance while realizing distributed convergence of the optimal forwarding strategy.

[0023] Specifically, the communication between drones is likely to be via the Loss of Space (LoS) channel, resulting in channel attenuation. It is directly related to the distance between the two drone nodes. This represents the channel attenuation coefficient from node n to node n′. Indicates a "proportional to" relationship. Indicates the distance from node n to node distance, This represents path loss, i.e., signal attenuation. It is inversely proportional to the power of q of the distance.

[0024] When the distance between the two drones Distance less than the signal-to-noise ratio and the interruption probability threshold Communication connections can be established at any time. To prevent collisions, the distance between each drone should be greater than the minimum safe distance. Circular area } and are the areas for maintaining links between drones, where and They represent drones and Location, and Let be a hyperparameter related to signal-to-noise ratio, interruption probability, and safe distance, and satisfy . .

[0025] When the distance between the two drones satisfy If a connection exists, indicating that the communication link between the two drones is on the verge of deterioration, then this connection is called a pre-deletion connection. In this case, the connection should be updated. To ensure network connectivity, it is assumed that at most one connection will be deleted at a time, and a distributed game theory approach is used to ensure that all drone nodes reach a consensus on the pre-deletion connection to be deleted.

[0026] Since the current connection not only connects two independent drone nodes but also serves a forwarding function within the entire network, when updating a connection, in addition to considering the connectivity of a single connection, its importance within the entire network should also be taken into account. This importance can be described by the eigenvector corresponding to the largest eigenvalue of the adjacency matrix in the graph network. Each element in this eigenvector represents the importance of the current connection to other nodes. Let the largest eigenvalue of the graph's adjacency matrix be denoted as... Its corresponding eigenvector is denoted as Then the eigenvector medium elements That is, to represent the nodes The importance of links. The importance is the product of the feature vectors of the two UAVs connected to this link, denoted as: .

[0027] The link update game can be represented as ,in A collection of drones participating in the game; To participate in drones The set of strategies This indicates that it is a current drone. With drones Establish a single-hop direct connection; otherwise, do not establish a connection. For drones participating in the game Take strategy The cost function is defined as follows: (3) Each drone node updates its strategy based on its link status to maximize its cost function. The first term of the cost function represents the cost due to the deletion of connections. The first term represents the loss of link importance, while the second term represents the connectivity gain brought to the entire network by the newly added links. and These are the connectivity weights and connection importance weights. Consistent with game theory, a smaller loss function is better. It should be noted that this only takes effect when both endpoints simultaneously make the decision to establish a link; therefore, for drones… n The cost function is the current drone's decision and the decisions of other drones. The function.

[0028] As mentioned earlier, by constructing the cost function (The existence of a Nash equilibrium solution for this game problem can be proven by a global situation function that satisfies a monotonic relationship. Based on this, a distributed network topology control strategy can be designed through distributed link updates and iterations of a single node, which can guarantee convergence to Nash equilibrium.)

[0029] Existing technologies only focus on time of arrival in UAV-assisted communication systems. This invention proposes a task-oriented design scheme in which the user equipment can adopt a continuous offloading strategy, meaning that some tasks can be executed locally, while the remaining tasks can be offloaded to a remote computing unit for parallel execution. Therefore, to minimize task freshness, a good balance needs to be struck between local execution latency and task transmission latency, which is even more challenging under limited energy conditions, as the UAV's relay service allocation, dynamic trajectory, relay power, task allocation, and local CPU computing frequency and transmission power are closely related.

[0030] Existing technologies assume that tasks should be transmitted or executed within a single time slot. This invention considers a more general case where tasks can be offloaded and computed across time slots. Therefore, arrival time is closely related to instantaneous transmission rates and local execution capabilities, factors that are unpredictable and hinder mathematical expression from a global perspective. Thus, this invention transforms the problem of minimizing the evolution of arrival time into a problem of minimizing residual tasks, thereby decomposing UAV relay service allocation and trajectory optimization into a series of subproblems that are easily solved using existing DRL and SCA methods.

[0031] Multi-agent systems are a novel distributed computing technology. In a multi-agent system, each agent learns and improves its policy by interacting with the environment to obtain rewards, thus acquiring the optimal policy for that environment. This process is called multi-agent reinforcement learning. In single-agent reinforcement learning, the environment in which the agent operates is stable and unchanging, but in multi-agent reinforcement learning, the environment is complex and dynamic.

[0032] In a multi-agent system, there are at least two agents with certain relationships, such as cooperation, competition, or both. In this example, the agents both cooperate and compete with each other, i.e., they compete for spectrum resources, leading to access collision problems.

[0033] In a multi-agent system, the reward each agent receives is not only related to its own actions, but also to the actions of other agents; that is, agents influence each other.

[0034] To overcome the convergence difficulties caused by high-dimensional action spaces, such as Figure 4 As shown, this invention proposes a novel Multi-Agent DRL (MADRL) method, combined with the SCA algorithm, to solve this non-convex problem. This method pre-solves the service allocation problem to reduce training complexity and employs the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) method to avoid local suboptimal solutions by calculating the contribution of each UAV and guiding more robust policy updates. Subsequently, other variables such as UAV trajectories and task offloading are handled using convex optimization theory methods. Table 1 shows a schematic diagram of the solution process proposed in this invention.

[0035] In this invention, each iteration in the iterative solution of the objective function includes: taking the user's computing task and the drone's trajectory as the state, the one-to-one service relationship between the drone and the user as the action, and the negative value of the objective function as the reward, and obtaining the service allocation instruction matrix based on reinforcement learning; and sequentially optimizing the task offloading ratio matrix, the user's computing power matrix, the transmission power matrix, and the global trajectory vector matrix of the human and the drone based on convex optimization methods. After each iteration, the objective function value is calculated twice. If the reduction in the two objective function values ​​is less than the tolerance, the calculation is terminated; otherwise, the iteration continues until the condition is met.

[0036] Table 1

[0037] This invention primarily addresses the AoT minimization problem in multi-UAV-assisted MEC systems, where UAVs act as mobile aerial relays, forwarding some tasks generated by user equipment to the MEC server, thereby fully utilizing the local communication capabilities of user equipment and the abundant computing resources of the MEC. Based on this, relay service allocation, user computing frequency, UAV trajectory, and transmission and relay power of user equipment and UAVs can be jointly optimized, while simultaneously satisfying energy constraints.

[0038] To address the networking requirements of UAV swarms, a multi-objective, multi-constraint dynamic model is constructed. For the matching optimization problem of discrete variables such as service indication relationships, convex optimization and other methods are used to solve the problem first and ensure equilibrium convergence, thereby achieving spectrum efficiency optimization under low interruption probability and realizing topology control and cooperative routing of UAV self-organizing networks.

[0039] This invention performs better in multi-agent systems. When the number of secondary users increases, the learning performance of traditional machine learning methods, such as Q-Learning and DQN, will decrease. However, due to the centralized training and distributed execution characteristics of MADDPG, the probability of a trained user (GU) finding an idle UAV in one go can reach 90%, effectively reducing the probability of collisions with UAVs and other UAVs, while also reducing communication latency and perception overhead.

[0040] like Figure 1 As shown, the system has M One ground user, NA UAV equipped with an omnidirectional antenna and a mobile edge server (MEC) with strong computing power. The UAV generates computationally intensive tasks. Due to the distance or congestion, the MEC with superior processing capabilities cannot communicate directly with the UAV. The UAV can choose the following methods to process the tasks: (1) use its own computing power to process the tasks, (2) use the UAV as a relay to offload the tasks to the MEC for processing, and (3) perform local computing and the second method simultaneously. In most cases, this method is chosen for task processing.

[0041] This scenario considers a discrete-time system, with a total time... Divided into T There are 1 time slot, and the duration of each time slot is 1. When the first m Each GU generates a computing task. Specifically, it is expressed as ,in, Indicates the bit size of the task. This represents the number of CPU cycles required to complete each bit of the task. This refers to the task generation time. This invention uses a centralized control architecture, allocating a portion of the task's computational data to local computing and offloading the remaining data to the MEC server via UAV relay for computation. This vector represents the proportion of tasks that are offloaded to the MEC, indicating the task splitting ratio. and These are the local computed data volume and the unloaded data volume, respectively.

[0042] When drones serve users Represents a binary service indicator variable. Indicates the first m The user was the first n A drone service This indicates that the drone is hovering and not serving any users. For a single user, only one drone can serve them per time slot; conversely, for a single drone, only one user can be served per time slot. (C1) (C2) (C3) in, Indicates a collection of drones. This represents a set of users.

[0043] use This represents the maximum flight speed of the drone. To prevent potential collisions, the minimum safe distance between drones is defined as... So, what is the maximum distance a drone can travel within a time slot? for Since drones are powered by built-in batteries or fuel, they must periodically return to fixed locations to recharge or refuel, thus requiring a given initial location for the drone. and final position . Represented as the drone in the The three-dimensional Cartesian coordinate position of each time slot. and Let represent the positions of the user and the MEC, respectively, assuming they are stationary. The kinematic constraints of the UAV are as follows: (C4) (C5) (C6) in, Indicates the location of other drones. T This represents the last time slot in the total time. The minimum safe distance constraint between drones is... , Indicates the first The position coordinates of the UAV in the t-th time slot Indicates the first The first drone t The location coordinates of each time slot Indicates except the first The serial number of the drone other than the drone itself; Typically, the calculated data (i.e., the MEC's ​​calculation results) is much smaller than the offloading task itself; therefore, the downlink communication latency and energy consumption (i.e., the communication latency and energy consumption of the MEC sending data back to the GU) can be ignored. In this case, only the uplink transmission latency (i.e., the transmission latency from the UAV to the MEC) needs to be considered. Assuming that the UAV and MEC use frequency division multiple access (FDMA) technology for transmission, and there is no interference between the UAV and GU, therefore, in the... t The first time slot, the first m The first GU and the first n The first UAV and the first n The channel gain between each UAV and MEC can be expressed as follows: (4) (5) in, At a distance of Channel power gain at that time For carrier frequency, It is the speed of light. Indicates the first t The first time slot m The first GU and the first n The distance of a UAV, Indicates the first t The first time slot n The distance of each UAV.

[0044] The first can be obtained from the channel gain. m The GU to the n The first UAV and the first n The task unloading rate from UAV to MEC is: (6) (7) in, For the bandwidth of each channel, Indicates the first m The transmit power of each GU For the first n The transmit power of each UAV, This is the noise power in the channel, and the transmit power of the GU and UAV should be subject to the peak value ( and Limitations: (C7) (C8) Because users tend to dynamically adjust voltage and frequency to save energy, the user's actual CPU frequency is at the [missing information - likely a specific frequency range]. t Each time slot is represented as And there is an upper limit. : (C9) In the t The data size for a local processing task in a time slot can be expressed as: (8) in, Calculate the frequency locally.

[0045] Similarly, in the t The amount of data offloaded to MEC via UAV in each time slot can be expressed as: (9) Let the two-dimensional indicator variable Indicator variable for whether a task is generated: (10) Set as task It only completes after both local computation and remote transmission are finished. Therefore, in the... t The remaining tasks for local computation and offloading transmission in each time slot can be represented as follows: (11) (12) Two-dimensional indicator variables and As an indicator variable for whether the local computation and offloading transfer tasks have been completed: (13) (14) To describe the freshness of task execution, this invention proposes the concept of AoT (Age of Task), which describes the freshness of information. GU m In the t Tasks generated in each time slot The freshness of the task increases over time and resets to 0 once the task is completed. The evolution over time can be written as: (15) like Figure 2 As shown, each GU in Generate a task and in Execution complete. The thickest line represents the total AoT of all users over time. (The remaining text appears to be incomplete and possibly contains errors.) Taking a time slot as an example, when GU1's task has just finished, but GU2 and GU3's tasks are still being processed, while other GUs' tasks have not yet been generated. Therefore, the instantaneous AoT is... Therefore, in the case of multiple GUs, in order to minimize AoT, priority should be given to processing those GUs whose tasks have taken time but have not yet been completed, as this dominates the size of the total AoT.

[0046] The user's computing energy consumption and transmission energy consumption can be calculated separately as follows: (16) (17) in, and Indicates the first m Different positive coefficients for each user's CPU model; The communication energy consumption of the UAV as a relay to offload tasks to the MEC can be expressed as: (18) Considering the small quadcopter drone in the system, the drone's propulsion energy consumption is in the first... t Each time slot can be calculated as: (19) in, Indicates horizontal flight speed. It is the tip speed of the rotor blades, and This represents the average speed while hovering. Additionally, This represents the fuselage drag ratio. This represents the rotor stiffness, where R is the rotor radius, in meters. and These are the number of blades and the airfoil chord length, respectively. Additionally, This indicates the rotor disk area, in square meters. It is air density, the unit is... . and These are two constants representing the blade profile power and the blade induced power of the UAV, respectively. It is the blade angular velocity, expressed in radians per second. This is the profile drag coefficient. This is the incremental correction factor for the induced power. The value is the weight of the drone, expressed in Newtons.

[0047] Generally, remote MEC servers integrated with the BS possess enormous computing power and sustainable power supplementation. Therefore, it is reasonable to ignore the computational latency and energy consumption of remotely executed tasks on the MEC server. Thus, the total energy consumption constraint of a multi-UAV assisted MEC system (i.e., the upper limit constraint of the total energy consumption of the centralized control architecture) is: (C10) in, This represents the system's maximum total energy consumption. Indicates the time slot sequence number, T Indicates the largest time slot sequence number. Indicates the first t The computational energy consumption of the m-th user in a given time slot. Indicates the first t The first time slot m Transmission energy consumption per user M This represents the largest user ordinal number. N This represents the ordinal number of the largest drone. Indicates the first t The first time slotn The communication energy consumption of a drone acting as a relay to offload tasks to the MEC server. Indicates the first t The first time slot n The simplified conversion of propulsion energy consumption for a single drone, and the upper limit of total energy consumption of a centralized control architecture.

[0048] Considering the service time of drones Minimizing the Information Age (AoT) of GU tasks (i.e., the objective function established by minimizing the task freshness of all users in a centralized control architecture) can be expressed as: (20) (C11) (C12) (C13) in, and These serve as the service allocation indication matrix and the task offload ratio matrix in the nth time slot, respectively. This is the computational capability matrix of the GU; This is a combined transmission power matrix for GU and UAV, where ; Let be the global trajectory vector matrix of the UAV.

[0049] because Since it is unrelated to other variables in P1, when solving the service allocation indicator matrix based on reinforcement learning, the constraints are the one-to-one correspondence between UAVs and users and the total energy consumption limit of the centralized control architecture. The sub-problems of the relevant service allocation variables can be: (twenty one) P2 remains a MINLP problem due to the discontinuity and non-differentiability of the binary vector A, especially in the dynamic motion environment of UAVs. Therefore, this embodiment extends DDPG to the multi-agent domain, employing the MADDPG method, allowing each user device to act as an independent agent, making appropriate action mappings given partial observations from the system. Specific steps can be found in [reference needed]. Figure 3 .

[0050] The states, actions, and rewards in this embodiment are as follows: state: This refers to the task vectors of all GUs and the trajectories of all UAVs. All states constitute the state space. .

[0051] action: , that is GU m Service allocation on time slot t, where All actions constitute the action space. ; award: The reward is set to a negative value of the AoT for each GU across all time slots.

[0052] Specifically, the methods include: Obtain each of the M GUs m The environmental parameters at the start of the current time frame, m=1, 2, ..., M, include each GU m Observation S m ; Each GU m At the start of the current time frame, input the environmental parameters into the deterministic policy deep gradient MADDPG model; Obtain the output of each GU from the MADDPG model. m The service allocation strategy for time frames is to select an idle drone as a relay service user. Will GU m The perception results are fused into the global state S m (t), and global action A m (t), reward r m (t), the state S at the next time step m (t+1) is sent to the experience replay buffer of the MADDPG model; The MADDPG model is trained using tuples (S, A, R) consisting of states, actions, and rewards, where state S includes each GU. m Integrating the perception results from its partners, Action A includes every GU m In the perception policy of the current time frame, the reward R is based on each GU. m The reward obtained from the action taken.

[0053] On the other hand, the MADDPG method is used to optimize the service allocation indicator variable. Since this variable is independent of the optimization problem, it can be solved in advance before optimizing the other optimization variables. After determining the service allocation of the UAV, the computation frequency and offloading ratio of the GU can be further optimized using convex optimization theory, and then the transmission power and trajectory can be further optimized using convex optimization theory.

[0054] By inputting the environmental parameters of each agent at the start of the current time frame into the deterministic policy deep gradient MADDPG model, which includes the observations of each agent at the start of the current time frame, the MADDPG model outputs each GU (Guided User Unit). m The service allocation strategy for each time frame selects an idle drone as a relay service user. Each agent (in this invention, an agent refers to a user) performs sensing and access according to a defined strategy in each time frame. Furthermore, based on the original MADDPG algorithm, convex optimization methods are further utilized to optimize the user's computation frequency, offload ratio, transmission power, and drone trajectory, effectively improving task freshness and reducing task age.

[0055] Since AoT will continue to increase linearly without any newly generated task updates, once service associations are determined, the problem of minimizing AoT in P1 can be equivalently transformed into minimizing the remaining tasks in each time slot: (twenty two) After determining the service allocation for the UAVs, the computation frequency and offloading ratio of the GUs can be further optimized using convex optimization theory. After optimization, when optimizing the task offloading ratio matrix using the convex optimization method, the objective function is transformed into a subproblem related to the computation frequency and offloading ratio (P3.1). (twenty three) This sub-problem is constrained by the user's CPU frequency limit and the total energy consumption limit of the centralized control architecture; This represents the task unloading ratio matrix. Represents the user's computing power matrix. Indicates the t-th time slot. m The remaining amount of tasks computed locally by each user. Indicates the t-th time slot. m The remaining amount of tasks that each user has unloaded and transferred.

[0056] It should be noted that, according to equations (5) and (8), Regarding B and F being linear, while The function is linear with respect to B and independent of F. Therefore, the objective function of P3.1 is convex with respect to both B and F. (C9) is a linear constraint on F, and (C13) is a linear constraint on B. (C11) and (C12) are convex constraints because the sublevel set of a convex function is convex. (C10) is obtained through equation (11). Regarding the convex constraint of F, because and The coefficients are positive. Therefore, P3.1 is a convex problem and can be solved efficiently using existing convex toolboxes such as CVX.

[0057] Based on the variables already optimized above, when optimizing the transmission power matrix, subproblem P3.1 is transformed into subproblem P3.2, which is the optimization of power control for both the GU and UAV: (twenty four) This sub-problem is constrained by the peak transmit power of the UAV, the peak transmit power of the user, and the total energy consumption limit of the centralized control architecture. This represents the transmission power matrix. Specifically, (C7), (C8), and (C10) are about... and Linear constraints.

[0058] It should be noted that P can only affect the objective function of P3.2. and It is no longer related to P, but and equivalence( Indicates the preceding text (For simplification, the abbreviation is omitted; similarly, other related symbols appearing later are abbreviations of their corresponding symbols in the preceding text.) and For respectively and It is convex. The maximization of a set of functions containing constants and convex functions is still convex. Therefore, the objective function is convex with respect to P. Similarly, the sublevel set of convex functions is also convex, thus making (C12) a convex constraint. Therefore, P3.2 is a convex problem with respect to P and can also be solved by CVX.

[0059] With all other variables fixed, when optimizing the transmission power matrix using the convex optimization method, subproblem P3.2 is transformed into subproblem P3.3 of UAV trajectory optimization: (25) In other words, this subproblem is constrained by the maximum distance a drone can travel in a time slot, the minimum safe distance between drones, and the total energy consumption limit of a centralized control architecture.

[0060] Minimizing the objective function P3.3 is equivalent to minimizing Because of the optimization of trajectory Q and Irrelevant. Based on the above analysis, it is noted that… and Related, although in rate R about It is convex, but in composite functions, -max{.} is a decreasing convex function, therefore about The unevenness or concavity of the surface cannot be determined.

[0061] This embodiment will optimize the drone's trajectory by following four steps: (1) To facilitate calculation, an auxiliary vector is introduced. and , and It can be restated as: (26) (27) in, .

[0062] because and To each and It is convex, therefore it can be passed through any given feasible point. and We apply the first-order Taylor expansion to obtain their global lower bounds.

[0063] (28) (29) because and about The second derivative of is always positive, therefore it is convex. and for It is a concave function, therefore, The upper bound can be obtained from the following: (30) (2) Regarding (C12), according to equation (30), it can be rewritten in the form of a convex constraint (C12.a): (C12.a) (3) Regarding (C5), although It is convex, but the resulting set is not a convex set because, in general, the hyperlevel set of a convex quadratic function is not convex. Since any convex function at any given feasible point... The first-order Taylor expansion at any point is a global lower bound for the function, therefore: (31) Therefore, (C5) can be accessed through its global lower bound. To strengthen, the lower bound is and The affine function. Therefore, Transform it into (C5.a) as the new minimum safe distance constraint between drones, i.e., the enhanced constraint is: (C5.a) in, express The global lower bound, express Given feasible points.

[0065] (4) Regarding constraint (C10), The second item about It is non-convex. This can be addressed by introducing slack variables. ,in The following equation holds true: (32) Relax it into an inequality constraint: (C13) Then It can then be re-represented as: (33) because and For respectively and It is a convex function, therefore we have: (34) for and It is a joint affine function, for any given feasible point. have: (C13.a) Therefore, constraint (C10) can be transformed into (C10.a): (10.a) In summary, P3.3 can be approximately transformed into a convex problem, denoted as P3.3.a, which can be efficiently solved using CVX: (35) (36) The present invention also included verification embodiments, with specific scenarios and related results shown below. Twelve GUs were deployed within a 2000m × 2000m area. Four UAVs collaboratively collected and unloaded data from ground users, maintaining fixed initial and final positions during flight before transmitting the data to the MEC. All UAVs flew at a fixed altitude of 150 meters. The MEC was located at (0,0) at an altitude of 140 meters.

[0067] The focus of this invention is to minimize the instantaneous AoT in each time slot. However, in performance evaluation, to compare the task freshness of the proposed Min-AoT method with the traditional Max-Rate method, the average peak AoT is generally used: (37) The Max-Rate method is as follows: (38) The system is configured with four drones, each serving three users. Figure 5 and Figure 6 Comparing the trajectories of the two schemes, it can be seen that in the Max-Rate scheme, the drones fly to nearby users sequentially, regardless of the task generation time ℓ. 𝑚 In contrast, the Min-AoT solution prioritizes drone services. 𝑚 Smaller users, i.e., self-generation time ℓ 𝑚 For users whose tasks have not yet been completed, the order in which they are served will differ from that of the Max-Rate scheme (the service order of user 9 and user 8). The Max-Rate scheme prioritizes user 8, while the Min-AoT scheme prioritizes user 9. Figure 7 As can be seen, the Min-AoT scheme significantly optimizes PAOT, not only because it prioritizes minimizing AoT, but also because it offloads more data to the MEC for computation. Figure 8 The results show that in the Min-AoT scheme, the drone follows the mission generation time ℓ 𝑚 Prioritizing speed over distance results in a greater cumulative travel distance, thus consuming more propulsion and transmission energy.

[0068] To ensure that computationally intensive tasks generated by ground users can be offloaded and executed promptly in the UAV-assisted mobile edge computing system, this embodiment uses the concept of AoT (Aspect-Oriented Time) to quantify task freshness—that is, the time elapsed from the generation of the latest task by the source node to its completion. Based on a continuous offloading strategy, each user can autonomously handle some tasks, while the rest are synchronously executed remotely to the MEC server via UAV relay. By transforming the minimization of AoT over multiple time slots into the problem of minimizing remaining tasks, this invention further proposes a MADDPG-SCA joint optimization algorithm. Specifically, each user interacts with the environment as an independent task agent within the MADDPG framework, intelligently planning feasible service allocation schemes between each unit and the UAV. Other closed-loop coupling optimization variables (such as task partitioning, user computation frequency, UAV trajectory, and energy consumption) are solved alternately using the SCA method. Numerical results show that the Min-AoT algorithm exhibits significant advantages in achieving minimum AoT.

[0069] The present invention also discloses a UAV networking calculation offloading device based on task age, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement any of the above methods.

[0070] The present invention also discloses an embodiment that provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0071] The present invention also provides a computer program product that, when run on a data storage device, enables the data storage device to implement the steps in the above-described method embodiments.

[0072] If the integrated unit module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a storage device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

Claims

1. A method for calculating unloading in UAV networking based on task age, characterized in that, This method is applied in a centralized control architecture, which includes several users, several drones, and a MEC server; the method includes the following steps: Obtain user-generated computing tasks; the computing tasks may be computed locally by the user, offloaded to the MEC server via the drone for computation, or computed jointly by the user and the MEC server. The user's task freshness is generated based on the computational task; the task freshness increases with the increase of time slots. The objective function is established with the one-to-one service relationship between UAVs and users, the peak transmission power of UAVs, the peak transmission power of users, the upper limit of CPU frequency of users, and the upper limit of total energy consumption of centralized control architecture as constraints, and the minimum task freshness of all users in the centralized control architecture. The objective function is solved iteratively to obtain the service allocation indication matrix, task offloading ratio matrix, user computing power matrix, transmission power matrix, and UAV global trajectory vector matrix for each time slot.

2. The method for calculating unloading in UAV networking based on task age as described in claim 1, characterized in that, Each iteration in the iterative solution of the objective function includes: The service allocation instruction matrix is ​​obtained by solving the problem based on reinforcement learning, with the user's computing task and the drone's trajectory as the state, the one-to-one service relationship between the drone and the user as the action, and the negative value of the objective function as the reward. The task offloading ratio matrix, user computing power matrix, transmission power matrix, and UAV global trajectory vector matrix are optimized sequentially using the convex optimization method.

3. The method for calculating unloading in UAV networking based on task age as described in claim 2, characterized in that, When solving the service allocation instruction matrix using reinforcement learning, constraints are imposed on the one-to-one service relationship between UAVs and users and the total energy consumption limit of the centralized control architecture.

4. The method for calculating unloading in UAV networking based on task age as described in claim 3, characterized in that, When optimizing the task unloading ratio matrix and the user's computing power matrix using convex optimization methods, the objective function is transformed into subproblems P3.1: , This sub-problem is constrained by the user's CPU frequency limit and the total energy consumption limit of the centralized control architecture; This represents the task unloading ratio matrix. Represents the user's computing power matrix. Indicates the t-th time slot. m The remaining amount of tasks computed locally by each user. Indicates the t-th time slot. m The remaining amount of tasks that each user has unloaded and transferred.

5. The method for calculating unloading in UAV networking based on task age as described in claim 4, characterized in that, When optimizing the transmission power matrix using the convex optimization method, subproblem P3.1 is transformed into subproblem P3.2: , This sub-problem is constrained by the peak transmit power of the UAV, the peak transmit power of the user, and the total energy consumption limit of the centralized control architecture. This represents the transmission power matrix.

6. The method for calculating unloading in UAV networking based on task age as described in claim 5, characterized in that, When optimizing the global trajectory vector matrix of a UAV using convex optimization methods, subproblem P3.2 is transformed into subproblem P3.3: , The sub-problem is constrained by the maximum distance a drone can travel in a time slot, the minimum safe distance between drones, and the total energy consumption limit of the centralized control architecture.

7. The method for calculating unloading in UAV networking based on task age as described in claim 6, characterized in that, The minimum safe distance constraint between drones is , Indicates the first The position coordinates of the UAV in the t-th time slot Indicates the first The position coordinates of the UAV in the t-th time slot Indicates except the first The serial number of the drone other than the drone itself; Will Convert to As a new minimum safe distance constraint between drones; among which... express The global lower bound, express Given feasible points.

8. The method for calculating unloading in UAV networking based on task age as described in claim 6, characterized in that, The total energy consumption limit constraint for a centralized control architecture is: , in, Indicates the time slot sequence number, T Indicates the largest time slot sequence number. Indicates the first t The computational energy consumption of the m-th user in a given time slot. Indicates the first t The first time slot m Transmission energy consumption per user M This represents the largest user ordinal number. N This represents the ordinal number of the largest drone. Indicates the first t The first time slot n The communication energy consumption of a drone acting as a relay to offload tasks to the MEC server. Indicates the first t The first time slot n A simplified conversion formula for the propulsion energy consumption of a drone. This indicates the upper limit of total energy consumption for a centralized control architecture.

9. A UAV networking calculation offloading device based on task age, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-8.