Unmanned aerial vehicle assisted edge computing multi-objective optimization method based on gradient conflict detection
By employing a gradient conflict detection mechanism and gradient projection correction, the problem of coordinating the optimization of information age and energy consumption in UAV-assisted mobile edge computing systems is solved, achieving a synergistic trade-off between information timeliness and energy efficiency, and improving the system's operational performance and learning stability.
Patent Information
- Application Number
- CN202610846320.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-25
AI Technical Summary
In existing UAV-assisted mobile edge computing systems, there is a lack of a stable and effective unified modeling and solution framework for the coordinated optimization of information age and energy consumption. Inconsistent update directions of different objectives during multi-objective learning lead to training instability and unsatisfactory performance trade-offs.
A gradient conflict detection mechanism is constructed. The gradient conflict between information age and energy consumption target is detected in the UAV-assisted edge computing system by reinforcement learning algorithm. The conflict gradient components are eliminated by gradient projection correction, so as to achieve a synergistic trade-off between information timeliness and energy efficiency.
While ensuring timely information updates, unnecessary energy consumption is suppressed, system performance is improved, information lag is reduced, and policy learning efficiency is increased, thus achieving a better information age-energy consumption trade-off.
Smart Images

Figure CN122640752A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile edge computing networks and UAV-assisted communication technology, and specifically relates to a multi-objective optimization method for UAV-assisted edge computing based on gradient conflict detection. Background Technology
[0002] With the rapid development of the Internet of Things (IoT), the Internet of Vehicles (IoV), and mobile internet services, the amount of data and computing demands generated by ground terminal equipment continue to grow. Limited by the battery capacity, processor performance, and heat dissipation conditions of ground terminal equipment, executing complex tasks entirely locally can easily lead to increased computational latency and energy consumption, making it difficult to meet real-time requirements. Mobile edge computing, by offloading computing and storage capabilities to the network edge, can reduce task processing latency to a certain extent and alleviate the computational pressure on ground terminal equipment, thus becoming an important technical path to support low-latency services. In actual deployments, scenarios such as terrain obstruction, insufficient base station coverage, disaster emergencies, or temporary hotspots can lead to insufficient service capabilities or unstable link quality for ground edge nodes. Drones, with their mobility, rapid deployment, and significant line-of-sight link advantages, can serve as aerial access points or aerial edge nodes, providing users with communication relay and computation offloading support, thereby improving the coverage and quality of edge services. Therefore, drone-assisted mobile edge computing is gradually becoming an important direction for academic research and engineering applications.
[0003] In UAV-assisted mobile edge computing systems, task offloading and resource allocation for ground terminal equipment typically require comprehensive consideration of factors such as wireless transmission, edge computing, queue backlog, and dynamic changes in network topology. Traditional research often uses system latency, throughput, or task completion rate as performance indicators, improving system efficiency by optimizing offloading decisions, transmission power, computing resource allocation, and variables such as UAV position / trajectory. However, for services requiring state updates and real-time perception and control (such as environmental monitoring, target tracking, vehicle-road cooperation, and industrial monitoring), latency alone is insufficient to accurately characterize the "freshness" of information. Information age measures the time difference between the information received by the receiver and the latest state, reflecting the timeliness of data updates, and thus has more direct engineering significance in the aforementioned applications. Existing technologies have developed task scheduling, sampling, and transmission optimization methods with information age as the core indicator, but most works tend to optimize the single information age indicator under fixed assumptions or only perform local optimization on a small number of control variables. On the other hand, both UAV platforms and mobile ground terminal equipment face energy constraints: the ground terminal equipment involves local computing and transmission energy consumption, while the UAV side involves flight energy consumption and communication / computing energy consumption. Pursuing only low information age or low latency often requires higher transmission power, more frequent data updates, or more aggressive scheduling strategies, which significantly increases system energy consumption and reduces continuous service capability. Therefore, in practical systems, a synergistic trade-off between information timeliness and energy consumption is needed, constructing a joint optimization framework of AoI (Information Age) and energy consumption to meet the engineering requirement of "timely updates and acceptable energy efficiency".
[0004] For the aforementioned joint optimization problem, due to the dynamic and stochastic nature of the system state and the fact that decision variables are often high-dimensional and strongly coupled, data-driven methods such as reinforcement learning are used to solve complex online decision problems. However, in multi-objective optimization scenarios, the optimization directions of different objectives for policy updates may be inconsistent. If simple weighted summation or fixed-weight methods are used for learning, problems such as mutual constraints between objectives during training, unstable convergence, or getting trapped in suboptimal compromise solutions can easily occur, making it difficult for the algorithm to simultaneously meet the requirements of system information timeliness and energy efficiency.
[0005] Therefore, existing technologies in UAV-assisted mobile edge computing systems still face the following shortcomings: First, there is a lack of a stable and effective unified modeling and solution framework for the coordinated optimization of information age and energy consumption; second, when the update directions of different objectives are inconsistent during multi-objective learning, it can easily lead to training instability and unsatisfactory performance trade-offs. To address these issues, it is necessary to propose a joint optimization method for UAV-assisted mobile edge computing that can achieve a coordinated trade-off between information age and energy consumption in dynamic environments and improve the stability of multi-objective learning. Summary of the Invention
[0006] To address the aforementioned problems, the present invention aims to provide a multi-objective optimization method for UAV-assisted edge computing based on gradient conflict detection.
[0007] To achieve the above objectives, the UAV-assisted edge computing multi-objective optimization method based on gradient collision detection provided by the present invention includes the following steps performed in sequence:
[0008] 1) Construct a drone-assisted mobile edge computing network system that includes drones equipped with edge servers and ground terminal equipment. This system is used to characterize the task arrival, link status changes and edge computing service process of ground terminal equipment, and initialize system parameters and constraints to provide a basic environment for subsequent communication modeling, computation modeling and optimization decision-making.
[0009] 2) Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model is established according to the positional relationship between the UAV and the ground terminal equipment to obtain the link transmission rate during the task upload process; based on the above link transmission rate, a local computing model and an edge computing model are established to obtain the local computing latency and the total offloading execution latency; based on the communication process and the computing process, a local computing energy consumption model and an edge computing energy consumption model are established to characterize the energy consumption generated during the task execution process.
[0010] 3) Based on the communication model and computing model established in step 2), obtain the task execution status and task completion time information, and construct an information age model accordingly. Then, decide whether to update the timestamp in the information age model based on whether the task is successfully completed, so as to characterize the freshness and timeliness of the system information.
[0011] 4) Based on the system parameters obtained in step 1), the energy consumption model established in step 2), and the information age model constructed in step 3), design a multi-objective reward function and construct a multi-objective joint optimization model of information age and energy consumption to achieve a synergistic trade-off between the timeliness of system information and energy consumption.
[0012] 5) Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the reinforcement learning algorithm is used to learn the task offloading and resource allocation strategy. During the strategy update process, the gradient conflict between the information age objective and the energy consumption objective is detected, and the conflict gradient components are eliminated by gradient projection correction, thereby improving the stability and convergence performance of the multi-objective optimization process, and finally obtaining the optimal strategy of the joint operation of UAV trajectory, task offloading and transmission scheduling.
[0013] In step 1), the method for constructing a drone-assisted mobile edge computing network system consisting of a drone equipped with an edge server and ground terminal equipment, and initializing system parameters and constraints, is as follows:
[0014] Construct a drone-assisted mobile edge computing network system consisting of one drone equipped with an edge server and N ground terminal devices; represent the positional relationship between the drone and the ground terminal devices in a three-dimensional Cartesian coordinate system: define the set of positions Q of the ground terminal devices. users = {Q u (1), Q u (2),…,Q u (N)}, where Q u (i) = {x u (i),y u {i, 0} represents the fixed horizontal coordinates of ground terminal device i. Each ground terminal device can establish a wireless communication connection with the UAV for uploading and unloading mission data. The position of the UAV in time slot t is Q. uav (t) = {x uav (t),y uav (t),H}, where H is the flight altitude of the UAV;
[0015] Discretize the system runtime into a set of time slots. Where T is the preset total number of time slots, each time slot has a length of τ, and it is assumed that the completion time of each task, whether executed locally or offloaded, does not exceed one time slot, to meet the requirements of time slotted decision-making and online optimization; the task set of the ground terminal equipment in time slot t is defined as S(t) = {S1(t), S2(t),…, S…} N (t)}, where S i (t) represents the task generated by ground terminal device i in time slot t, used to characterize the characteristics of terminal task arrival and task size change over time; when the ground terminal device chooses to offload the task to the edge server on the UAV, due to the limited capacity of the wireless link, only a portion of the task can be uploaded in a single time slot. Therefore, tasks that are not scheduled for upload / not completed for upload wait in the transmission queue on the UAV side, thus forming a service process of "task arrival - offload request - task queue waiting - scheduling transmission - edge computing".
[0016] The system parameters include the number of terminals N, the flight altitude of the UAV H, the time slot length τ, the total number of time slots T, and the terminal task arrival / task data scale parameters; the constraints include wireless link capacity constraints, task completion time limit assumptions, and the remaining energy / energy budget of the ground terminal equipment, to ensure that subsequent offloading and scheduling decisions are executed within the feasible domain.
[0017] In step 2), the UAV-assisted mobile edge computing network system constructed in step 1) establishes a communication model based on the positional relationship between the UAV and the ground terminal equipment to obtain the link transmission rate during the task upload process; based on the above link transmission rate, a local computing model and an edge computing model are established to obtain the local computing latency and the total offloading execution latency; the method for establishing the local computing energy consumption model and the edge computing energy consumption model based on the communication process and the computing process is as follows:
[0018] Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model between ground terminal equipment and UAV upper edge server is constructed.
[0019] The wireless transmission link is affected by both small-scale fading and large-scale path loss. Its instantaneous channel gain is expressed as:
[0020] ;
[0021] Among them, g i,u L is a random variable that follows a Rayleigh distribution and is used to characterize small-scale fading. i,u The large-scale path loss described by the logarithmic distance model is expressed as:
[0022] ;
[0023] Where parameters A and B are the constant term and attenuation factor in the path loss, respectively, d i,u Let i be the distance between the drone i and the ground terminal equipment u.
[0024] Under the above channel conditions, based on the above channel gain, the link transmission rate is given by Shannon's formula:
[0025] ;
[0026] Where B is the bandwidth allocated to the link, and P... i Let σ be the transmit power of UAV i. 2 Noise power;
[0027] Based on the system parameters obtained in step 1), a local computing model is constructed:
[0028] Let the amount of task data generated by ground terminal equipment u in time slot t be D. u (t), the task computation density is C, and the CPU frequency of the ground terminal equipment is f. u The local computation latency is:
[0029] ;
[0030] Based on the above link transmission rate, an edge computing model is constructed: when the ground terminal device u offloads the task to the edge server of the drone for execution in time slot t, its transmission latency is:
[0031] ;
[0032] Assume the CPU frequency of the edge server on the drone is f. i The execution latency of the task on the edge server is then:
[0033] ;
[0034] Therefore, the total execution delay for unloading is:
[0035] ;
[0036] In terms of energy consumption, a local computing energy consumption model is established, represented as follows:
[0037] ;
[0038] Where, k u The terminal energy consumption coefficient is determined by the processor architecture.
[0039] Establish an edge computing energy consumption model: decompose the energy consumption caused by terminal offloading into two parts: transmission energy consumption and edge execution energy consumption; when the ground terminal device u performs task offloading, its transmission energy consumption is determined by the transmission power of the ground terminal device u. The transmission delay mentioned above determines:
[0040] ;
[0041] When a task is executed on an edge server on a drone, its execution energy consumption is expressed as follows:
[0042] ;
[0043] Where, k i The energy consumption coefficient of the edge server on the drone;
[0044] Therefore, the energy consumption model for edge computing is as follows:
[0045] ;
[0046] in, This represents the total energy consumption for unloading execution.
[0047] In step 3), the method for obtaining task execution status and task completion time information based on the communication model and computing model established in step 2), constructing an information age model accordingly, and then determining whether to update the timestamp in the information age model based on whether the task was successfully completed is as follows:
[0048] Let τ be the timestamp corresponding to the task being received and processed in time slot t. u (t), then the information age model of the ground terminal equipment u in time slot t is defined as:
[0049] ;
[0050] in, This indicates the information age of the ground terminal equipment u in time slot t;
[0051] When a new task generated by ground terminal equipment u in time slot t is successfully received and processed, the timestamp in the information age model is updated to the latest timestamp, denoted as AoI. u (t+1)=1; When the new task is not completed, keep the original timestamp unchanged, represented as AoI. u (t+1)= AoI u (t)+1.
[0052] In step 4), the method for designing a multi-objective reward function and constructing a multi-objective joint optimization model of information age and energy consumption based on the system parameters obtained in step 1), the energy consumption model established in step 2), and the information age model constructed in step 3) is as follows:
[0053] Introducing system state s t ∈S、Action a t ∈A and the strategy π(a|s); where the system state s t System information including the current information age of ground terminal equipment, task queue status, task data volume, wireless channel status, remaining energy of ground terminal equipment, and edge server computing resource usage; Action a t This represents the task execution method and resource allocation decisions, including the choice between local or offloaded execution, the offload ratio, and the allocation strategy for communication and computing resources; policy π(a|s) represents the probability distribution of choosing an action given a state; the goal is to learn the optimal policy π. * This allows the system to achieve the maximum cumulative return throughout its entire operating cycle;
[0054] To achieve a balanced optimization of information age and energy consumption, a multi-objective reward function consisting of information age reward and energy consumption reward is constructed:
[0055] Information Age Reward: To maximize the reward for younger information ages, the information age is negatively factored out, and the information age reward is defined as follows:
[0056] ;
[0057] Energy Consumption Reward: To maximize the reward for lower energy consumption, the energy consumption reward is defined as the total energy consumption of the aforementioned unloading execution. Negative values:
[0058] ;
[0059] By linearly weighting the two types of rewards using weight coefficients ω∈[0,1], a multi-objective reward function is constructed to obtain the weighted total reward:
[0060] ;
[0061] Based on the aforementioned multi-objective reward function, a multi-objective joint optimization model is constructed that maximizes the cumulative discount return over the entire operating cycle as the optimization objective for the UAV, considering both information age and energy consumption.
[0062] ;
[0063] Where γ∈[0,1] is the discount factor, and T is the maximum number of time slots in the system.
[0064] In step 5), the multi-objective joint optimization model for information age and energy consumption constructed in step 4) uses reinforcement learning algorithms to learn task offloading and resource allocation strategies; during the strategy update process, gradient conflicts between the information age objective and the energy consumption objective are detected, and conflicting gradient components are eliminated through gradient projection correction. The method for finally obtaining the optimal joint strategy for UAV trajectory, task offloading, and transmission scheduling is as follows:
[0065] Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the task unloading decision process is modeled as a reinforcement learning process using a reinforcement learning algorithm: the UAV observes the system state s(t) in each time slot t, outputs action a(t), and obtains reward r(t) based on environmental feedback. The optimal policy π(a|s) is approximated by iteratively updating the policy network parameters θ in the reinforcement learning algorithm to improve the expected cumulative reward.
[0066] Construct corresponding target loss functions L for information age target and energy consumption target respectively. (aoi) With L (e) The objective loss function L (aoi) With L (e) The information obtained in step 4) is the age reward r. (aoi) and energy consumption reward r (e)The gradient vectors ∇ for the two objectives are obtained through the policy gradient method. During the policy network update process, the gradient vectors ∇ for the two objectives are calculated using the policy gradient method. θ L (aoi) With ∇ θ L (e) , used to characterize the influence of different optimization objectives on the update direction of the policy network parameter θ;
[0067] In the multi-objective update process, the gradient vector dot product is used to determine whether the optimization directions of different objectives are consistent: when the gradient vector dot product of two objectives is positive, it indicates that the optimization directions are consistent; when the gradient vector dot product of two objectives is negative:
[0068] ;
[0069] This indicates a gradient conflict, requiring the PCGrad strategy to correct the conflicting gradients and avoid gradient cancellation and training instability caused by direct weighting. The method involves projecting and correcting one of the gradient vectors.
[0070] ;
[0071] The correction term corresponds to the negative projection along the direction of the other gradient vector, thus eliminating the conflicting components. After the correction is completed, the conflict-free gradient vectors are weighted and synthesized.
[0072] ;
[0073] Ultimately, the parameters θ←θ-ηg are used to update the policy network, where η is the learning rate; finally, the optimal policy combining the UAV trajectory, task offloading, and transmission scheduling is obtained.
[0074] The UAV-assisted edge computing multi-objective optimization method based on gradient collision detection provided by this invention has the following beneficial effects:
[0075] Compared with existing technologies, this invention targets UAV-assisted mobile edge computing systems, incorporating information age and energy consumption into a unified optimization framework. It constructs information age rewards and energy consumption rewards, combining them through weighted coefficients, enabling the system to achieve a synergistic trade-off between information timeliness and energy efficiency under different business preferences. Furthermore, this invention introduces a gradient conflict detection mechanism during the multi-objective reinforcement learning strategy update process. It identifies the adversarial relationship between the information age objective and the energy consumption objective by determining the consistency of the target gradient direction. When a gradient conflict is detected, a projection correction mechanism is used to correct the conflicting gradients, weakening the unstable updates caused by gradient cancellation, thereby improving the stability and convergence quality of the multi-objective training process.
[0076] Based on the above mechanism, this invention can dynamically adjust the task execution mode according to the real-time status of the system (such as information age level, task data volume and computing requirements, channel transmission capacity, computing resources and energy constraints, etc.), while ensuring the timeliness of information updates and suppressing unnecessary energy consumption, thus achieving a better information age-energy consumption trade-off performance. Compared with multi-objective learning or fixed weighting methods without conflict handling, this invention can effectively reduce information lag and improve policy learning efficiency, thereby improving the overall operating performance of the UAV-assisted mobile edge computing system. Attached Figure Description
[0077] Figure 1 This is a schematic diagram of the drone-assisted mobile edge computing network system constructed in this invention.
[0078] Figure 2 This is a framework diagram of the conflict detection mechanism based on the PCGrad strategy in this invention.
[0079] Figure 3 The flowchart of the UAV-assisted edge computing multi-objective optimization method based on gradient conflict detection provided by the present invention is shown.
[0080] Figure 4 This is a Pareto curve diagram under the conflict-free detection mechanism of this invention.
[0081] Figure 5 This is a Pareto curve diagram for different conflict mitigation strategies in this invention.
[0082] Figure 6 This is a graph showing the total reward value under the conflict-free detection mechanism of this invention.
[0083] Figure 7 This is the total reward curve compared with different gradient mitigation strategies in this invention. Detailed Implementation
[0084] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0085] like Figure 3 As shown, the UAV-assisted edge computing multi-objective optimization method based on gradient collision detection provided by this invention includes the following steps performed in sequence:
[0086] 1) Construct a drone-assisted mobile edge computing network system that includes drones equipped with edge servers and ground terminal equipment. This system is used to characterize the task arrival, link status changes and edge computing service process of ground terminal equipment, and initialize system parameters and constraints to provide a basic environment for subsequent communication modeling, computation modeling and optimization decision-making.
[0087] The method is as follows:
[0088] like Figure 2 As shown, a drone-assisted mobile edge computing network system is constructed, consisting of one drone equipped with an edge server and N ground terminal devices. The drone acts as a mobile edge computing node, providing task offloading and edge computing services to the ground terminal devices. The ground terminal devices continuously generate computing tasks and can establish wireless communication connections with the drone to achieve task data uploading and offloading. The positional relationship between the drone and the ground terminal devices is represented in a three-dimensional Cartesian coordinate system: the set of positions Q of the ground terminal devices is defined. users = {Q u (1), Q u (2),…,Q u (N)}, where Q u (i) = {x u (i),y u {i, 0} represents the fixed horizontal coordinates of ground terminal device i. Each ground terminal device can establish a wireless communication connection with the UAV for uploading and unloading mission data. The position of the UAV in time slot t is Q. uav (t) = {x uav (t),y uav (t),H}, where H is the flight altitude of the UAV;
[0089] Discretize the system runtime into a set of time slots. Where T is the preset total number of time slots, each time slot has a length of τ, and it is assumed that the completion time of each task, whether executed locally or offloaded, does not exceed one time slot, to meet the requirements of time slotted decision-making and online optimization; the task set of the ground terminal equipment in time slot t is defined as S(t) = {S1(t), S2(t),…, S…} N (t)}, where S i (t) represents the task generated by ground terminal device i in time slot t, used to characterize the characteristics of terminal task arrival and task size change over time; when the ground terminal device chooses to offload the task to the edge server on the UAV, due to the limited capacity of the wireless link, only a portion of the task can be uploaded in a single time slot. Therefore, tasks that are not scheduled for upload / not completed for upload wait in the transmission queue on the UAV side, thus forming a service process of "task arrival - offload request - task queue waiting - scheduling transmission - edge computing".
[0090] The system parameters include the number of terminals N, the flight altitude of the UAV H, the time slot length τ, the total number of time slots T, and the terminal task arrival / task data scale parameters; the constraints include wireless link capacity constraints, task completion time limit assumptions, and the remaining energy / energy budget of the ground terminal equipment, to ensure that subsequent offloading and scheduling decisions are executed within the feasible domain.
[0091] 2) Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model is established according to the positional relationship between the UAV and the ground terminal equipment to obtain the link transmission rate during the task upload process; based on the above link transmission rate, a local computing model and an edge computing model are established to obtain the local computing latency and the total offloading execution latency; based on the communication process and the computing process, a local computing energy consumption model and an edge computing energy consumption model are established to characterize the energy consumption generated during the task execution process.
[0092] The method is as follows:
[0093] Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model between ground terminal equipment and UAV upper edge server is constructed.
[0094] The wireless transmission link is affected by both small-scale fading and large-scale path loss. Its instantaneous channel gain is expressed as:
[0095] ;
[0096] Among them, g i,u L is a random variable that follows a Rayleigh distribution and is used to characterize small-scale fading. i,u The large-scale path loss described by the logarithmic distance model is expressed as:
[0097] ;
[0098] Where parameters A and B are the constant term and attenuation factor in the path loss, respectively, d i,u Let i be the distance between the drone i and the ground terminal equipment u.
[0099] Under the above channel conditions, based on the above channel gain, the link transmission rate is given by Shannon's formula:
[0100] ;
[0101] Where B is the bandwidth allocated to the link, and P... i Let σ be the transmit power of UAV i. 2 The above formula shows that the link transmission rate depends not only on the instantaneous state of the wireless channel, but also on factors such as distance, bandwidth, and transmit power.
[0102] Based on the system parameters obtained in step 1), a local computing model is constructed:
[0103] Let the amount of task data generated by ground terminal equipment u in time slot t be D. u (t), the task computation density is C (CPU cycle bits), and the CPU frequency of the ground terminal equipment is f. uThe local computation latency is:
[0104] ;
[0105] Based on the above link transmission rate, an edge computing model is constructed: when the ground terminal device u offloads the task to the edge server of the drone for execution in time slot t, its transmission latency is:
[0106] ;
[0107] Assume the CPU frequency of the edge server on the drone is f. i The execution latency of the task on the edge server is then:
[0108] ;
[0109] Therefore, the total execution delay for unloading is:
[0110] ;
[0111] In terms of energy consumption, a local computing energy consumption model is established, represented as follows:
[0112] ;
[0113] Where, k u The terminal energy consumption coefficient is determined by the processor architecture.
[0114] Establish an edge computing energy consumption model: decompose the energy consumption caused by terminal offloading into two parts: transmission energy consumption and edge execution energy consumption; when the ground terminal device u performs task offloading, its transmission energy consumption is determined by the transmission power of the ground terminal device u. The transmission delay mentioned above determines:
[0115] ;
[0116] When a task is executed on an edge server on a drone, its execution energy consumption is expressed as follows:
[0117] ;
[0118] Where, k i The energy consumption coefficient of the edge server on the drone;
[0119] Therefore, the energy consumption model for edge computing is as follows:
[0120] ;
[0121] in, This represents the total energy consumption for unloading execution.
[0122] 3) Based on the communication model and computing model established in step 2), obtain the task execution status and task completion time information, and construct an information age model accordingly. Then, decide whether to update the timestamp in the information age model based on whether the task is successfully completed, so as to characterize the freshness and timeliness of the system information.
[0123] The method is as follows:
[0124] Let τ be the timestamp corresponding to the task being received and processed in time slot t. u (t), then the information age model of the ground terminal equipment u in time slot t is defined as:
[0125] ;
[0126] in, This indicates the information age of the ground terminal equipment u in time slot t;
[0127] When a new task generated by ground terminal equipment u in time slot t is successfully received and processed, the timestamp in the information age model is updated to the latest timestamp, denoted as AoI. u (t+1)=1; When the new task is not completed, keep the original timestamp unchanged, represented as AoI. u (t+1)= AoI u (t)+1. Through the above update mechanism, the freshness of the status information of the ground terminal equipment can be dynamically reflected, and it can serve as an important evaluation indicator for subsequent joint optimization.
[0128] 4) Based on the system parameters obtained in step 1), the energy consumption model established in step 2), and the information age model constructed in step 3), design a multi-objective reward function and construct a multi-objective joint optimization model of information age and energy consumption to achieve a synergistic trade-off between the timeliness of system information and energy consumption.
[0129] The method is as follows:
[0130] Based on the model established in steps 1) to 3), the task unloading and scheduling process needs to reduce the system information age and system energy consumption simultaneously under dynamic channel and random task arrival conditions. Therefore, this problem can be abstracted into a multi-objective optimization problem, incorporating information age and energy consumption into a unified optimization framework, and optimizing information timeliness and energy consumption indicators simultaneously throughout the entire operation cycle.
[0131] To facilitate solving within the reinforcement learning framework, a system state s is introduced. t ∈S、Action a t ∈A and the strategy π(a|s); where the system state s tSystem information including the current information age of ground terminal equipment, task queue status, task data volume, wireless channel status, remaining energy of ground terminal equipment, and edge server computing resource usage; Action a t This represents the task execution method and resource allocation decisions, including the choice between local or offloaded execution, the offload ratio, and the allocation strategy for communication and computing resources; policy π(a|s) represents the probability distribution of choosing an action given a state; the goal is to learn the optimal policy π. * This allows the system to achieve the maximum cumulative return throughout its entire operating cycle.
[0132] To achieve a balanced optimization of information age and energy consumption, a multi-objective reward function consisting of information age reward and energy consumption reward is constructed:
[0133] Information Age Reward: To maximize the reward for younger information ages, the information age is negatively factored out, and the information age reward is defined as follows:
[0134] ;
[0135] Energy Consumption Reward: To maximize the reward for lower energy consumption, the energy consumption reward is defined as the total energy consumption of the aforementioned unloading execution. Negative values:
[0136] ;
[0137] By linearly weighting the two types of rewards using weight coefficients ω∈[0,1], a multi-objective reward function is constructed to obtain the weighted total reward:
[0138] ;
[0139] The weighting coefficient ω reflects the system's preference for optimizing information age and energy consumption at different times: a larger ω indicates a greater emphasis on reducing information age; a smaller ω indicates a greater emphasis on reducing energy consumption.
[0140] Based on the aforementioned multi-objective reward function, a multi-objective joint optimization model is constructed that maximizes the cumulative discount return over the entire operating cycle as the optimization objective for the UAV, considering both information age and energy consumption.
[0141] ;
[0142] Where γ∈[0,1] is the discount factor, and T is the maximum number of time slots in the system operation; by learning the optimal strategy to maximize the weighted total reward, a multi-objective joint optimization of information age and energy consumption is achieved.
[0143] 5) Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the reinforcement learning algorithm is used to learn the task offloading and resource allocation strategy. During the strategy update process, the gradient conflict between the information age objective and the energy consumption objective is detected, and the conflict gradient components are eliminated by gradient projection correction, thereby improving the stability and convergence performance of the multi-objective optimization process, and finally obtaining the optimal strategy of the joint operation of UAV trajectory, task offloading and transmission scheduling.
[0144] The method is as follows:
[0145] Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the task unloading decision process is modeled as a reinforcement learning process using a reinforcement learning algorithm: the UAV observes the system state s(t) in each time slot t, outputs action a(t), and obtains reward r(t) based on environmental feedback. The optimal policy π(a|s) is approximated by iteratively updating the policy network parameters θ in the reinforcement learning algorithm to improve the expected cumulative reward.
[0146] To avoid the problems of mutual cancellation of gradients among multiple objectives, unstable training, or even falling into suboptimal solutions caused by directly using "weighted total reward", corresponding objective loss functions L are constructed for the information age objective and the energy consumption objective, respectively. (aoi) With L (e) The objective loss function L (aoi) With L (e) The information obtained in step 4) is the age reward r. (aoi) and energy consumption reward r (e) The gradient vectors ∇ for the two objectives are obtained through the policy gradient method. During the policy network update process, the gradient vectors ∇ for the two objectives are calculated using the policy gradient method. θ L (aoi) With ∇ θ L (e) , used to characterize the influence of different optimization objectives on the update direction of the policy network parameter θ;
[0147] In the multi-objective update process, the gradient vector dot product is used to determine whether the optimization directions of different objectives are consistent: when the gradient vector dot product of two objectives is positive, it indicates that the optimization directions are consistent; when the gradient vector dot product of two objectives is negative:
[0148] ;
[0149] This indicates a gradient conflict has occurred, such as Figure 2 As shown, the PCGrad strategy is needed to correct conflicting gradients to avoid gradient cancellation and training instability caused by direct weighting. The method is to project and correct one of the gradient vectors (e.g., the gradient vector of the information age target).
[0150] ;
[0151] The correction term corresponds to the negative projection along the direction of another gradient vector (the energy consumption target gradient vector), thus eliminating the conflicting components between the two. After the correction is completed, the conflict-free gradient vectors are weighted and synthesized.
[0152] ;
[0153] Ultimately, the parameters θ←θ-ηg are used to update the policy network, where η is the learning rate; finally, the optimal policy combining the UAV trajectory, task offloading, and transmission scheduling is obtained.
[0154] The inventors conducted the following simulation experiments to verify the effectiveness of the method of the present invention. Figure 4 and Figure 5 It can be seen that by adding a conflict detection mechanism and adopting the PCGrad strategy, this method achieves a lower AoI at the same energy consumption level, or lower energy consumption at the same AoI level. This indicates that the introduced conflict detection mechanism and PCGrad mitigation mechanism effectively coordinate the two optimization objectives of latency and energy consumption, avoiding the performance degradation caused by gradient conflicts in traditional multi-objective optimization methods. Figure 6 It can be seen that PCGrad can not only identify conflicts between different objectives, but also effectively eliminate interference components in conflict gradients through projection correction, thereby enabling more coordinated joint optimization of the AoI objective and the energy consumption objective during parameter update. Figure 7 It can be seen that, compared with the baselines PPO and GD, the proposed method considering gradient conflict mitigation has a significant advantage in convergence. Among several typical conflict mitigation methods, PCGrad performs best in terms of convergence speed, training stability, and final reward level, making it more suitable for solving the joint optimization problem of AoI and energy consumption.
Claims
1. A multi-objective optimization method for UAV-assisted edge computing based on gradient collision detection, characterized in that: The method includes the following steps performed in sequence: 1) Construct a drone-assisted mobile edge computing network system that includes drones equipped with edge servers and ground terminal equipment. This system is used to characterize the task arrival, link status changes and edge computing service process of ground terminal equipment, and initialize system parameters and constraints to provide a basic environment for subsequent communication modeling, computation modeling and optimization decision-making. 2) Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model is established according to the positional relationship between the UAV and the ground terminal equipment to obtain the link transmission rate during the task upload process; Based on the above link transmission rate, a local computing model and an edge computing model are established to obtain the local computing latency and the total offloading execution latency. Local computing energy consumption model and edge computing energy consumption model are established based on communication process and computing process to characterize the energy consumption generated during task execution; 3) Based on the communication model and computing model established in step 2), obtain the task execution status and task completion time information, and construct an information age model accordingly. Then, decide whether to update the timestamp in the information age model based on whether the task is successfully completed, so as to characterize the freshness and timeliness of the system information. 4) Based on the system parameters obtained in step 1), the energy consumption model established in step 2), and the information age model constructed in step 3), design a multi-objective reward function and construct a multi-objective joint optimization model of information age and energy consumption to achieve a synergistic trade-off between the timeliness of system information and energy consumption. 5) Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the reinforcement learning algorithm is used to learn the task offloading and resource allocation strategy. During the strategy update process, the gradient conflict between the information age objective and the energy consumption objective is detected, and the conflict gradient components are eliminated by gradient projection correction, thereby improving the stability and convergence performance of the multi-objective optimization process, and finally obtaining the optimal strategy of the joint operation of UAV trajectory, task offloading and transmission scheduling.
2. The UAV-assisted edge computing multi-objective optimization method based on gradient conflict detection according to claim 1, characterized in that: In step 1), the method for constructing a drone-assisted mobile edge computing network system consisting of a drone equipped with an edge server and ground terminal equipment, and initializing system parameters and constraints, is as follows: Construct a drone-assisted mobile edge computing network system consisting of one drone equipped with an edge server and N ground terminal devices; represent the positional relationship between the drone and the ground terminal devices in a three-dimensional Cartesian coordinate system: define the set of positions Q of the ground terminal devices. users = {Q u (1), Q u (2),…,Q u (N)}, where Q u (i) = {x u (i),y u {i, 0} represents the fixed horizontal coordinates of ground terminal device i. Each ground terminal device can establish a wireless communication connection with the UAV for uploading and unloading mission data. The position of the UAV in time slot t is Q. uav (t) = {x uav (t),y uav (t),H}, where H is the flight altitude of the UAV; Discretize the system runtime into a set of time slots. Where T is the preset total number of time slots, each time slot has a length of τ, and it is assumed that the completion time of each task, whether executed locally or offloaded, does not exceed one time slot, to meet the requirements of time slotted decision-making and online optimization; the task set of the ground terminal equipment in time slot t is defined as S(t) = {S1(t), S2(t),…,S...} N (t)}, where S i (t) represents the task generated by ground terminal device i in time slot t, used to characterize the characteristics of terminal task arrival and task size change over time; when the ground terminal device chooses to offload the task to the edge server on the UAV, due to the limited capacity of the wireless link, only a portion of the task can be uploaded in a single time slot. Therefore, tasks that are not scheduled for upload / not completed for upload wait in the transmission queue on the UAV side, thus forming a service process of "task arrival - offload request - task queue waiting - scheduling transmission - edge computing". The system parameters include the number of terminals N, the flight altitude of the UAV H, the time slot length τ, the total number of time slots T, and the terminal task arrival / task data scale parameters; the constraints include wireless link capacity constraints, task completion time limit assumptions, and the remaining energy / energy budget of the ground terminal equipment, to ensure that subsequent offloading and scheduling decisions are executed within the feasible domain.
3. The UAV-assisted edge computing multi-objective optimization method based on gradient collision detection according to claim 1, characterized in that: In step 2), the UAV-assisted mobile edge computing network system built based on step 1) establishes a communication model according to the positional relationship between the UAV and the ground terminal equipment to obtain the link transmission rate during the task upload process; Based on the aforementioned link transmission rate, a local computing model and an edge computing model are established to obtain the local computing latency and the total offloading execution latency. The method for establishing the local computing energy consumption model and the edge computing energy consumption model based on the communication process and the computing process is as follows: Based on the UAV-assisted mobile edge computing network system constructed in step 1), a communication model between ground terminal equipment and UAV upper edge server is constructed. A wireless transmission link is affected by both small-scale fading and large-scale path loss. Its instantaneous channel gain is expressed as: ; Among them, g i,u L is a random variable that follows a Rayleigh distribution and is used to characterize small-scale fading. i,u The large-scale path loss described by the logarithmic distance model is expressed as: ; Where parameters A and B are the constant term and attenuation factor in the path loss, respectively, d i,u Let i be the distance between the drone i and the ground terminal equipment u. Under the above channel conditions, based on the above channel gain, the link transmission rate is given by Shannon's formula: ; Where B is the bandwidth allocated to the link, and P... i Let σ be the transmit power of UAV i. 2 Noise power; Based on the system parameters obtained in step 1), a local computing model is constructed: Let the amount of task data generated by ground terminal equipment u in time slot t be D. u (t), the task computation density is C, and the CPU frequency of the ground terminal equipment is f. u The local computation latency is: ; Based on the above link transmission rate, an edge computing model is constructed: when the ground terminal device u offloads the task to the edge server of the drone in time slot t, its transmission latency is: ; Assume the CPU frequency of the edge server on the drone is f. i The execution latency of the task on the edge server is then: ; Therefore, the total execution delay for unloading is: ; In terms of energy consumption, a local computing energy consumption model is established, represented as follows: ; Where, k u The terminal energy consumption coefficient is determined by the processor architecture. Establish an edge computing energy consumption model: decompose the energy consumption caused by terminal offloading into two parts: transmission energy consumption and edge execution energy consumption; when the ground terminal device u performs task offloading, its transmission energy consumption is determined by the transmission power of the ground terminal device u. The transmission delay mentioned above determines: ; When a task is executed on an edge server on a drone, its execution energy consumption is expressed as follows: ; Where, k i The energy consumption coefficient of the edge server on the drone; Therefore, the energy consumption model for edge computing is as follows: ; in, This represents the total energy consumption for unloading execution.
4. The UAV-assisted edge computing multi-objective optimization method based on gradient collision detection according to claim 1, characterized in that: In step 3), the method for obtaining task execution status and task completion time information based on the communication model and computing model established in step 2), constructing an information age model accordingly, and then determining whether to update the timestamp in the information age model based on whether the task was successfully completed is as follows: Let τ be the timestamp corresponding to the task being received and processed in time slot t. u (t), then the information age model of the ground terminal equipment u in time slot t is defined as: ; in, This indicates the information age of the ground terminal equipment u in time slot t; When a new task generated by ground terminal equipment u in time slot t is successfully received and processed, the timestamp in the information age model is updated to the latest timestamp, denoted as AoI. u (t+1)=1; When the new task is not completed, keep the original timestamp unchanged, represented as AoI. u (t+1)= AoI u (t)+1.
5. The UAV-assisted edge computing multi-objective optimization method based on gradient conflict detection according to claim 1, characterized in that: In step 4), the method for designing a multi-objective reward function and constructing a multi-objective joint optimization model of information age and energy consumption based on the system parameters obtained in step 1), the energy consumption model established in step 2), and the information age model constructed in step 3) is as follows: Introducing system state s t ∈S、Action a t ∈A and the strategy π(a|s); where the system state s t System information including the current information age of ground terminal equipment, task queue status, task data volume, wireless channel status, remaining energy of ground terminal equipment, and edge server computing resource usage; Action a t This represents the task execution method and resource allocation decisions, including the choice between local or offloaded execution, the offload ratio, and the allocation strategy for communication and computing resources; policy π(a|s) represents the probability distribution of choosing an action given a state; the goal is to learn the optimal policy π. * This allows the system to achieve the maximum cumulative return throughout its entire operating cycle; To achieve a balanced optimization of information age and energy consumption, a multi-objective reward function consisting of information age reward and energy consumption reward is constructed: Information Age Reward: To maximize the reward for younger information ages, the information age is negatively factored out, and the information age reward is defined as follows: ; Energy Consumption Reward: To maximize the reward for lower energy consumption, the energy consumption reward is defined as the total energy consumption of the aforementioned unloading execution. Negative values: ; By linearly weighting the two types of rewards using weight coefficients ω∈[0,1], a multi-objective reward function is constructed to obtain the weighted total reward: ; Based on the aforementioned multi-objective reward function, a multi-objective joint optimization model is constructed that maximizes the cumulative discount return over the entire operating cycle as the optimization objective for the UAV, considering both information age and energy consumption. ; Where γ∈[0,1] is the discount factor, and T is the maximum number of time slots in the system.
6. The UAV-assisted edge computing multi-objective optimization method based on gradient collision detection according to claim 1, characterized in that: In step 5), the multi-objective joint optimization model of information age and energy consumption constructed in step 4) uses reinforcement learning algorithm to learn task offloading and resource allocation strategies. The method for detecting gradient conflicts between the information age target and the energy consumption target during the policy update process, and eliminating conflicting gradient components through gradient projection correction, ultimately obtaining the optimal policy jointly considering UAV trajectory, task offloading, and transmission scheduling is as follows: Based on the multi-objective joint optimization model of information age and energy consumption constructed in step 4), the task unloading decision process is modeled as a reinforcement learning process using a reinforcement learning algorithm: the UAV observes the system state s(t) in each time slot t, outputs action a(t), and obtains reward r(t) based on environmental feedback. The optimal policy π(a|s) is approximated by iteratively updating the policy network parameters θ in the reinforcement learning algorithm to improve the expected cumulative reward. Construct corresponding target loss functions L for information age target and energy consumption target respectively. (aoi) With L (e) The objective loss function L (aoi) With L (e) The information obtained in step 4) is the age reward r. (aoi) and energy consumption reward r (e) The gradient vectors ∇ for the two objectives are obtained through the policy gradient method. During the policy network update process, the gradient vectors ∇ for the two objectives are calculated using the policy gradient method. θ L (aoi) With ∇ θ L (e) , used to characterize the influence of different optimization objectives on the update direction of the policy network parameter θ; In the multi-objective update process, the gradient vector dot product is used to determine whether the optimization directions of different objectives are consistent: when the gradient vector dot product of two objectives is positive, it indicates that the optimization directions are consistent; when the gradient vector dot product of two objectives is negative: ; This indicates a gradient conflict, requiring the PCGrad strategy to correct the conflicting gradients and avoid gradient cancellation and training instability caused by direct weighting. The method involves projecting and correcting one of the gradient vectors. ; The correction term corresponds to the negative projection along the direction of another gradient vector, thus eliminating the conflicting components between the two. After the correction is completed, the conflict-free gradient vectors are weighted and synthesized: ; Ultimately, the parameters θ←θ-ηg are used to update the policy network, where η is the learning rate; finally, the optimal policy combining the UAV trajectory, task offloading, and transmission scheduling is obtained.