Unmanned aerial vehicle unloading and charging path planning method for Internet of Vehicles mobile edge computing

By using drones to carry edge servers and deep reinforcement learning algorithms in the Internet of Vehicles scenarios, the unloading and charging paths of drones are optimized, and the Internet of Vehicles' needs for low-latency computing and real-time user interaction are solved, and efficient and dynamic computing resource management and battery life extension are achieved.

CN120046922APending Publication Date: 2025-05-27GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127395.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-04
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the demands of the Internet of Vehicles for low-latency computing and real-time user interaction, especially when fixed edge servers cannot be dynamically adjusted and the battery life of the drone is insufficient.

Method used

The drone is equipped with an edge server, and the unloading and charging paths of the drone are optimized through path planning methods and deep reinforcement learning algorithms, and the charging station location is dynamically adjusted to reduce the mobile cost of the drone.

Benefits of technology

It realizes efficient computing and offloading and charging management of drones in vehicle network scenarios, dynamically adjusts service scope and computing resource allocation, and reduces the mobile cost and battery consumption of drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046922A_ABST
    Figure CN120046922A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle unloading and charging path planning method for Internet of Vehicles mobile edge computing, an unmanned aerial vehicle is used for unloading an Internet of Vehicles user, and the specific path planning method can be summarized as comprising the following steps: step 1, modeling an unloading and charging path planning scene of the unmanned aerial vehicle; step 2, constructing an objective function about the movement and calculation of the unmanned aerial vehicle carrying the edge server and the deployment cost of the charging station; and step 3, determining the position of the charging station by using constraints, and solving the target function by using a multi-agent depth deterministic strategy gradient algorithm, thereby obtaining an optimal solution of unmanned aerial vehicle unloading and charging path planning. According to the method, the unmanned aerial vehicle is used for carrying the edge server to track a user of the Internet of Vehicles in real time, the problem that a fixed station cannot meet the calculation unloading requirement according to time-space change is solved, and the problem that the service life of a battery is short when the unmanned aerial vehicle is used for deploying the edge server is solved according to a charging station set by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the construction planning of vehicle - to - everything (V2X) mobile edge computing facilities, and the goal focuses on a method for deploying urban mobile edge servers based on unmanned aerial vehicle (UAV) offloading and charging path planning. Background Art

[0002] In recent years, with the development of 5G, the increasingly mature intelligent transportation, intelligent driving, and industrial control will surely change all aspects of human life. Edge computing provides powerful data - processing capabilities for Internet of Things (IoT) devices and supports the development of IoT applications such as smart home, smart city, and industrial automation. Mobile edge computing, as a research branch of edge computing, uses network service facilities closer to devices to provide computing services. It can provide faster response speed and better service quality. Since it moves computing and storage capabilities to the network edge, it can serve users closer, thereby reducing latency and increasing response speed, and supporting a wider range of V2X application scenarios.

[0003] In today's highly information - based era, the rapid development of V2X is accompanied by new applications with requirements for low - latency computing, real - time user interaction, and continuous and efficient services. In terms of using UAVs carrying edge servers for computing offloading, domestic and foreign research is still in its initial stage. Although there has been some theoretical research before, expanding it to actual deployment still faces many challenges, and more research work is needed to achieve a solution that is efficient, sustainable, low - cost, and widely applicable. In terms of edge - server - assisted V2X computing offloading, it is divided into fixed edge servers and mobile edge servers. For fixed edge servers, they have problems such as location - deployment limitations, energy consumption, and difficulty in implementing computing - capacity adjustment. Although they have the characteristics of being easy to deploy and implement, they are usually deployed at specific physical locations, covering a limited range of V2X users and being immobile; they are configured with continuous power supply and cooling systems, which will result in high energy consumption and operating costs. For mobile edge servers, they are divided into edge servers carried by buses and edge servers carried by UAVs. Edge servers carried by buses are affected by traffic flow and bus routes, making it impossible for the edge servers to cover all V2X users, and the risk of disconnection at any time leads to a poor experience for V2X users. While UAV - carried edge servers can solve the above problems, and combined with charging stations, they can solve the UAV endurance problem. UAVs can also make corresponding solutions for the characteristics of the highly spatio - temporal dynamic changes of V2X users. Summary of the Invention

[0004] To fill the above technical gap, the present invention discloses a method for unmanned aerial vehicle (UAV) offloading and charging path planning for vehicle-to-everything (V2X) mobile edge computing.

[0005] The technical solution adopted in this invention patent is as follows:

[0006] A method for UAV offloading and charging path planning for V2X mobile edge computing. The steps for path planning of using UAVs to offload the calculations of V2X users specifically include:

[0007] Model the scenario of UAV offloading and charging path planning;

[0008] Construct an objective function regarding UAV movement, computing, and charging station deployment costs;

[0009] Use constraints to determine the positions of charging stations, and then solve the objective function using the method of deep reinforcement learning to obtain the optimal solution for UAV offloading and charging path planning.

[0010] A further optimized solution is as follows.

[0011] The steps for modeling the scenario of UAV offloading and charging path planning are specifically as follows:

[0012] Let \(W = \{w_{i}|i = 1,\cdots,N\}\) represent the set of \(N\) tasks of the UAV; i |i=1,...,N} represents the set of N tasks of the UAV;

[0013] Let \((G,B)\) be the computing task issued by the user, where \(G\) is the amount of input data of the task and \(B\) is the number of cycles required for the computing task. The computing tasks are randomly distributed in the modeled scenario;

[0014] Suppose there are \(K\) charging stations, and these \(K\) charging stations need to be deployed into the scenario through constraints. Use \(R = \{r_{i}|i = 1,\cdots,K\}\) to represent the set of UAV charging stations; i |i=1,...,K} represents the set of UAV charging stations;

[0015] Suppose there are \(M\) UAVs. Use \(U = \{u_{i}|i = 1,\cdots,M\}\) to represent the set of UAVs. The positions of the UAVs are initialized at the positions of the charging stations. i |i=1,...,M} represents the set of UAVs, and the positions of the UAVs are initialized at the positions of the charging stations.

[0016] A further optimized solution is as follows.

[0017] The constructed objective function regarding UAV movement, computing, and charging station deployment costs is:

[0018] min:C=C all uav +C all charge

[0019]

[0020] Among them, C all uav represents the movement and computing costs of all drones, and C all charge represents the deployment cost of the drone charging station;

[0021]

[0022] represents the cost of the i-th drone moving a unit distance, and d i represents the distance the i-th drone moves, represents the movement cost of all drones;

[0023] represents the computing cost of the i-th drone, represents the computing cost of all drones;

[0024] The computing cost of the i-th drone The formula can be expressed as:

[0025]

[0026] Among them, k represents the number of tasks within the coverage of drone i, represents task w j the number of input bits required, and each input bit requires CPU cycles for processing, that is represents the number of CPU cycles required for each input of task w j ;

[0027]

[0028] Among them, represents the maximum computing power of the drone. This formula means that the computing power corresponding to the tasks processed by the drone must be less than the maximum computing power of the drone.

[0029] A further optimized solution is,

[0030] Use constraints to determine the charging station location, solve the objective function based on the method of deep reinforcement learning, and obtain the optimal solution for the drone offloading and charging path planning. The specific steps are as follows:

[0031] a1. Initialization: Determine the drone charging station location through constraints, randomly generate multiple drones at the charging station location, and set the path planning cost of the drone carrying the edge server to infinity;

[0032] a2. Input the task set W of the vehicle network users;

[0033] a3. Determine the specific locations of the charging stations by setting the distance constraints between each UAV charging station and the boundary constraints with the simulation environment;

[0034] a4. Set the range of the number of UAV charging stations and the maximum range of the map as limiting conditions;

[0035] a5. Combine the information freshness and the total energy consumption of multiple UAVs to set the reward function r t :

[0036]

[0037] where r A represents the reward for the UAV to process the task freshness, which consists of the reward of the task freshness itself and the penalty λ after a certain duration; r A represents the reward for the UAV energy consumption, which consists of the energy consumption of the UAV itself E and the constraint penalty λ for the UAV flight. and the constraint penalty λ for the UAV flight. E consists of.

[0038] a6. For the information freshness mentioned in step a5, every time a time interval passes, the information freshness of the task decreases by one point. After the task is processed and a new task is issued, after another time interval, the information freshness decreases by one point again, and so on;

[0039] a7. Learn the optimal paths for UAV offloading and charging through the exploration mechanism of deep reinforcement learning, and increase the overall reward R of the system to approach the optimal path with the minimum combination of the information freshness and energy consumption of the UAV to process tasks.

[0040] The further optimized solution is that

[0041] the input of the deep learning method is the state of the current environment, including:

[0042] The positions of the current states of M UAVs are respectively

[0043] The energy states of M UAVs are respectively E 1 , …, E M ,

[0044] The positions of K UAV charging stations are respectively

[0045] The positions of N tasks are respectively

[0046] The maximum number of iterations T max , the coverage radius R of the UAV;

[0047] Input the current environmental status information into the Actor-Critic network. First, output the actions for the UAV through the policy function. After the UAV takes actions, evaluate the impact on the environment by the value function and output the reward r.

[0048] The further optimized solution is as follows.

[0049] Set the system state s(t) = {s UAV (t), s CS (t)}, where s UAV (t) represents the UAV state and s CS (t) represents the charging station state.

[0050] The state of the UAV

[0051] The state of the charging station

[0052] Define the action space as a(t) = {a f (t), c(t)}.

[0053] Introduce the AoI reward R AoI ,

[0054]

[0055] where β and γ are parameters controlling the reward and punishment intensities, M is the threshold of AoI, and AoI(t) represents the AoI value based on the system state.

[0056] Introduce the energy consumption reward R energy ,

[0057] R energy = -α × E(t)

[0058] where α is the weight coefficient of energy consumption and E(t) represents the flight energy consumption and task energy consumption of the UAV.

[0059] Introduce the overall reward R: By combining the rewards of AoI and energy consumption, the overall reward function can be obtained:

[0060] R = R AoI + R energy .

[0061] The further optimized solution is as follows.

[0062] The state of the nth UAV is located as where E n (t) represents the current battery level, P n (t) represents the current position, gn (t) represents the current working state,

[0063] The remaining power E(t) of each drone belongs to [0, 1]. When E(t) = 1, the drone's battery is fully charged;

[0064] The working state g(t) belongs to {0, 1, 2}. When g(t) = 0, the drone is in the return state and does not act as a server, directly returning to the charging station; when g(t) = 1, the drone is in the directional service stage, that is, it still acts as a server during the return journey; when g(t) = 2, the drone is working normally.

[0065] A further optimized solution is,

[0066] where s CS (t) is equal to 0 or 1, indicating whether the charging station can provide charging service.

[0067] A further optimized solution is,

[0068] In each time slot, the drone has five action choices, af(t) ∈ {0, 1, 2, 3, 4}, which are hovering, moving forward, moving backward, moving left, and moving right respectively.

[0069] When hovering a f (t) = 0, it means the drone hovers (lands), and lands when it coincides with the selected target position; c(t) = {1, 2,..., n} represents the selected target.

[0070] The present invention uses drones carrying edge servers to track users of the vehicle network in real time, solves the problem that fixed sites cannot meet the computing offloading requirements according to spatio-temporal changes, and solves the problem of insufficient battery life of edge servers deployed by drones according to the charging stations set by itself. The combination of the charging station and the drone can maximize the role of the drone, continuously offload the computing requirements of vehicle network users, and the movement of the drone allows them to dynamically adjust the service range and the allocation of computing resources according to user needs and environmental changes. Specify the path planning method, establish a mathematical model to describe the movement cost of the drone carrying the edge server, constrain the location deployment of the charging station, propose a reinforcement learning algorithm to solve the optimal solution, and the proposed path planning method for drones carrying edge servers can greatly reduce the movement cost of the drones. Brief Description of the Drawings

[0071] Figure 1 It is the overall flowchart of the drone offloading and charging path planning method for mobile edge computing of the vehicle network in the present invention;

[0072] Figure 2 It is a schematic diagram of the invention for the change of freshness of drone processing tasks;

[0073] Figure 3 To invent two modes (directional service mode and return mode) during the return process of the drone. Detailed implementation manners

[0074] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0075] The purpose of the present invention is to overcome the gap in the current technology and provide a method for planning the unloading and charging paths of drones for the society.

[0076] To achieve the above objective, the technical solution provided by the present invention is as follows:

[0077] The present invention discloses a method for planning the unloading and charging paths of drones for mobile edge computing in the vehicle network, using drones to unload vehicle network users. The specific path planning method can be generally summarized as including the following steps, as Figure 1 shown:

[0078] Step 1: Model the scenario of the unloading and charging path planning of the drone;

[0079] Step 2: Construct an objective function regarding the movement and calculation of the drone carrying the edge server and the deployment cost of the charging station;

[0080] Step 3: Use constraints to determine the location of the charging station, and then use the multi-agent deep deterministic policy gradient algorithm to solve the objective function, so as to obtain the optimal solution of the unloading and charging path planning of the drone.

[0081] In Step 1, specifically modeling the scenario of the unloading and charging path planning of the drone includes:

[0082] Let W = {w i | i = 1,..., N} represent the set of N tasks of the drone;

[0083] Let (G, B) be the computing task issued by the user, where G is the amount of input data of the task and B is the number of cycles required for the computing task. The computing tasks are randomly distributed in the modeled scenario;

[0084] Suppose there are K charging stations, and these K charging stations need to be deployed into the scenario through constraints. Use R = {r i | i = 1,..., K} to represent the set of drone charging stations;

[0085] Suppose there are M drones. Use U = {u i | i = 1,..., M} to represent the set of drones. The positions of the drones are initialized at the positions of the charging stations.

[0086] In step 2, the objective function for the movement and computing of the drones carrying edge servers and the deployment cost of charging stations (i.e., the objective function for the path planning cost of the drones carrying edge servers) is as follows:

[0087] min: C = C all uav + C all charge

[0088]

[0089] where C all uav represents the movement and computing costs of all drones, and C all charge represents the deployment cost of the drone charging stations;

[0090]

[0091] represents the cost per unit distance of the i-th drone's movement, and d i represents the distance of the i-th drone's movement, representing the movement costs of all drones;

[0092] represents the computing cost of the i-th drone, representing the computing costs of all drones;

[0093] The computing cost of the i-th drone can be expressed by the formula:

[0094]

[0095] where k represents the number of tasks within the coverage of drone i, represents the number of input bits required for task w j Each input bit requires CPU cycles for processing, that is represents the number of CPU cycles required for each input of task w j (changing the uppercase K to lowercase k to distinguish it from the previous one).

[0096] where represents the maximum computing power of the drone. This formula means that the computing power corresponding to the tasks processed by the drone must be less than the maximum computing power of the drone.

[0097] In step 3, constraints are used to determine the location of the charging station, and then the objective function is solved based on the method of deep reinforcement learning to obtain the optimal solution for the UAV offloading and charging path planning. In the solution of the present invention, the multi-agent deep deterministic policy gradient algorithm is selected. The specific implementation steps of step 3 are as follows:

[0098] a1. Initialization: Determine the location of the UAV charging station through constraints, randomly generate multiple UAVs at the charging station location, and the initial energy of the UAVs is the upper limit of the UAV battery storage;

[0099] a2. Input the task set W of the vehicle network users, and randomly initialize the positions and speeds of the vehicle network users on the simulated road through vehicle constraints;

[0100] a3. Determine the specific location of the charging station by setting the distance constraints of each UAV charging station and the boundary constraints with the simulated environment;

[0101] a4. Set the range of the number of UAV charging stations and the maximum range of the map as limiting conditions;

[0102] a5. Plan the offloading and charging paths of the UAVs through the method of deep reinforcement learning, combine the information freshness and the total energy consumption of multiple UAVs, and set the reward function r t :

[0103]

[0104] Among them, r A represents the reward for the UAV to process the task freshness, which consists of the reward of the task freshness itself and the penalty λ A after a certain duration; r E represents the reward for the UAV energy consumption, which consists of the energy consumption of the UAV itself and the constraint penalty λ E for the UAV flight.

[0105] a6. For the information freshness mentioned in step a5, every time a time interval passes, the information freshness of the task decreases by one point. After the task is processed and a new task is released, after another time interval, the information freshness decreases by one point again, and so on, as Figure 2 shown.

[0106] a7. Learn the optimal path for UAV offloading and charging through the exploration mechanism of deep reinforcement learning, and increase the reward of the system to a certain value to approximate the optimal path with the minimum comprehensive information freshness and energy consumption of the UAV to process tasks.

[0107] Use the deep reinforcement learning algorithm to determine the optimal paths for multiple drones to unload and charge. Each drone is an independent agent, and its position is sent to each drone by global state awareness (SDN), including:

[0108] The input is the state information of the current environment, specifically including the following: The positions of the current states of M drones are respectively The energy states of M drones are respectively E 1 , …, E M , and the positions of K drone charging stations are respectively The positions of N tasks are respectively The maximum number of iterations T max , and the coverage radius R of the drone;

[0109] One cycle step of deep reinforcement learning is as follows: Input the current environmental state information into the Actor-Critic network. First, output actions for the drones through the policy function (Actor). After the drones take actions, the impact on the environment is evaluated by the value function (Critic), and the reward r is output;

[0110] The specific steps of deep reinforcement learning for the unloading and charging path planning of drones are as follows:

[0111] 1) Problem definition

[0112] During the process of drones performing loading and unloading tasks, limited by the battery capacity of the drones and the energy consumption of the tasks, it is difficult to complete all tasks with a single charge. Therefore, it is necessary to deploy charging stations within the task execution area to expand the cruising radius of the drones and reduce the energy consumption of round trips to the task points. However, the construction cost of charging stations is relatively high. Therefore, it is necessary to minimize the number of charging stations on the premise of not affecting the completion of drone tasks, and select the charging stations for the drones to return and optimize the path planning.

[0113] The core objectives of this problem include:

[0114] 1. Minimize the number of charging stations: Ensure that the drones can successfully complete all tasks while reducing the deployment cost of charging stations.

[0115] 2. Optimize path planning: Design a reasonable cruising path for the drones to ensure the lowest energy consumption during task execution and return.

[0116] 3. Charging station selection strategy: Select appropriate charging stations between task points to prevent the drones from being unable to complete tasks due to exhausted power.

[0117] 2) State representation

[0118] A. State space

[0119] The system state includes the drone state s UAV (t) and the charging station state s CS (t), collectively referred to as the system state s(t) = {s UAV (t), s CS (t)};

[0120] The state of the drone

[0121] The current battery level E n (t) of the drone, the current position P n (t), and the working state g n (t) are defined as

[0122] The remaining battery level E(t) of each drone belongs to [0, 1]. When E(t) = 1, the drone's battery is fully charged.

[0123] The working state g(t) ∈ {0, 1, 2}. When g(t) = 0, the drone is in the return flight state and does not act as a server, directly returning to the charging station; when g(t) = 1, the drone is in the directional service stage, that is, it still acts as a server during the return flight; when g(t) = 2, the drone is working normally.

[0124] The state of the charging station Among them, s CS (t) is equal to 0 or 1, indicating whether the charging station can provide charging services.

[0125] B. Action space

[0126] When each drone completes the task of the current node, it can select the next node or fly to the charging station according to its state. Therefore, the action space is defined as a(t) = {a f (t), c(t)}. In each time slot, the drone has five action choices, a f (t) ∈ {0, 1, 2, 3, 4}, which are hovering, forward, backward, left, and right respectively. When hovering a f (t) = 0, it means the drone hovers (lands), and lands when it coincides with the selected target position. c(t) = {1, 2,..., n} represents the selected target.

[0127] 3) Policy definition

[0128] 1. Policy definition of the drone

[0129] As Figure 3 shown, for the drone, there are two modes that can be selected during the return flight process: the directional service mode and the return flight mode. The probabilities of the drone selecting the two modes are related to the remaining battery level of the drone.

[0130] (1) In the directional service mode, the drone will detour to nearby nodes in need of support on its way back to maximize the utilization rate of the battery between two charges. Suppose the drone departs from node i and returns to the jth charging station, and the state of this charging station is The set of nodes it can choose to visit on the way is I, and the path can be planned through an optimization problem:

[0131]

[0132] where d ij represents the distance between node i and node j, and the constraint is

[0133]

[0134] E ij is the energy consumption of the drone between node i and node j, and E(t) is the remaining battery power at the current moment.

[0135] (2) In the return mode, when the battery power is insufficient or there are no nodes in need of support nearby, the drone directly chooses to return to the charging station.

[0136] 2. Strategies and Modeling of Charging Stations

[0137] For the location selection of charging stations, the goal is to minimize the return distance of the drone. However, considering the actual situation, charging stations cannot be built at any location. Therefore, candidate locations can be selected first, and then according to the optimization algorithm, locations with high reward values can be selected as needed to build charging stations. Suppose there are m candidate locations P = {P 1 ,..., P m} in an area, and k locations need to be selected from them to build charging stations.

[0138] This problem should first predict the drone trajectory based on the node offloading data, and then model the trajectory as a location selection optimization problem, with the goal of minimizing the total distance of all drones returning from nodes to the charging stations:

[0139]

[0140] where, represents the distance from node i to the location of charging station c j , and n is the number of all nodes.

[0141] The constraint is the number k of charging stations to be built:

[0142]

[0143] where, P i = 1 indicates that a charging station is built at the location, P i= 0 indicates no construction.

[0144] 4) Model Selection

[0145] (1) Energy Consumption Model

[0146] The energy consumption of the UAV includes the flight energy consumption of the UAV and the task energy consumption of the UAV;

[0147] i) UAV flight energy consumption: Using a rotary-wing UAV to carry a mobile edge computing (MEC) server, the UAV flight energy consumption can be expressed as:

[0148]

[0149] P 0 is the hovering power, indicating the power consumption of the UAV in the hovering state. ||v n || is the modulus of the flight speed of the nth UAV. U tip is the propeller tip speed. Induced power coefficient P i , indicating the power consumption related to aerodynamic characteristics. v 0 The descent speed of the UAV when hovering. d 0 is the air resistance coefficient. ρ is the air density. S 0 is the frontal area of the UAV. A is the total area of the propellers. τ n is the flight time of the nth UAV.

[0150] ii) UAV task energy consumption: Assume that the CPUs of the ground equipment and the mobile edge computing (MEC) server on the UAV both adopt dynamic voltage and frequency scaling (DVFS) technology to control the computing frequency in an adaptive manner. This technology will dynamically adjust the computing frequency according to the computing requirements of the actual task, so as to optimize the energy consumption efficiency on the premise of meeting the computing performance. In this case, the UAV task energy consumption is:

[0151]

[0152] τ n is the execution time required for the nth UAV to complete the task, κ u is the computing energy consumption coefficient, related to the hardware characteristics of the UAV. is the CPU computing frequency of the nth UAV in task k.

[0153] (2) AoI Model

[0154] In a multi-agent system, the state of each agent will affect the performance of other agents. Therefore, an AoI model based on the system state should be set:

[0155]

[0156] where: N is the number of agents, and AoI i (t) is the freshness of information of the i-th agent at time t, E(t) is the total energy consumption of the system at time t, and Emax is the upper limit of the maximum energy consumption of the system

[0157] An AoI threshold can be set. When the AoI of a task exceeds this threshold, the value of the task information drops sharply. Therefore, a penalty term is added to the reward function to encourage the system to process older tasks as soon as possible.

[0158] 5) Definition of loss function

[0159] (1) AoI reward: When processing tasks, the fresher the task is processed (lower AoI), the more rewards are given. If the AoI of a certain task is less than the threshold M, the reward is positive; otherwise, a penalty is introduced.

[0160]

[0161] where β and γ are parameters that control the intensity of rewards and penalties. M is the threshold of AoI

[0162] (2) Energy consumption reward: The drone consumes energy during flight and task calculation. To encourage the drone to complete tasks while minimizing energy consumption, rewards can be added according to the energy consumption situation. Assume that the total energy consumption of the drone is E(t), then the energy consumption reward is:

[0163] R energy =-α×E(t)

[0164] where α is the weight coefficient of energy consumption.

[0165] (3) Overall reward function: By combining the rewards of AoI and energy consumption, the overall reward function can be obtained:

[0166] R=R AoI +R energy

[0167] This reward function can balance the timeliness of tasks and the energy consumption of the drone, and encourage the drone to minimize energy consumption while keeping information fresh.

[0168] 6) Training process

[0169] The multi-agent deep reinforcement learning (DDPG) is used to train the system to minimize the overall AoI and energy consumption of the system.

[0170] Step 1: Initialization:

[0171] Set the system environment, including the initial position of the drone, the position of the charging station, the task generation area, etc. And randomly generate tasks, and initialize AoI = 0.

[0172] Step 2: Define the state space:

[0173] (1) UAV state space: including the position, remaining power, and the status of whether a mission is in progress for each UAV

[0174] (2) Charging station state space: Its s CS (t) is equal to 0 or 1, indicating whether the charging station can provide charging services.

[0175] Step 3: Action space:

[0176] (1) Action definition: The action a f (t) of each UAV includes five modes: hovering, moving forward, moving backward, turning left, and turning right;

[0177] (2) Select the flight target c(t) according to the current state.

[0178] Step 4: Reward calculation

[0179] (1) After executing the action, calculate the AoI of each task and update the reward R accordingly;

[0180] (2) Calculate the energy consumption penalty for this action and combine it with the AoI reward to obtain the overall reward.

[0181] Step 5: Update the policy:

[0182] Use the reward value R to update the parameters of the policy network. Through the policy, the UAV can maintain the freshness of the mission while minimizing energy consumption.

[0183] Step 6: Training stop condition:

[0184] When the total system reward reaches the set stable value or reaches the maximum number of training steps, stop the training.

[0185] 7) Evaluate the performance

[0186] UAV path planning optimization

[0187] Indicator: Reward value obtained by the UAV system

[0188] Goal: When completing the mission, minimize the power consumption as much as possible to improve the overall mission efficiency.

[0189] Definition: max R = R AoI +R energy

[0190] The present invention uses a drone carrying an edge server to track users of the vehicle network in real time, solves the problem that fixed sites cannot meet the computing offloading requirements according to spatio-temporal changes, and solves the problem of insufficient battery life of the edge server deployed by the drone according to the set charging stations. The combination of the charging stations and the drones can maximize the role of the drones, continuously offload the computing requirements of the vehicle network users, and the movement of the drones allows them to dynamically adjust the service range and the allocation of computing resources according to user needs and environmental changes. The path planning method is specified, a mathematical model is established to describe the movement cost of the drone carrying the edge server, the location deployment of the charging stations is restricted, a reinforcement learning algorithm is proposed to solve the optimal solution, and the proposed path planning method for the drone carrying the edge server can greatly reduce the movement cost of the drone.

Claims

1. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles, characterized by: The steps of path planning using drones to offload vehicle network user computing include: Modeling scenarios for unloading and charging paths for drones; Construct an objective function for the cost of UAV movement and computing and charging station deployment; Constraints are used to determine the location of the charging station, and then deep reinforcement learning methods are used to solve the objective function to obtain the optimal solution for drone unloading and charging path planning.

2. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 1, characterized in that: The steps to model the scenario of unloading and charging path planning of drones are as follows: Let W = {w i |i=1,...,N} represents the set of N missions of the drone; Let (G, B) be the computing task issued by the user, where G is the amount of input data of the task, B is the number of cycles required for the computing task, and the computing tasks are randomly distributed in the modeling scenario; Suppose there are K charging stations, which need to be deployed into the scene through constraints, using R = {r i |i=1,...,K} represents the set of drone charging stations; Suppose there are M drones, using U = {u i |i=1,...,M} represents the set of drones, and the positions of the drones are initialized at the locations of the charging stations.

3. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 2, characterized in that: The objective function constructed for the UAV movement and computing and charging station deployment costs is: min:C=C alluav +C allcharge Among them, C alluav represents the movement and computation cost of all drones, C allcharge represents the deployment cost of drone charging stations; represents the cost of the i-th drone moving a unit distance, d i represents the distance moved by the i-th drone, represents the movement cost of all drones; represents the computational cost of the i-th UAV, represents the computational cost of all drones; The computational cost of the i-th drone The formula can be expressed as: Where k represents the number of missions within the coverage area of ​​UAV i, Represents task w j The number of input bits required, each input bit requires CPU cycles to process, that is Represents task w j Each input of is the number of CPU cycles required; in, Represents the maximum computing power of the drone. This formula means that the computing power corresponding to the task processed by the drone must be less than the maximum computing power of the drone.

4. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 3, characterized in that: The location of the charging station is determined using constraints, and the objective function is solved based on the deep reinforcement learning method to obtain the optimal solution for drone unloading and charging path planning. The specific steps include the following: a1. Initialization: determine the location of the drone charging station through constraints, randomly generate multiple drones at the charging station location, and set the path planning cost of the drone carrying the edge server to infinity; a2. Input the task set W of the Internet of Vehicles user; a3. Determine the specific location of the charging station by setting the distance constraints of each UAV charging station and the boundary constraints with the simulation environment; a4. Set the range of the number of drone charging stations and the maximum range of the map as restrictions; a5. Combine the information freshness and the total energy consumption of multiple drones to set the reward function r t : Among them, r A Represents the reward for the freshness of the drone's task processing, which is composed of the reward of the task freshness itself And the penalty λ after exceeding a certain time A Composed of; r E The reward for the drone's energy consumption, which is determined by the drone's own energy consumption and the constraint penalty λ for drone flight E Consists of. a6. For the information freshness mentioned in step a5, the information freshness of the task decreases by one point after each time interval. After the task is processed, when a new task is released, another time interval passes and the information freshness decreases by one point again, and so on. a7. Learn the optimal path for unloading and charging the drone through the exploration mechanism of deep reinforcement learning, and improve the overall reward R of the system to approach the optimal path with the minimum information freshness and energy consumption for the drone to process the task.

5. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 4, characterized in that: The input of the deep learning method is the state of the current environment, including: The current positions of the M drones are The energy states of the M drones are E1,…,E M , The locations of the K drone charging stations are The locations of N tasks are Maximum number of iterations T max , the coverage radius R of the drone; The current environmental state information is input into the Actor-Critic network. The action of the drone is first output through the policy function. After the drone takes action, the impact on the environment is evaluated by the value function, and the reward r is output.

6. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 5, characterized in that: Set the system state s(t) = {s UAV (t),s CS (t)}, where s UAV (t) represents the state of the drone, s CS (t) represents the charging station status, Drone status Charging station status The action space is defined as a(t) = {a f (t),c(t)}, Introducing AoI Rewards R AoI , Among them, β and γ are parameters that control the intensity of rewards and penalties, M is the threshold of AoI, and AoI(t) represents the AoI value based on the system state; Introducing energy consumption reward R energy , R energy =-α×E(t) Among them, α is the weight coefficient of energy consumption, and E(t) represents the flight energy consumption and mission energy consumption of the UAV. Introducing the overall reward R: Combining the rewards of AoI and energy consumption, we can get the overall reward function: R=R AoI +R energy 。 7. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 6, characterized in that: The status of the nth drone is located as Among them, E n (t) represents the current power, P n (t) indicates the current position, g n (t) indicates the current working status, The remaining power of each drone is E(t)∈[0,1]. When E(t)=1, the drone battery is fully charged; Working state g(t)∈{0,1,2}, when g(t)=0, the UAV is in the return state, does not act as a server, and directly returns to the charging station; when g(t)=1, the UAV is in the directional service stage, that is, it still acts as a server on the way back; when g(t)=2, the UAV works normally.

8. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 7, characterized in that: where s CS (t) is equal to 0 or 1, indicating whether the charging station can provide charging service.

9. A method for unloading and charging paths of drones for mobile edge computing in Internet of Vehicles according to claim 8, characterized in that: In each time slot, the drone has five action options, a f (t)∈{0,1,2,3,4}, which are hovering, forward, backward, left and right. When hovering a f When (t) = 0, it means that the drone is hovering (landing), and when it coincides with the selected target position, it is landing; c(t) = {1, 2, ..., n} represents the selected target.