Unmanned aerial vehicle task unloading method and device based on deep reinforcement learning
Through the deep reinforcement learning framework, optimized the task offloading, resource allocation and trajectory planning of drones in the Internet of Vehicles environment, solving the problems of task offloading efficiency and energy consumption in the Internet of Vehicles, and achieving efficient and low-latency task processing.
Patent Information
- Application Number
- CN202510117790.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the Internet of Vehicles environment, heterogeneous and time-varying tasks have strict requirements on unmanned aerial vehicle offload services, and the unloading decisions between UDs are complex, which affects the efficiency and energy consumption of computing tasks.
The UAV collaborative multi-device task offloading framework based on deep reinforcement learning is adopted to optimize the total system time delay and total energy consumption by jointly optimizing task offloading, resource allocation and drone trajectory planning.
It realizes efficient offloading of tasks, optimizes the utilization of computing resources, reduces latency and energy consumption, improves system throughput, and enhances robustness.
Smart Images

Figure CN119997105A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle networking technology, and in particular to a method and device for unloading unmanned aerial vehicle tasks based on deep reinforcement learning. Background Art
[0002] With the rapid development of technologies such as autonomous driving and intelligent transportation, the massive data generated in modern VANETs (Vehicular Ad-Hoc Networks) has rapidly increased the demand for computing power and communication resources. This puts higher demands on the existing network infrastructure, especially in multi-device and multi-task scenarios. How to efficiently handle these tasks has become an urgent problem to be solved. In order to meet these challenges, the combination of MEC (Multi-Access Edge Computing) and UAV (Unmanned Aerial Vehicle) technology provides a new solution for vehicle networking task offloading.
[0003] MEC can effectively reduce the transmission delay of computing tasks and reduce the bandwidth burden of the core network by deploying computing resources at the edge of the network. Drones, with their flexible maneuverability and wide coverage capabilities, have further expanded the application scenarios of MEC, especially in the face of dynamically changing Internet of Vehicles environments. The integration of drones and MEC provides a high-probability LoS (Line-of-Sight) communication link, which can significantly improve the communication coverage, capacity and connection reliability of the network. In addition, the flexibility of drones allows them to be quickly deployed in areas lacking ground infrastructure, providing on-demand computing resource support for multi-device tasks in the Internet of Vehicles.
[0004] Although drone-assisted MEC technology has many advantages, it still faces several challenges in practical applications: (1) In the Internet of Vehicles environment, the tasks performed by different UDs (User Devices) are usually heterogeneous and time-varying. The requirements for offloading services for these tasks are often very strict. Due to the limited battery capacity of drones, their flight time is greatly limited. This requires not only optimizing the use time of drones, but also balancing energy consumption and task completion efficiency while ensuring mission quality. Therefore, how to effectively allocate resources under resource constraints to meet different task requirements is a key issue. (2) The offloading decisions between UDs are complexly coupled. The offloading decision of each device not only affects its own computing tasks, but is also affected by the offloading decisions of other devices. This increases the computational complexity of the task offloading process. (3) Although the mobility of drones improves the flexibility of the system, it also brings complexity to its trajectory planning. In order to optimize task offloading, how to reasonably plan the flight path of drones becomes a challenge.
[0005] To address these issues, a UAV cooperative IoV multi-device task offloading framework based on DRL (Deep Reinforcement Learning) is proposed. By jointly optimizing task offloading, resource allocation, and UAV trajectory planning, the framework aims to minimize the total system time delay and minimize the total energy consumption of the system.
[0006] In recent years, with the widespread application of drones and multi-access edge computing technologies in the Internet of Vehicles, researchers have proposed a variety of optimization methods for task offloading and resource allocation, aiming to improve system performance and meet the vehicles' growing demand for computing resources.
[0007] Prior art [L. He et al., "An Online Joint Optimization Approach for QoE Maximization in UAV-Enabled Mobile Edge Computing," IEEE INFOCOM 2024-IEEE Conference on Computer Communications, Vancouver, BC, Canada, 2024, pp. 101-110, doi: 10.1109 / INFOCOM52122.2024.10621306.] He et al. considered a UAV-assisted MEC scenario, in which the UAV acts as an air edge server to provide computing services to multiple ground UDs. A joint task offloading, resource allocation, and UAV trajectory planning optimization problem is proposed to maximize the QoE (Quality of Experience) of the UD under the UAV energy consumption constraint.
[0008] Prior art [X. Dai, Z. Xiao, H. Jiang and JCS Lui, "UAV-Assisted Task Offloading in Vehicular Edge Computing Networks," in IEEE Transactions on Mobile Computing, vol. 23, no. 4, pp. 2520-2534, April 2024, doi: 10.1109 / TMC.2023.3259394.] Dai et al. solved the VEC overload problem by introducing UAVs. Specifically, a new online UAV-assisted vehicle task offloading problem is proposed, with the goal of minimizing vehicle task delays under long-term UAV energy constraints. To solve this problem, the long-term energy constraints are first decoupled based on the Lyapunov optimization technique, so that the problem can be solved in real time without the need for future information.
[0009] Prior art [D. Yang, J. Wang, F. Wu, L. Xiao, Y. Xu and T. Zhang, "Energy Efficient Transmission Strategy for Mobile Edge Computing Network in UAV-Based Patrol Inspection System," in IEEE Transactions on Mobile Computing, vol. 23, no. 5, pp. 5984-5998, May 2024, doi: 10.1109 / TMC.2023.3315477.] Yang et al. studied the energy efficient transmission strategy for mobile edge computing network in UAV-Based Patrol Inspection System. The goal is to minimize the total energy consumption by jointly optimizing the task completion time, communication scheduling, computing resource allocation and UAV trajectory.
[0010] Prior art [PA Apostolopoulos, G. Fragkos, EET siropoulou and S. Papavassiliou,"Data Offloading in UAV-Assisted Multi-Access Edge Computing Systems Under Resource Uncertainty," in IEEE Transactions on Mobile Computing, vol. 22, no. 1, pp. 175-190, 1Jan. 2023, doi: 10.1109 / TMC.2021.3069911.] Apostolopoulos et al. studied a MEC network supported by drones, and used game theory to model and analyze user computing offloading strategies and MEC server deployment, considering multi-user computing offloading and edge server deployment, with the goal of minimizing system-level computing costs in a dynamic environment.
[0011] However, many edge computing scenarios are dynamically changing over time, such as real-time video analysis, which means that computing tasks arrive randomly, the computing requirements of vehicles are time-varying, and vehicles are moving dynamically. Most of the existing UAV MEC research focuses on UAV trajectory optimization, without considering that vehicles are also moving dynamically. Summary of the invention
[0012] In order to solve the problem that most existing UAV MEC research focuses on UAV trajectory optimization, without considering the technical problem that vehicles also move dynamically, the embodiment of the present invention provides a method and device for unloading UAV tasks based on deep reinforcement learning. The technical solution is as follows:
[0013] On the one hand, a method for unloading a UAV task based on deep reinforcement learning is provided, the method is implemented by a UAV task unloading device, and the method includes:
[0014] S1. Construct a drone MEC system model; wherein the drone MEC system model includes multiple dynamically moving vehicles and multiple edge servers, and the multiple edge servers include multiple drones and multiple intelligent roadside facilities.
[0015] S2. After a task is generated, any one of the multiple dynamically moving vehicles determines, based on the computing capability of the vehicle, whether the generated task is processed by the vehicle or the task is offloaded to an edge server for processing by the edge server.
[0016] S3. When the task is processed by the edge server, a Gauss-Markov mobility model of the vehicle is constructed, the vehicle position is obtained according to the Gauss-Markov mobility model of the vehicle, and the distance between the vehicle and the intelligent roadside facility is obtained according to the vehicle position.
[0017] Determine whether the drone’s computing resources are sufficient to complete the mission.
[0018] If it is sufficient, the vehicle will offload the task to the drone, which will perform the task. The obtained drone group size and the preset number of iterations will be input into the DGBCO model, and the path planning result of the drone will be output.
[0019] If it is not enough, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if it is less, the vehicle will unload the task to the intelligent roadside facility, and the intelligent roadside facility will perform the task; if it is not less than, the drone will be used as a relay node to forward the task to the intelligent roadside facility for execution.
[0020] Among them, drones are used as relay nodes to forward tasks to intelligent roadside facilities for execution, including:
[0021] S31. Based on the computing reuse technology, tasks are preferentially assigned to edge servers with high request probabilities.
[0022] S32. Construct a UAV energy constraint problem based on Lyapunov optimization.
[0023] S33. The MADDPG algorithm is used to solve the UAV energy constraint problem and obtain the task offloading strategy. According to the task offloading strategy, the UAV is used as a relay node to forward the task to the intelligent roadside facilities for execution. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
[0024] Optionally, a Gauss-Markov mobility model of the vehicle is constructed, a vehicle position is obtained according to the Gauss-Markov mobility model of the vehicle, and a distance between the vehicle and the intelligent roadside facility is obtained according to the vehicle position, including:
[0025] Construct the vehicle's speed update at time slot t+1 as shown in the following equation (1):
[0026]
[0027] In the formula, v v (t+1) represents the speed of the vehicle at time slot t+1, α represents the memory level, represents the velocity vector at time slot t, represents the speed of the vehicle in the x-axis direction, represents the speed of the vehicle in the y-axis direction, represents the asymptotic mean of the velocity, W v (t) represents an uncorrelated random Gaussian process.
[0028] The vehicle position update at time slot t+1 is constructed based on the speed update, as shown in equation (2):
[0029] P v [t+1]=P v [t]+v v [t] (2)
[0030] Where P v [t+1] represents the vehicle position at time slot t+1, P v [t]=(x v [t],y v [t]) represents the position of vehicle v at time slot t, x v [t] represents the position of the vehicle in the x-axis direction, y v [t] represents the position of the vehicle in the y-axis direction.
[0031] The distance between the vehicle and the intelligent roadside facility is constructed based on the vehicle position update, as shown in the following formula (3):
[0032]
[0033] Where, d v,r (t) represents the distance between the vehicle and the intelligent roadside facility at time slot t, x r Indicates the position of the intelligent roadside facility in the x-axis direction, y r Indicates the position of the intelligent roadside facility on the y-axis.
[0034] Optionally, the computing reuse technology in S31 includes:
[0035] Obtain each vehicle's comments on each task, transmit the comments to the intelligent roadside device, obtain the comment sentiment value of each task based on the comments and the BERT model, normalize the comment sentiment values of all tasks, and obtain the request probability matrix of the task.
[0036] Optionally, the construction in S32 is based on the Lyapunov optimization of the UAV energy constraint problem, including:
[0037] Define the first virtual energy queue Represents the calculation energy queue of time slot t, and defines the second virtual energy queue represents the propulsion energy queue at time slot t.
[0038] According to the first virtual energy queue And the second virtual energy queue Define the Lyapunov function L(Qu (t)), the Lyapunov function represents a scalar measure of queue backlog.
[0039] According to the Lyapunov function, the conditional Lyapunov drift ΔL(Q u (t)).
[0040] Penalize drift and construct an upper bound on the drift penalty.
[0041] According to the upper bound, the UAV energy constraint problem based on Lyapunov optimization is obtained as shown in the following formula (4):
[0042]
[0043] In the formula, A t represents the task offloading strategy for time slot t, F t represents the computing resource allocation of time slot t, W t Denotes the communication resource allocation for time slot t, D′ u =D u (t+1) represents the position of the vehicle and the UAV at time slot t+1, D u (t+1) represents the position of the vehicle and the UAV at time slot t, represents the computing energy of the drone in time slot t, represents the propulsion energy of the UAV in time slot t, H represents the parameter that weighs the total cost and queue stability, V represents the set of all vehicles, and C(t) represents the mission cost.
[0044] Optionally, the vehicle state in the MADDPG algorithm in S33 is as shown in the following equation (5):
[0045]
[0046] In the formula, represents the task status, t represents the time slot, z v (t) represents the data size of the task, c v (t) represents the computational intensity of the task, τ v (t) represents the priority of the task, represents the normalized channel gain state, represents the transmission budget, Indicates the maximum computing capacity of the vehicle, represents the local resource allocation budget, Indicates the maximum computing resource of the vehicle.
[0047] Optionally, the behavior in the MADDPG algorithm in S33 includes client proxy behavior and master proxy behavior.
[0048] Among them, the client proxy behavior is shown in the following formula (6):
[0049]
[0050] In the formula, A c (t) represents the client proxy behavior, a v (t) represents the decision of vehicle v to unload the task, p c,v (t) represents the client action that determines the transmission power, f c,v (t) represents the action that determines the allocation of local computing resources, θ v represents the parameterized policy function of the client agent, S v (t) represents the state of vehicle v, t represents the time slot, and V represents the set of all vehicles.
[0051] The master agent behavior is shown in equation (7):
[0052]
[0053] In the formula, A m (t) represents the master agent behavior, x m,v (t) represents the decision of the master agent regarding the client agent at time slot t, represents the main policy, S represents the state set of all client agents, A represents the set of actions of all client agents, S c Represents the state set of the client agent, A c Represents a set of actions for a client agent.
[0054] Optionally, the reward function in the MADDPG algorithm in S33 is as shown in the following formula (8):
[0055]
[0056] Where r(S(t), A(t)) represents the reward function, S(t) represents the state set at time slot t, A(t) represents the action set at time slot t, V represents the set of all vehicles, and P′ represents the energy constraint problem of the UAV based on Lyapunov optimization.
[0057] On the other hand, a UAV task offloading device based on deep reinforcement learning is provided, and the device is applied to a UAV task offloading method based on deep reinforcement learning, and the device comprises:
[0058] A building module is used to build a drone MEC system model; wherein the drone MEC system model includes multiple dynamically moving vehicles and multiple edge servers, and the multiple edge servers include multiple drones and multiple intelligent roadside facilities.
[0059] The judgment module is used for any vehicle among a plurality of dynamically moving vehicles to judge, after a task is generated, whether the generated task is processed by the vehicle or the task is offloaded to an edge server for processing by the edge server according to the computing capacity of the vehicle.
[0060] The output module is used to construct a Gauss-Markov mobility model of the vehicle when the task is processed by the edge server, obtain the vehicle position according to the Gauss-Markov mobility model of the vehicle, and obtain the distance between the vehicle and the intelligent roadside facility according to the vehicle position.
[0061] Determine whether the drone’s computing resources are sufficient to complete the mission.
[0062] If it is sufficient, the vehicle will offload the task to the drone, which will perform the task. The obtained drone group size and the preset number of iterations will be input into the DGBCO model, and the path planning result of the drone will be output.
[0063] If it is not enough, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if it is less, the vehicle will unload the task to the intelligent roadside facility, and the intelligent roadside facility will perform the task; if it is not less than, the drone will be used as a relay node to forward the task to the intelligent roadside facility for execution.
[0064] Among them, drones are used as relay nodes to forward tasks to intelligent roadside facilities for execution, including:
[0065] S31. Based on the computing reuse technology, tasks are preferentially assigned to edge servers with high request probabilities.
[0066] S32. Construct a UAV energy constraint problem based on Lyapunov optimization.
[0067] S33. The MADDPG algorithm is used to solve the UAV energy constraint problem and obtain the task offloading strategy. According to the task offloading strategy, the UAV is used as a relay node to forward the task to the intelligent roadside facilities for execution. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
[0068] On the other hand, a drone task offloading device is provided, which includes: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned drone task offloading methods based on deep reinforcement learning is implemented.
[0069] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned methods for unloading drone tasks based on deep reinforcement learning.
[0070] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0071] In the present invention, efficient task offloading: through the collaborative work of vehicles and drones, efficient task offloading is achieved and the utilization of computing resources is optimized.
[0072] Reduce latency and energy consumption: The introduction of computing reuse and caching technology reduces the waste of computing resources, and optimizes the task offloading strategy, computing resources and communication resource allocation through the Lyapunov optimization method, significantly reducing task processing latency and system energy consumption.
[0073] Improve system throughput: The multi-agent deep reinforcement learning algorithm (MADDPG) improves the task completion rate and system throughput in complex Internet of Vehicles environments.
[0074] Enhanced robustness: Compared with traditional deep reinforcement learning algorithms, the algorithm of the present invention exhibits better robustness under high load and dynamic scenarios.
[0075] Flexibility and efficiency: The dual-stage bee swarm optimization algorithm (DGBCO) demonstrates excellent flexibility and efficiency in UAV trajectory planning.
[0076] Theoretical basis and technical support: It provides a new solution for drone-assisted multi-access edge computing (MEC) of the Internet of Vehicles, and provides a theoretical basis and technical support for resource optimization and scheduling in future Internet of Vehicles application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0078] Figure 1 It is a flow chart of a method for unloading UAV tasks based on deep reinforcement learning provided by an embodiment of the present invention;
[0079] Figure 2 It is a design diagram of a UAV task offloading system based on deep reinforcement learning provided by an embodiment of the present invention;
[0080] Figure 3 It is a scene diagram of the drone MEC system provided by an embodiment of the present invention;
[0081] Figure 4 It is a block diagram of a drone task offloading device based on deep reinforcement learning provided by an embodiment of the present invention;
[0082] Figure 5 It is a structural schematic diagram of a UAV task offloading device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0083] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0084] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0085] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0086] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0087] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0088] The embodiment of the present invention provides a method for unloading UAV tasks based on deep reinforcement learning. The method can be implemented by a UAV task unloading device, which can be a terminal or a server. Figure 1 , Figure 2 The flowchart of the method for unloading UAV tasks based on deep reinforcement learning is shown in the figure. The processing flow of the method may include the following steps:
[0089] S1. Construct a drone MEC system model; wherein the drone MEC system model includes multiple dynamically moving vehicles and multiple edge servers, and the multiple edge servers include multiple drones and multiple intelligent roadside facilities.
[0090] In a feasible implementation, Figure 3As shown, the present invention designs a random drone MEC system model, which includes multiple drones and multiple dynamically moving ground vehicles. The drones act as air edge servers to provide computing services for vehicles with time-varying computing requirements. The system model takes into account energy and resource constraints.
[0091] Specifically, in the proposed IoV multi-device task offloading framework, it is assumed that SRI (Smart Roadside Infrastructure) is deployed as an edge server beside urban roads, and drones act as edge servers or relay nodes over the city to provide services for vehicles. In the simulation environment, the present invention assumes that the vehicle application chooses to communicate with the server within its communication range.
[0092] Furthermore, each task is modeled as an independent and indivisible object that can be completely processed by the vehicle terminal or the edge server. When a task is generated, the vehicle terminal decides whether to process the task locally or offload it to the edge server based on the task offloading strategy. The present invention designs and analyzes the communication and computing model in the vehicle network-drone-edge server environment, with the goal of minimizing the system's latency, energy consumption, and total cost.
[0093] S2. After a task is generated, any one of the multiple dynamically moving vehicles determines, based on the computing capability of the vehicle, whether the generated task is processed by the vehicle or the task is offloaded to an edge server for processing by the edge server.
[0094] S3. When the task is processed by the edge server, a Gauss-Markov mobility model of the vehicle is constructed, the vehicle position is obtained according to the Gauss-Markov mobility model of the vehicle, and the distance between the vehicle and the intelligent roadside facility is obtained according to the vehicle position.
[0095] In a feasible implementation, the importance of mobility models needs to be carefully considered because they describe the way mobile devices move, including direction, speed, and acceleration, all of which are critical for making appropriate scheduling decisions. In the context of the vehicle-based edge computing offloading model, in order to optimize latency and energy consumption, tasks are offloaded to nearby edge servers, which requires the use of accurate mobility models to predict the location of the vehicle-based terminal and offload tasks to the nearest edge server when appropriate.
[0096] Specifically, the mobility of the vehicle is modeled as a Gauss-Markov mobility model, which is widely used in cellular communication networks. Specifically, the speed of vehicle v at time slot t+1 is updated as:
[0097]
[0098] In the formula, v v(t+1) represents the speed of the vehicle at time slot t+1, α represents the memory level, which reflects the degree of time dependence, represents the velocity vector at time slot t, represents the speed of the vehicle in the x-axis direction, represents the speed of the vehicle in the y-axis direction, represents the asymptotic mean of the velocity, W v (t) represents an uncorrelated random Gaussian process where σ v represents the asymptotic standard deviation of the velocity.
[0099] Similarly, the speed of drone u at time slot t+1 is updated as:
[0100]
[0101] Therefore, the vehicle position at time slot t+1 is updated as:
[0102] P v [t+1]=P v [t]+v v [t] (3)
[0103] Where P v [t+1] represents the vehicle position at time slot t+1, P v [t]=(x v [t],y v [t]) represents the position of vehicle v at time slot t, x v [t] represents the position of the vehicle in the x-axis direction, y v [t] represents the position of the vehicle in the y-axis direction, P v [1] represents the initial position of vehicle v. At the beginning of time slot t, the present invention assumes that the vehicle position feedback knows the current position P v [t], but the future trajectory is currently unknown.
[0104] Similarly, the position of the drone at time slot t+1 is updated as:
[0105] P u [t+1]=P u [t]+v u [T] (4)
[0106] When the distance between the vehicle and the edge server is d v,r (t) is less than the longest communication distance d max When , the vehicle directly offloads the task to the edge server:
[0107]
[0108] Where, d v,r (t) represents the distance between the vehicle and the intelligent roadside facility at time slot t, x r Indicates the position of the intelligent roadside facility in the x-axis direction, y r Indicates the position of the intelligent roadside facility on the y-axis.
[0109] When the distance is greater than the longest communication distance, it involves the trajectory optimization problem of the vehicle and the drone:
[0110]
[0111] The distance between drones and smart roadside facilities is similar:
[0112]
[0113] Furthermore, it is determined whether the computing resources of the drone are sufficient to complete the task.
[0114] If it is sufficient, the vehicle will offload the task to the drone, which will perform the task. The obtained drone group size and the preset number of iterations will be input into the DGBCO model, and the path planning result of the drone will be output.
[0115] In one possible implementation, the proposed model combines different UAV path planning metaheuristic algorithms. In this invention, two phases of the DGBCO algorithm are discussed in detail, namely the exploration phase and the exploitation phase for solving complex path planning environments. Initialization DGBCO randomly assigns the position of particles in the entire search space. The equation used to specify the initial value is:
[0116] X(i)-Min X +rand*(Max X -Min X )(8)
[0117] Where X(i) is the particle position, Min X is the minimum value, Max X is the maximum value in the search space.
[0118] Furthermore, since DGBCO has sub-groups in the exploration and development stage, accounting for 70%-30% of the total participants. In the initial stage, a larger search space is needed to explore the global optimum, so the exploration group is 70% and the mining group is 30%. With iterations, the development group will increase to carry out more development.
[0119] The swarm plays an important role in finding the global optimal solution in the entire search space. Therefore, the exploration phase of the algorithm needs to effectively solve the engineering problem. In DGBCO, the positions of the particles in the exploration swarm are updated through two techniques.
[0120] The first is to search for the best solution around the current position, using the following formula:
[0121]
[0122] Where X(i) and X(i+1) are the current and updated positions of the drone. r1 and r2 are numbers from a random normal distribution. d is the distance of the solution search. The second method is mutation, which uses genetic operators to create and maintain diversity in the population.
[0123] In this group, particles find the optimal solution of the group and converge to the global optimal solution. Therefore, two strategies are also used in the utilization phase to update the particle position. In the first strategy, particles search for the global optimal solution in a ring form. This technology can be used in DGBCO:
[0124]
[0125] Where G(i) represents the optimal solution and r3 is a random number. The second method to update the particle position is to search around the global optimal solution.
[0126] This technique can be used:
[0127]
[0128] where r4 and r5 are randomly generated numbers k that decrease exponentially from 2 to 0.
[0129] Elitism is another property used by DGBCO, where the global best is found in each iteration and passed to the next iteration.
[0130] The pseudo code of DGBCO is shown in Algorithm 1. The UAV population and the number of iterations are taken as input parameters. First, the population is initialized, and the fitness value of each UAV is evaluated to obtain the optimal solution until the maximum iteration is reached. Next, if the optimal solution remains unchanged in the first two iterations, the position of the UAV in the group is updated and the number of particles in the exploration group is increased. The present invention considers both the exploration group and the development group to update the previous fitness value. For the exploration group, the parameters r1, r2 and P are updated, and the solution mutates when P≥0.5, otherwise the search is performed around the current solution.
[0131] For the development group, the parameters r2, r3, r4 and P are updated, if P ≥ 0.5, it will identify the best solution, otherwise it will search for the best solution.
[0132]
[0133]
[0134] If it is not enough, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if it is less, the vehicle will unload the task to the intelligent roadside facility, and the intelligent roadside facility will perform the task; if it is not less than, the drone will be used as a relay node to forward the task to the intelligent roadside facility for execution.
[0135] Among them, drones are used as relay nodes to forward tasks to intelligent roadside facilities for execution, including:
[0136] S31. Based on the computing reuse technology, tasks are preferentially assigned to edge servers with high request probabilities.
[0137] Optionally, the computing reuse technology in S31 includes:
[0138] Obtain each vehicle's comments on each task, transmit the comments to the intelligent roadside device, obtain the comment sentiment value of each task based on the comments and the BERT model, normalize the comment sentiment values of all tasks, and obtain the request probability matrix of the task.
[0139] In a feasible implementation, in the prior art, the conditions of the vehicle mission content are not carefully considered. Usually, it is assumed that the probability of the vehicle mission content follows a uniform distribution (P∝1 / N) or a Zipf distribution (P∝1 / i γ ). However, the above distribution cannot represent the actual situation of these tasks. Therefore, in the present invention, the sentiment analysis of vehicle task feedback is used as a guide for obtaining the task request probability matrix.
[0140] like Figure 3 As shown in the figure, the feedback from the vehicles will first be collected and transmitted to the edge cloud server, where the feedback from the vehicles will be fully analyzed by the deep learning neural network. Then, the results of the sentiment analysis will be used as the basis for obtaining the probability matrix of task requests. Since it is assumed that the operation of task content caching can be completed in the idle time of the local device, the calculation speed of these neural networks does not need to be considered. The selection of the deep learning neural network depends only on the accuracy of its sentiment analysis results.
[0141] Commonly used text analysis methods include TextCNN, TextRNN, FastText, Bert, etc. TextCNN is suitable for short text analysis, while TextRNN can further consider the correlation between words that are far apart in a sentence. Compared with the first two methods, Fasttext can achieve faster calculation speed at the expense of accuracy. Bert uses the self-attention mechanism in Transformer to enable each word to interact with other words, thereby capturing richer contextual information and achieving higher accuracy, although it may take more computing time. Further experimental results based on the IMDB dataset are shown in Table 1, where the number of samples in the training set, validation set, and test set are 70,000, 15,000, and 15,000, respectively.
[0142] Table 1
[0143]
[0144] As mentioned above, the invention only considers the accuracy of sentiment analysis, without considering the calculation time caused by the idle time transmission strategy. Through simple theoretical analysis and rigorous experiments of the above methods, Bert is the most suitable method for sentiment analysis in the invention. Therefore, Bert is selected to obtain the user's request probability matrix to solve the request probability sub-problem. The specific steps are as follows:
[0145] First, the Bert model is used to perform sentiment analysis on the collected comments of each vehicle on each task to obtain the sentiment value of the comments on the corresponding task; the sentiment value is between 0 and 1. The closer it is to 1, the more interested the vehicle is in the corresponding task content, and vice versa.
[0146] Subsequently, the sentiment value of each vehicle is normalized for all task contents so that their sum is equal to 1. After normalization, the sentiment value of each user for each content is used as the request probability to obtain the request probability matrix. According to the request probability matrix, the tasks of the vehicles are preferentially assigned to the edge servers with high request probabilities. This ensures that the edge servers process the most interesting tasks first, improving the efficiency of task processing.
[0147] S32. Construct a UAV energy constraint problem based on Lyapunov optimization.
[0148] In a feasible implementation, since the drone energy constraint problem P depends on the future, an online method is needed to make real-time decisions without predicting the future. Similarly, the Lyapunov-based optimization framework is a commonly used online algorithm design method, which has the advantages of being simple and effective. To this end, the present invention first transforms the problem P into a slot-by-slot real-time optimization problem based on the Lyapunov optimization framework.
[0149] Optionally, the above step S32 may include the following steps S321-S326:
[0150] S321. To meet the energy constraints of the UAV, the first virtual energy queue is defined based on the Lyapunov optimization technology. Represents the calculation energy queue of time slot t, and defines the second virtual energy queue represents the propulsion energy queue of time slot t. The present invention assumes that the queue is in the initial time slot, i.e. Therefore, the virtual energy queue can be updated as:
[0151]
[0152] in and Represent the computation and propulsion energy budgets for each slot,
[0153] S322: According to the first virtual energy queue And the second virtual energy queue Define the Lyapunov function L(Q u (t)), the Lyapunov function represents a scalar measure of queue backlog:
[0154]
[0155] in is the vector of the current queue backlog.
[0156] S323, according to the Lyapunov function, define the conditional Lyapunov drift ΔL(Q u (t)):
[0157]
[0158] S324. Penalty for drifting:
[0159] D(Q u (t)) = ΔL(Q u (t))+HE{C S (t)|Q u (t)} (15)
[0160] in is the total cost of all vehicles at time slot t, and H is a parameter that weighs the total cost and queue stability.
[0161] S325. Construct an upper bound of drift plus penalty.
[0162] Theorem 1: For all t and all possible queue backlogs Q u(t), the upper bound of drift plus penalty is:
[0163]
[0164] in is a finite constant.
[0165] According to the Lyapunov optimization framework, the right side of the inequality is minimized. Therefore, the problem P that depends on future information is transformed into a real-time optimization problem P′ that can be solved only by current information, as shown below:
[0166]
[0167] In the formula, A t represents the task offloading strategy for time slot t, F t represents the computing resource allocation of time slot t, W t Denotes the communication resource allocation for time slot t, D′ u =D u (t+1) represents the position of the vehicle and the UAV at time slot t+1, D u (t+1) represents the position of the vehicle and the UAV at time slot t, represents the computing energy of the drone in time slot t, represents the propulsion energy of the UAV in time slot t, H represents the parameter that weighs the total cost and queue stability, V represents the set of all vehicles, and C(t) represents the mission cost.
[0168]
[0169] S33. The MADDPG algorithm is used to solve the UAV energy constraint problem and obtain the task offloading strategy. According to the task offloading strategy, the UAV is used as a relay node to forward the task to the intelligent roadside facilities for execution. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
[0170] In a feasible implementation, in order to solve the optimization problem of cost minimization in equation (17), the present invention transforms the optimization problem into a reward maximization problem and applies LyBMADRL_MEC. The expressions of state, behavior and reward function are as follows:
[0171] S331, state: The state S(t) of the MEC environment at time t, which includes the state set of vehicle v, described as Constant values such as the number of subchannels K, the number of processing units Ue on the server, the processing capacity fe, and the storage capacity ze of the server are not included in the state information. v (t) is composed of four parts: task status Normalized channel gain state Transmission Budget Local resource allocation budget Defined in the equation:
[0172]
[0173] In the formula, represents the task status, t represents the time slot, z v (t) represents the data size of the task, c v (t) represents the computational intensity of the task, τ v (t) represents the priority of the task, Represents the normalized channel gain state, represented by g n Decision, g n =h n / σ 2 , g n represents the normalized channel gain state, h n represents the channel gain, σ 2 represents the background noise variance, represents the transmission budget, Indicates the maximum computing capacity of the vehicle, represents the local resource allocation budget, Indicates the maximum computing resource of the vehicle.
[0174] S332, Action: At the beginning of each time step, the UDs terminal devices use the client agent to decide their resource allocation. Then, the SDN controller collects the status and action information of the UDs and performs one of the following three processes through the master agent: 1) For the UDs that decide to be processed locally, the server does not intervene. 2) If the number of UDs proposed for offloading is greater than the number of subchannels, or the sum of the sizes of the subchannels is greater than the storage capacity of the server, the server will make a combined decision on which requests are approved and which requests are rejected. 3) If the proposed requests are less than the constraints, the server accepts all requests. Finally, subchannels are assigned to the accepted UDs, and then the task offloading and processing process begins.
[0175] Existing deep reinforcement learning-based task offloading algorithms include the number of subchannels in their state and action spaces. However, from the perspective of one UD, the transmission capacity of the subchannels is equal. If a channel is restricted to be used by only one UD at a time, it does not matter which channel a UD uses. Therefore, including subchannels in the state and action space creates a dimensionality problem but does not play any important role. The present invention excludes channel information from the state and action space and uses it as a constraint for the selection of combined actions. Note that by removing the constraint that a channel must be used by one UD, a channel can be reused by multiple UDs one after another. In this case, only the storage capacity of the server becomes a constraint in the selection of combined actions.
[0176] The operation of the client agent and the master agent is as follows.
[0177] Client behavior. At each time step, each client agent generates three actions, which are all continuous valued actions in the range [0,1]. The action space can be represented as:
[0178]
[0179] In the formula, S v (t) is the state of vehicle v in equation (18), A c (t) represents the client proxy behavior, a v (t) represents the decision of vehicle v to unload the task, p c,v (t) represents the client action that determines the transmission power, f c,v (t) represents the action that determines the allocation of local computing resources, θ v represents the parameterized policy function of the client agent, S v (t) represents the state of vehicle v, t represents the time slot, and V represents the set of all vehicles.
[0180] The client agent determines p v and f v The actions are as follows:
[0181]
[0182] Further, master actions: For client actions, the master takes a combination of client state and action and provides a binary output for combined decision making on which should be distributed locally and which should be accepted by the MEC server for processing:
[0183]
[0184] In the formula, A m (t) represents the master agent behavior, x m,v(t) represents the decision of the master agent regarding the client agent at time slot t, represents the main policy, S represents the state set of all client agents, A represents the set of actions of all client agents, S c Represents the state set of the client agent, A c Represents a set of actions for a client agent.
[0185] S333, reward function: Since the design of the cooperative learning formula of the present invention is based on the cost minimization problem in formula (17), the system reward function of the present invention is equal to the negative value of the system cost function. Therefore, the system reward function can be expressed as:
[0186]
[0187] Where r(S(t), A(t)) represents the reward function, S(t) represents the state set at time slot t, A(t) represents the action set at time slot t, V represents the set of all vehicles, and P′ represents the energy constraint problem of the UAV based on Lyapunov optimization.
[0188] The power and resource allocation decisions made by vehicle v in the current step affect its operating lifetime in the next time step by affecting its energy consumption. Therefore, DRL must consider both immediate and long-term rewards using the Bellman equation, as shown in the formula:
[0189]
[0190] In the formula is the learning rate, γ is the discount factor, Strategies for critics, is the target critic’s strategy, S = {S1,…,S N} and S′={S1′,…,S′ N} is the combination of the current state and the next state, A={A1,…,A N} and A′={A′1,…,A′ N} is the combination of the current and next actions of the client agent, S′ N and A′ N The state and action corresponding to the agent.
[0191] Furthermore, the multi-agent deep reinforcement learning algorithm learns how to maximize global and local rewards. It contains two main critics, one is the global critic, which is shared by all agents and inputs the states and behaviors of all agents to maximize the global reward; the other is the local critic, which receives the states and behaviors of a specific agent and evaluates its reward. In addition, the present invention also considers the impact of policy approximation error and value update on the global critic of the MADPG algorithm. In order to improve the MADPG algorithm, the present invention adopts a double-delayed deterministic policy gradient instead of the global critic.
[0192] Specifically, the present invention considers a vehicle environment with V vehicles (agents), and the strategies of all agents are π = {π1, π2, ..., π V}. The policy π of vehicle v v , Q function and double-delay double-Q value functions are represented by θ k , ψ1, ψ2 are parameterized. For each agent, the modified policy gradient can be written as:
[0193]
[0194] where s=(s1,s2,…,s V ), a=(a1,a2,…,a V ) is the total state and action vector, D is the replay buffer, a v =π v (s v ) is the agent v according to its own strategy π v The action selected. Then the double-delay double Q value function is updated as:
[0195]
[0196] in
[0197] The detailed MARL algorithm is shown in Table 3.
[0198]
[0199]
[0200] In a feasible implementation, in the model of the present invention, after the task is generated, the vehicle will select an execution strategy based on its computing power. If the local computing power of the vehicle is sufficient, the task will be processed locally; if the computing resources are insufficient, the vehicle will offload the task to the edge server served by the drone or intelligent roadside facilities. At this time, the communication and computing resources between the vehicle and the edge server need to be considered. When the computing resources of the drone are sufficient, the drone is preferred to perform the task because the drone can dynamically adjust its position to keep pace with the vehicle. When the drone resources are insufficient or the edge server is far away, the drone will act as a relay node to forward the task to the SRI.
[0201] Furthermore, to ensure the efficiency of task processing, the edge server will schedule tasks based on their priorities when queuing them. High-priority tasks are processed first, and when the priorities are the same, a first-in-first-out queue mechanism is used. In addition, the edge server can use computation reuse technology to store processed computation results in advance, and use the LSH (Locality Sensitive Hashing) algorithm to achieve approximate matching and accelerate the processing of similar tasks.
[0202] Specifically, first, the vehicle generates a task. When the vehicle has sufficient computing memory, the task is calculated directly locally in the vehicle. When the vehicle memory is insufficient, the vehicle will offload the task to a drone or an edge server played by an intelligent roadside facility. At this time, the communication between the vehicle and the edge server and the computing resources of the drone and the edge server must be considered. When the computing resources of the drone are sufficient, the drone should be given priority because the drone can be synchronized with the movement of the vehicle in real time. When the drone resources are insufficient, the intelligent roadside facilities will be considered. When the drone resources are insufficient and the edge server is too far away, the vehicle will use the drone as a relay node to offload the task to the edge server played by the intelligent roadside facility.
[0203] This paper proposes a collaborative optimization algorithm based on deep reinforcement learning to solve the problem of task offloading and resource allocation. By introducing computational reuse and joint caching technology, combined with the Lyapunov optimization method, it further addresses the problem of long-term energy constraints.
[0204] Furthermore, the effectiveness of the proposed method is verified through experiments in a simulated dynamic environment of the Internet of Vehicles. The results show that the method has significant advantages in reducing the total system cost and improving task processing efficiency, and exhibits good robustness in complex environments.
[0205] In the embodiment of the present invention, efficient task offloading: through the collaborative work of the vehicle and the drone, efficient task offloading is achieved and the utilization of computing resources is optimized.
[0206] Reduce latency and energy consumption: The introduction of computing reuse and caching technology reduces the waste of computing resources, and optimizes the task offloading strategy, computing resources and communication resource allocation through the Lyapunov optimization method, significantly reducing task processing latency and system energy consumption.
[0207] Improve system throughput: The multi-agent deep reinforcement learning algorithm (MADDPG) improves the task completion rate and system throughput in complex Internet of Vehicles environments.
[0208] Enhanced robustness: Compared with traditional deep reinforcement learning algorithms, the algorithm of the present invention exhibits better robustness under high load and dynamic scenarios.
[0209] Flexibility and efficiency: The dual-stage bee swarm optimization algorithm (DGBCO) demonstrates excellent flexibility and efficiency in UAV trajectory planning.
[0210] Theoretical basis and technical support: It provides a new solution for drone-assisted multi-access edge computing (MEC) of the Internet of Vehicles, and provides a theoretical basis and technical support for resource optimization and scheduling in future Internet of Vehicles application scenarios.
[0211] Figure 4 1 is a block diagram of a drone task offloading device based on deep reinforcement learning according to an exemplary embodiment, wherein the device is used in a drone task offloading method based on deep reinforcement learning. Figure 4 The device includes a construction module 310, a judgment module 320 and an output module 330. Among them:
[0212] The construction module 310 is used to construct a drone MEC system model; wherein the drone MEC system model includes multiple dynamically moving vehicles and multiple edge servers, and the multiple edge servers include multiple drones and multiple intelligent roadside facilities.
[0213] The judgment module 320 is used for any vehicle among the multiple dynamically moving vehicles to judge whether the generated task is processed by the vehicle or to offload the task to the edge server for processing according to the computing capacity of the vehicle after the task is generated.
[0214] The output module 330 is used to construct a Gauss-Markov mobility model of the vehicle when the task is processed by the edge server, obtain the vehicle position according to the Gauss-Markov mobility model of the vehicle, and obtain the distance between the vehicle and the intelligent roadside facility according to the vehicle position.
[0215] Determine whether the drone’s computing resources are sufficient to complete the mission.
[0216] If it is sufficient, the vehicle will offload the task to the drone, which will perform the task. The obtained drone group size and the preset number of iterations will be input into the DGBCO model, and the path planning result of the drone will be output.
[0217] If it is not enough, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if it is less, the vehicle will unload the task to the intelligent roadside facility, and the intelligent roadside facility will perform the task; if it is not less than, the drone will be used as a relay node to forward the task to the intelligent roadside facility for execution.
[0218] Among them, drones are used as relay nodes to forward tasks to intelligent roadside facilities for execution, including:
[0219] S31. Based on the computing reuse technology, tasks are preferentially assigned to edge servers with high request probabilities.
[0220] S32. Construct a UAV energy constraint problem based on Lyapunov optimization.
[0221] S33. The MADDPG algorithm is used to solve the UAV energy constraint problem and obtain the task offloading strategy. According to the task offloading strategy, the UAV is used as a relay node to forward the task to the intelligent roadside facilities for execution. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
[0222] In the embodiment of the present invention, efficient task offloading: through the collaborative work of the vehicle and the drone, efficient task offloading is achieved and the utilization of computing resources is optimized.
[0223] Reduce latency and energy consumption: The introduction of computing reuse and caching technology reduces the waste of computing resources, and optimizes the task offloading strategy, computing resources and communication resource allocation through the Lyapunov optimization method, significantly reducing task processing latency and system energy consumption.
[0224] Improve system throughput: The multi-agent deep reinforcement learning algorithm (MADDPG) improves the task completion rate and system throughput in complex Internet of Vehicles environments.
[0225] Enhanced robustness: Compared with traditional deep reinforcement learning algorithms, the algorithm of the present invention exhibits better robustness under high load and dynamic scenarios.
[0226] Flexibility and efficiency: The dual-stage bee swarm optimization algorithm (DGBCO) demonstrates excellent flexibility and efficiency in UAV trajectory planning.
[0227] Theoretical basis and technical support: It provides a new solution for drone-assisted multi-access edge computing (MEC) of the Internet of Vehicles, and provides a theoretical basis and technical support for resource optimization and scheduling in future Internet of Vehicles application scenarios.
[0228] Figure 5 is a structural diagram of a drone task offloading device provided by an embodiment of the present invention, such as Figure 5 As shown, the drone task offloading device may include the above Figure 4 The drone task offloading device based on deep reinforcement learning is shown. Optionally, the drone task offloading device 410 may include a first processor 2001.
[0229] Optionally, the drone task offloading device 410 may also include a memory 2002 and a transceiver 2003 .
[0230] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0231] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0232] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0233] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0234] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0235] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0236] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0237] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0238] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0240] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0241] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for unloading UAV tasks based on deep reinforcement learning, characterized in that: The method comprises: S1. Constructing a drone MEC system model; wherein the drone MEC system model includes multiple dynamically moving vehicles and multiple edge servers, and the multiple edge servers include multiple drones and multiple intelligent roadside facilities; S2. After a task is generated, any one of the multiple dynamically moving vehicles determines, based on the computing capability of the vehicle, whether the generated task is processed by the vehicle or offloads the task to an edge server for processing by the edge server; S3. When the task is processed by the edge server, a Gauss-Markov mobility model of the vehicle is constructed, a vehicle position is obtained according to the Gauss-Markov mobility model of the vehicle, and a distance between the vehicle and the intelligent roadside facility is obtained according to the vehicle position; Determining whether the computing resources of the drone are sufficient to complete the task; If it is sufficient, the vehicle offloads the task to the drone, which performs the task, and inputs the obtained drone group size and the preset number of iterations into the DGBCO model, and outputs the path planning result of the drone; If not, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if less, the vehicle unloads the task to the intelligent roadside facility, and the intelligent roadside facility performs the task; if not, use the drone as a relay node to forward the task to the intelligent roadside facility for execution; The method of using a drone as a relay node to forward the task to the intelligent roadside facility for execution includes: S31, according to the computing reuse technology, preferentially assigning the task to the edge server with a high request probability; S32, construct the UAV energy constraint problem based on Lyapunov optimization; S33. The MADDPG algorithm is used to solve the energy constraint problem of the UAV to obtain a task offloading strategy. According to the task offloading strategy, the task is forwarded to the intelligent roadside facility for execution by using the UAV as a relay node. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
2. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The constructing of the Gauss-Markov mobility model of the vehicle, obtaining the vehicle position according to the Gauss-Markov mobility model of the vehicle, and obtaining the distance between the vehicle and the intelligent roadside facility according to the vehicle position include: Construct the vehicle's speed update at time slot t+1 as shown in the following equation (1): In the formula, v v (t+1) represents the speed of the vehicle at time slot t+1, α represents the memory level, represents the velocity vector at time slot t, represents the speed of the vehicle in the x-axis direction, represents the speed of the vehicle in the y-axis direction, represents the asymptotic mean of the velocity, W v (t) represents an uncorrelated random Gaussian process; The vehicle position update at time slot t+1 is constructed based on the speed update, as shown in equation (2): P v [t+1]=P v [t]+v v [t] (2) Where P v [t+1] represents the vehicle position at time slot t+1, P v [t]=(x v [t],y v [t]) represents the position of vehicle v at time slot t, x v [t] represents the position of the vehicle in the x-axis direction, y v [t] represents the position of the vehicle in the y-axis direction; The distance between the vehicle and the intelligent roadside facility is constructed based on the vehicle position update, as shown in the following formula (3): Where, d v,r (t) represents the distance between the vehicle and the intelligent roadside facility at time slot t, x r Indicates the position of the intelligent roadside facility in the x-axis direction, y r Indicates the position of the intelligent roadside facility on the y-axis.
3. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The computing reuse technology in S31 includes: Obtain each vehicle's comments on each task, transmit the comments to the intelligent roadside device, obtain the comment sentiment value of each task based on the comments and the BERT model, normalize the comment sentiment values of all tasks, and obtain the request probability matrix of the task.
4. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The construction of S32 is based on Lyapunov optimization of the UAV energy constraint problem, including: Define the first virtual energy queue Represents the calculation energy queue of time slot t, and defines the second virtual energy queue represents the propulsion energy queue at time slot t; According to the first virtual energy queue And the second virtual energy queue Define the Lyapunov function L(Q u (t)), the Lyapunov function represents a scalar measure of queue backlog; According to the Lyapunov function, the conditional Lyapunov drift ΔL(Q u (t)); Penalize drift and construct an upper bound on the drift penalty. According to the upper bound, the UAV energy constraint problem based on Lyapunov optimization is obtained as shown in the following formula (4): In the formula, A t represents the task offloading strategy for time slot t, F t represents the computing resource allocation for time slot t, W t Denotes the communication resource allocation for time slot t, D′ u =D u (t+1) represents the position of the vehicle and the UAV at time slot t+1, D u (t+1) represents the position of the vehicle and the UAV at time slot t, represents the computing energy of the drone in time slot t, represents the propulsion energy of the UAV in time slot t, H represents the parameter that weighs the total cost and queue stability, V represents the set of all vehicles, and C(t) represents the mission cost.
5. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The vehicle state in the MADDPG algorithm in S33 is shown in the following formula (5): In the formula, represents the task status, t represents the time slot, z v (t) represents the data size of the task, c v (t) represents the computational intensity of the task, τ v (t) represents the priority of the task, represents the normalized channel gain state, represents the transmission budget, Indicates the maximum computing capacity of the vehicle, represents the local resource allocation budget, Indicates the maximum computing resource of the vehicle.
6. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The behaviors in the MADDPG algorithm in S33 include client agent behaviors and master agent behaviors; The client proxy behavior is as shown in the following formula (6): In the formula, A c (t) represents the client proxy behavior, a v (t) represents the decision of vehicle v to unload the task, p c,v (t) represents the client action that determines the transmission power, f c,v (t) represents the action that determines the allocation of local computing resources, θ v represents the parameterized policy function of the client agent, S v (t) represents the state of vehicle v, t represents the time slot, and V represents the set of all vehicles; The master agent behavior is shown in the following formula (7): In the formula, A m (t) represents the master agent behavior, x m,v (t) represents the decision of the master agent regarding the client agent at time slot t, represents the main policy, S represents the state set of all client agents, A represents the action set of all client agents, S c Represents the state set of the client agent, A c Represents a set of actions for a client agent.
7. The method for unloading unmanned aerial vehicle tasks based on deep reinforcement learning according to claim 1, characterized in that: The reward function in the MADDPG algorithm in S33 is shown in the following formula (8): In the formula, r(S(t), A(t)) represents the reward function, S(t) represents the state set at time slot t, A(t) represents the action set at time slot t, V represents the set of all vehicles, and P ′ Represents the UAV energy constraint problem based on Lyapunov optimization.
8. A drone task offloading device based on deep reinforcement learning, wherein the drone task offloading device based on deep reinforcement learning is used to implement the drone task offloading method based on deep reinforcement learning as claimed in any one of claims 1 to 7, characterized in that: The device comprises: A construction module for constructing a drone MEC system model; wherein the drone MEC system model includes a plurality of dynamically moving vehicles and a plurality of edge servers, and the plurality of edge servers include a plurality of drones and a plurality of intelligent roadside facilities; a judgment module, configured to judge, after a task is generated by any one of the plurality of dynamically moving vehicles, whether the generated task is to be processed by the vehicle or to offload the task to an edge server for processing by the edge server according to the computing capacity of the vehicle; an output module, for constructing a Gauss-Markov mobility model of the vehicle when the task is processed by the edge server, obtaining a vehicle position according to the Gauss-Markov mobility model of the vehicle, and obtaining a distance between the vehicle and the intelligent roadside facility according to the vehicle position; Determining whether the computing resources of the drone are sufficient to complete the task; If it is sufficient, the vehicle offloads the task to the drone, which performs the task, and inputs the obtained drone group size and the preset number of iterations into the DGBCO model, and outputs the path planning result of the drone; If not, determine whether the distance between the vehicle and the intelligent roadside facility is less than the longest communication distance; if less, the vehicle unloads the task to the intelligent roadside facility, and the intelligent roadside facility performs the task; if not, use the drone as a relay node to forward the task to the intelligent roadside facility for execution; The method of using a drone as a relay node to forward the task to the intelligent roadside facility for execution includes: S31, according to the computing reuse technology, preferentially assigning the task to the edge server with a high request probability; S32, construct the UAV energy constraint problem based on Lyapunov optimization; S33. The MADDPG algorithm is used to solve the energy constraint problem of the UAV to obtain a task offloading strategy. According to the task offloading strategy, the task is forwarded to the intelligent roadside facility for execution by using the UAV as a relay node. The global criticism of the MADDPG algorithm adopts a double-delay deterministic policy gradient.
9. A drone task unloading device, characterized in that: The drone task offloading device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Reliable vehicle-mounted edge calculation unloading method based on reinforcement learning
CN112929849A
Unmanned aerial vehicle assisted V2I network task offloading method
CN114650567A
URLLC (Uniform Resource Logical Link Control) perception air-ground Internet of Vehicles cooperative computing problem decoupling method
CN117412262A
MADDPG-based air-ground vehicle networking energy consumption minimum unloading method
CN118102254A
Edge calculation optimization method based on cooperative game and multi-target whale algorithm
CN118200878A