Service function deployment and UAV edge node resource allocation method based on deep reinforcement learning
Through a method based on deep reinforcement learning, a two-dimensional task scheduling model is built, and the resource allocation and service function deployment of drone edge nodes is optimized, which solves the problem of complex task scheduling in multi-UAV edge computing networks, and realizes low-cost and high-efficiency service function deployment.
Patent Information
- Application Number
- CN202410794724.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-06-19
AI Technical Summary
The prior art is difficult to optimize the deployment of drone edge nodes in multi-UAV edge computing networks and efficiently utilize their service resources. Especially when handling complex tasks, it is impossible to effectively schedule tasks composed of multiple subtasks.
Using a method based on deep reinforcement learning, a two-dimensional task scheduling model is built, state space and action space are set through the agent framework, and reward functions are designed to optimize resource allocation and service function deployment of drone edge nodes.
The number of drone edge nodes is minimized under the constraints on DAG execution performance, reducing deployment costs, and improving the stability and scope of application of the system.
Smart Images

Figure CN118784650B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle technology, and more specifically to a service function deployment based on deep reinforcement learning and a method for allocating resources of unmanned aerial vehicle edge nodes. Background Art
[0002] Equipping drones with communication, computing, and storage units as edge nodes can fully utilize the high mobility of drones and the proximal processing advantages of edge computing, and provide efficient integrated communication and computing services in scenarios where infrastructure is insufficient. This can be applied to scenarios such as disaster relief and emergency response. Ground equipment will offload tasks to drone edge nodes for processing via ground-to-air links, and return the processing results to the ground equipment, significantly reducing task processing latency and reducing the energy consumption of mobile devices.
[0003] At present, in order to solve the shortcomings of limited coverage and insufficient service capabilities of a single drone edge node, multiple drone edge nodes collaborate to form a multi-drone edge computing network, which can significantly improve coverage and service capabilities. How to optimize the deployment of drone edge nodes and efficiently utilize their service resources is the core of improving service efficiency.
[0004] However, current scheduling optimization methods for simple or indivisible single tasks cannot be used to schedule complex tasks composed of multiple subtasks; in addition, current methods often abstract task execution into several CPU cycles without considering the software and data environment (called service functions) on which they depend, which is contrary to the actual situation of software execution.
[0005] In this regard, the invention patent "Multi-UAV edge computing service deployment and scheduling method and system" with patent number ZL202110821000.8 designed a service-dependent and topology-aware microservice deployment method; the invention patent "Multi-UAV edge computing path optimization and dependent task scheduling optimization method and system" with patent number ZL202310255675.X determines the flight trajectory of the drone edge node based on the multi-agent reinforcement learning method, and decides whether to execute a given DAG task at each flight point. However, these methods can only handle a fixed number of drone edge nodes and cannot achieve the joint optimization of deployment cost and task processing quality.
[0006] Therefore, how to provide a service function deployment and drone edge node resource allocation method with low deployment cost and wide application range, so as to improve service quality and reduce costs, is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0007] In view of this, the present invention provides a service function deployment and drone edge node resource allocation method based on deep reinforcement learning to solve the technical problems mentioned in the background technology.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A method for service function deployment and drone edge node resource allocation based on deep reinforcement learning, comprising the following steps:
[0010] S101. Deployment problem modeling: Based on the deployment scenarios of ground tasks and drones, a two-dimensional task scheduling model is constructed; the deployment scenario is that multiple mobile devices are distributed in the area, each mobile device has a computing-intensive task DAG, and the DAG task is offloaded to the drone edge node for execution. Each drone deploys different types and numbers of service functions SF according to resource capacity to execute the subtasks of the DAG task; the two-dimensional task scheduling model is based on the drone position, subtask offloading position and offloading delay, and the number of drones. In each time slot of the task offloading process, the subtask is offloaded to the drone edge node to execute the subtask of the DAG task;
[0011] S102. State space setting: According to the deep reinforcement learning framework, the set of all subtasks and drones is regarded as an agent, and the state of the agent is defined;
[0012] S103. Action space setting: setting the action of the agent to select the drone number for the subtask to be unloaded in each time slot and the action of each drone in each time slot. The action of the drone includes speed and angle;
[0013] S104. Reward function setting: Based on the weighted sum of reducing the offloading delay of DAG tasks and the deployment cost of drone edge nodes, define the system overhead and design the reward function of the agent action;
[0014] S105. Training model design: Based on the set agent states, actions, and action rewards, select a deep reinforcement learning model for the multi-agent and configure the model structure;
[0015] S106. Scheduling model training: Use multi-agent reinforcement learning method to train the scheduling model, maximize the cumulative discount reward, and obtain a trained agent network model;
[0016] S107. Scheduling model deployment: Deploy the trained intelligent agent network model to the ground control station responsible for service function deployment and scheduling in the multi-UAV edge computing system for use in scheduling decisions by the ground control station.
[0017] Preferably, the coordinates of the drone's position are:
[0018]
[0019] Where t∈T={1,2,...,T} is the time slots into which the task offloading process is divided, T= is the number of time slots, τ is the length of each time slot, l max In order to limit the flight range of the drone to the side length of the square target area, is the initial coordinate of the drone, H is the flight altitude, is a fixed value, and v t is the flight speed of the UAV at time slot t, θ t is the flight angle, v t ∈[0,v max ],θ t ∈[0,θ max ],v max is the maximum flight speed, θ max is the maximum turning angle;
[0020] At time slot t, drone u r With drone u r′ The distance between them is:
[0021]
[0022] in, d min is the safe distance between two drones.
[0023] Preferably, the unloading delay of the mobile device to unload the subtask to the drone includes data transmission time, task execution time and waiting time;
[0024] The data transmission time includes the data transmission time of the entry subtask and the non-entry subtask:
[0025] For the entry subtask, in the first time slot, subtask v m,1 Transmitted from the mth mobile device to drone u r The data transfer time is:
[0026]
[0027] in, is the data volume of the entry subtask, For the mth mobile device to drone u r The transmission rate of the wireless link between
[0028] For non-entry subtasks, the data generated by the execution of all predecessor subtasks needs to be transmitted from the drone edge node where the predecessor subtask is unloaded to the drone edge node where the non-entry subtask is unloaded. The transmission time is:
[0029]
[0030] in, For drone u r to u r′ The transmission bandwidth between, k represents the kth subtask executed in sequence in time slot t, data i,j For the predecessor task v m,i The result data generated by the execution;
[0031] The execution time is:
[0032]
[0033] in, Respectively represent the subsequent subtask v m,j The amount of data, the number of CPU cycles required to calculate each bit of data, For the subsequent subtask v m,j Calculate the number of CPU cycles required, obtained by executing the program analyzer;
[0034] The waiting time is the time from when the input data of the subtask reaches the start time of execution:
[0035]
[0036] in, v m,i In drone u r The queuing time on depends on the corresponding SF deployed on, that is, the queue length of s(m, i), is the execution time of the kth subtask in the queue, β r v m,i The queue position in the queue;
[0037] Subtask v m,j The total unloading delay is the sum of the transmission time, task execution time and waiting time, expressed as:
[0038]
[0039] Preferably, the state of the agent includes the state of all subtasks in each time slot and the state of the drone;
[0040] The status of each subtask at each time slot, including the data size of the subtask, the number of CPU cycles required for calculation, the horizontal and vertical coordinates of the mobile device position where the subtask is located in the initial state, the horizontal and vertical coordinates of the mobile device position where the subtask is located after training starts, or the horizontal and vertical coordinates of the drone position where the latest completed subtask among all the predecessor subtasks of the subtask is unloaded, the position of the subtask in the DAG task, whether the subtask is unloaded at each time slot, and the number of the drone edge node to which the subtask is unloaded;
[0041] The state of the drone includes the set of states of each drone edge node, the remaining resources of the drone in each time slot, and the horizontal and vertical coordinates of the drone position.
[0042] Preferably, the total system overhead in time slot t is:
[0043]
[0044] in, is the total deployment cost of all UAVs in time slot t, L t is the minimum delay for all subtasks to complete in time slot t, ω 0 and ω 1 Are the corresponding two weight parameters;
[0045] The reward function consists of three parts, specifically:
[0046] The total deployment cost of all drones at time slot t:
[0047]
[0048] Among them, Q r is the initial deployment cost of each drone, is the number of times each UAV is deployed before time slot t, and s is the number of UAVs deployed in time slot t;
[0049] The minimum delay for all subtasks to complete in time slot t, that is, the total unloading delay:
[0050]
[0051] in, is the completion delay of N subtasks in time slot t;
[0052] If a subtask that needs to be unloaded in time slot t cannot find a suitable drone to unload, a penalty will be given; if more than n subtasks cannot find a suitable drone to unload for consecutive times, a larger penalty will be given; the penalty is the reward value in time slot t minus a fixed value.
[0053] Preferably, the deep reinforcement learning model includes but is not limited to a deep Q network, a deep double Q network, and a proximal strategy optimization algorithm model.
[0054] Preferably, the discount reward is specifically to select a discount factor γ, γ∈(0,1). When γ=0, the agent pays more attention to the immediate reward. When γ is close to 1, the agent pays more attention to the long-term reward and considers the contribution of each action to achieving the ultimate goal.
[0055] Preferably, the specific content of the scheduling model training is:
[0056] S1061. Variable assignment: Initialize the variables of deep reinforcement learning, including initializing network parameters, initializing training step size, and setting learning rate, number of network updates and discount factor;
[0057] S1062. Action selection: In the tth time slot, after interacting with the environment, the agent state is input into the policy network, the mean and variance of the action probability density function are output, the normal distribution of the action probability density function is obtained, an action is obtained from the normal distribution, and the log value of the probability density corresponding to the action is calculated;
[0058] S1063. Environmental interaction: After completing the action selection for the t-th time slot, obtain the drone number selected by the subtask, subtract the resources required for the subtask calculation from the original resources to obtain the drone's computing resources and update the drone resources in the agent state, use the drone's flight speed and angle to update the coordinates of the drone in the t-th time slot in the agent state, select the subtask that meets the unloading conditions in the next time slot and update the agent state;
[0059] S1064. Experience accumulation: In the process of interaction between the agent and the environment, the environment state, action space, rewards obtained from the environment, and new states generated together constitute the agent state, which is stored in the experience pool as training samples. When the number of training samples in the experience pool is an integer multiple of the number of training steps, a new sample is used to replace an old training sample. In subsequent training, samples are continuously randomly selected from the experience pool and input into the neural network for training to break the correlation between data;
[0060] S1065. Loss calculation: After all subtasks are unloaded, the training is terminated, and the last state obtained is input into the value network. After obtaining the corresponding agent state value, the discounted reward for each step is calculated; all states are input into the value network to obtain the agent state, the difference between the two agent state values is calculated and the square difference is calculated, and the value network is updated through gradient back propagation; the stored state is input into the new policy network to obtain the ratio of the behavior probabilities of the new and old policy networks, respectively, and the loss function is calculated, back propagation is performed, and the policy network is updated. This step is repeated, and then the old policy network is updated with the weights of the new policy network;
[0061] S1066. Strategy export: After a period of training, the intelligent agent network model is obtained. Given a certain input state, the intelligent agent network model outputs the optimal action with the maximum expected reward.
[0062] Preferably, the new and old strategy networks are in state s t The ratio of the probability of taking the action is:
[0063]
[0064] Control the strategy update amplitude by calculating the ratio of probabilities;
[0065] Among them, s t = {v t ,u t},s t is the environmental state, v t is the status of all subtasks, u t is the state of the drone, a t is the action space, a t ={a_d t , a_c t}, a_d t is the discrete action value, a_c t is the continuous action value;
[0066] The function for scheduling model training is as follows:
[0067]
[0068] Among them, L CLIP (θ) is the training loss function, is the estimate of the advantage function, δ t =r t +γv(s t+1 )-v(s t ), e t =(s t , a t , r t ,st+1 ) is the state of the agent, and the state of the agent is used as a training sample e = (e 1 , e 2 , ..., e T ) is stored in the experience pool, r t is the reward obtained from the environment, s t+1 is the new state generated, γ is the discount factor, ∈ is a small positive number, and the clip function is used to limit the probability ratio r t The range of variation of (θ) prevents the update step from being too large.
[0069] Preferably, the specific content of the ground control station scheduling decision is:
[0070] According to the actual deployment scenario, the agent state is abstracted as the input of the agent network model, and the output is transmitted to the mobile device and the drone edge node respectively. The drone edge node to which the subtask is offloaded and the flight angle and speed of the drone are determined based on the model output.
[0071] Through the above technical solutions, it can be known that compared with the prior art, the present invention discloses a service function deployment and UAV edge node resource allocation method based on deep reinforcement learning, which solves the problem of balancing the UAV deployment cost and task unloading delay under multi-UAV edge computing, comprehensively considers the complexity of ground tasks, the correlation between task scheduling and service deployment, and the UAV's own flight status, and uses a deep reinforcement learning method to combine the number of UAVs, task scheduling, and service function deployment. Under the constraints of limited computing resources of UAVs, the UAV deployment cost and task unloading delay are balanced during the movement of UAVs, thereby improving the stability of the system;
[0072] The present invention can minimize the number of UAV edge nodes required to be deployed and reduce deployment costs while meeting the DAG execution performance constraints. It uses a deep reinforcement learning method to jointly optimize the number and posture of UAVs, service function deployment and DAG task offloading, which can adapt to different deployment scenarios and has a wide range of uses. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0074] Figure 1 A flow chart of a method for service function deployment and UAV edge node resource allocation based on deep reinforcement learning provided by the present invention;
[0075] Figure 2 A schematic diagram of a DAG task provided by the present invention;
[0076] Figure 3 A schematic diagram of a multi-UAV edge computing scenario provided by the present invention;
[0077] Figure 4 A schematic diagram of the state space of a certain time slot provided by the present invention;
[0078] Figure 5 A schematic diagram of the action space of a certain time slot provided by the present invention;
[0079] Figure 6 A schematic diagram of a multi-agent deep reinforcement learning training model provided by the present invention;
[0080] Figure 7 A schematic diagram of the Actor network in deep reinforcement learning provided by the present invention;
[0081] Figure 8 This is a flow chart of the scheduling model training method provided by the present invention. DETAILED DESCRIPTION
[0082] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0083] The embodiment of the present invention discloses a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning, such as Figure 1 , including the following steps:
[0084] S101. Deployment problem modeling: Based on the deployment scenarios of ground tasks and drones, a two-dimensional task scheduling model is constructed; the deployment scenario is that multiple mobile devices are distributed in the area, each mobile device has a computing-intensive task DAG, and the DAG task is offloaded to the drone edge node for execution. Each drone deploys different types and numbers of service functions SF according to resource capacity to execute the subtasks of the DAG task; the two-dimensional task scheduling model is based on the drone position, subtask offloading position and offloading delay, and the number of drones. In each time slot of the task offloading process, the subtask is offloaded to the drone edge node to execute the subtask of the DAG task;
[0085] S102. State space setting: According to the deep reinforcement learning framework, the set of all subtasks and drones is regarded as an agent, and the state of the agent is defined;
[0086] S103. Action space setting: setting the action of the agent to select the drone number for the subtask to be unloaded in each time slot and the action of each drone in each time slot. The action of the drone includes speed and angle;
[0087] S104. Reward function setting: Based on the weighted sum of reducing the offloading delay of DAG tasks and the deployment cost of drone edge nodes, define the system overhead and design the reward function of the agent action;
[0088] S105. Training model design: Based on the set agent states, actions, and action rewards, select a deep reinforcement learning model for the multi-agent and configure the model structure;
[0089] S106. Scheduling model training: Use multi-agent reinforcement learning method to train the scheduling model, maximize the cumulative discount reward, and obtain a trained agent network model;
[0090] S107. Scheduling model deployment: Deploy the trained intelligent agent network model to the ground control station responsible for service function deployment and scheduling in the multi-UAV edge computing system for use in scheduling decisions by the ground control station.
[0091] In this embodiment, the complex tasks generated by the ground equipment are represented as a directed acyclic graph (DAG), each of which contains multiple interdependent subtasks, and there is an execution order relationship between the subtasks, for example, Figure 2 , a schematic diagram of the DAG structure of the video playback task is given: the task contains 11 subtasks, and the edges between two subtasks represent data dependencies; for a subtask, its operation usually needs to rely on specific data and software environment, collectively referred to as service function (SF), and this subtask can only be offloaded to the drone edge node where the required SF is deployed for execution; in addition, in the multi-drone edge computing scenario, due to the limited computing resources of a single node and the SF operation constraints, multiple subtasks contained in a DAG task may be offloaded to different drone edge nodes.
[0092] The service function SF is that the drone loads the software and data required for the service function from the storage device it carries, and runs an instance of the service function in the form of a container; the mobile device is a removable device with limited computing power, such as a communication radio, an IoT sensor, etc.; the drone edge node includes but is not limited to a power supply module, a power module, a control module, a communication module, a computing module, and a positioning module, which are used to receive the DAG tasks generated by the mobile device and return the running results; the deployment scenario is a specific scenario for the deployment of multi-drone edge computing, including drone edge nodes, mobile devices, DAG tasks, etc.
[0093] The deployment scenarios are as follows:
[0094] like Figure 3 , there are R drone edge nodes providing computing services for M mobile devices, each of which has a computing-intensive task to perform: the drone and the ground device are represented by the set U = {u 1 ,u 2 ...,u r ...,u R} and V={V 1 , V 2 ..., V m ..., V M} means, u r represents the rth drone, V m represents the mth DAG task; each drone deploys different types and numbers of SFs according to its resource capacity to facilitate the execution of the subtasks of the DAG task; the subtask of the mobile device selects the drone edge node to which the subtask needs to be offloaded according to the type of SF it needs to offload;
[0095] The task offloading process is divided into multiple time slots, represented as t∈T={1,2,...,T}. It is assumed that the position of the mobile device remains unchanged in each time slot. In the process of subtask offloading, in order to reduce the deployment cost of drones, it is necessary to reduce the number of drone edge nodes required to be deployed. The DAG task that the mth mobile device needs to offload is represented as G m =(V m ,ε), is a set of subtasks, α m is the number of subtasks of the mth DAG task, ε is the amount of dependent data between subtasks, and the position of the mth mobile device is expressed as The set of all SFs is F = {f 1 , f 2 , ..., f Q}, where Q is the total number of SF categories.
[0096] In this embodiment, scheduling refers to offloading the subtasks of the DAG task to the edge nodes of the drone. Specifically:
[0097] Define the decision variables for unloading subtasks in time slot t and service function deployment variables They represent the offloading of subtasks and the deployment of service functions, respectively. Both are 0-1 decision variables. m,i Represents subtask v m,i Uninstall the required SF type, Indicates that the subtask v m,i Unload to UAV u at time slot t r superior, Indicates that the subtask v is not m,i Unload to the drone during the time slot, Representatives m,i Deployed on drones r superior, Representatives m,i Not deployed on drones r Given that the execution of subtasks depends on the deployment of SF, the subtask offloading decision variables Deploy decision variables with service functions There is also a dependency relationship between them. When the subtask v m,i Unload to drone u in time slot r When the drone is on, the corresponding service function must be deployed on it. m,i ,but
[0098] In order to further implement the above technical solution, each drone edge node has an initial position at the beginning. In each time slot, the position of the drone edge node can change. The mobile model determines the position of the drone in the next time slot according to the position, movement angle and speed of the drone in the current time slot:
[0099] At time slot t, drone u r The coordinates of the location are:
[0100]
[0101] Where t∈T={1,2,...,T} is the time slots into which the task offloading process is divided, T= is the number of time slots, τ is the length of each time slot, l max In order to limit the flight range of the drone to the side length of the square target area, is the initial coordinate of the drone, H is the flight altitude, is a fixed value, and v t is the flight speed of the UAV at time slot t, θ t is the flight angle, vt ∈[0,V max ],θ t ∈[0,θ max ],v max is the maximum flight speed, θ max is the maximum turning angle;
[0102] At time slot t, drone u r With drone u r′ The distance between them is:
[0103]
[0104] in, d min is the safe distance between two drones.
[0105] In order to offload the subtask to the drone, the mobile device needs to transmit the data corresponding to the subtask to the drone edge node through a wireless communication link. The time required for transmission is called the offloading delay.
[0106] To further implement the above technical solution, the offloading delay of the mobile device to offload the subtask to the drone includes data transmission time, task execution time and waiting time;
[0107] The data transmission time includes the data transmission time of the entry subtask and the non-entry subtask:
[0108] For the entry subtask, in the first time slot, subtask v m,1 Transmitted from the mth mobile device to drone u r The data transfer time is:
[0109]
[0110] in, is the data volume of the entry subtask, For the mth mobile device to drone u r The transmission rate of the wireless link between
[0111] For each subsequent time slot, a subtask that meets the unloading condition is selected for execution. Meeting the unloading condition means that the subtask is an entry subtask or all the predecessor subtasks of the subtask have been executed;
[0112] For non-entry subtasks, all data generated by the execution of the predecessor subtasks need to be transmitted from the drone edge node where the predecessor subtasks are unloaded to the drone edge node where the non-entry subtasks are unloaded. Specifically:
[0113] Current driver task v m,i Unload to drone rExecute on, its successor subtask v m,j Unload to drone r′ When executing on m,i Execution result data i,j Need to get from u r Transfer to u r′ The transmission time is:
[0114]
[0115] in, For drone u r to u r′ The transmission bandwidth between, k represents the kth subtask executed in sequence in time slot t, data i,j For the predecessor task v m,i The result data generated by the execution;
[0116] When the subtask v m,j When the subtask v is unloaded to the drone and the execution result data of the predecessor subtask required by it has been transmitted, the subtask v m,j Can be used in drones r′ Start execution on
[0117] The execution time is:
[0118]
[0119] in, Respectively represent the subsequent subtask v m,j The amount of data, the number of CPU cycles required to calculate each bit of data, For the subsequent subtask v m,j Calculate the number of CPU cycles required, obtained by executing the program analyzer;
[0120] Waiting time refers to the time from the input data of the subtask to the start of execution. When multiple subtasks are unloaded to the same SF on the same drone, queuing time will be generated. Each subtask must wait for the previous subtask to be completed before it can start execution. The execution time is:
[0121]
[0122] in, v m,i In drone u r The queuing time on depends on the corresponding SF deployed on, that is, the queue length of s(m, i), is the execution time of the kth subtask in the queue, β r v m,i The queue position in the queue;
[0123] Subtask v m,j The total unloading delay is the sum of the transmission time, task execution time and waiting time, expressed as:
[0124]
[0125] In order to further implement the above technical solution, Figure 4 , the state of the agent Includes the status of all subtasks in each time slot and the status of the drone The status of a subtask is the collection of all subtask statuses
[0126] The status of each subtask in each time slot Include subtask v m,i The amount of data Calculate the number of CPU cycles required The horizontal coordinate of the mobile device where the subtask is located in the initial state and the vertical coordinate The horizontal and vertical coordinates of the mobile device location where the subtask is located after the training starts, or the horizontal and vertical coordinates of the drone location where the latest completed subtask among all the predecessor subtasks of the subtask is unloaded (after the training starts, The value of v changes. m,i is the entry subtask, then Represents subtask v m,i The horizontal and vertical coordinates of the mobile device's location, otherwise, Represents subtask v m,i The horizontal and vertical coordinates of the drone location where the latest completed subtask among all predecessor subtasks is unloaded), the position of the subtask in the DAG task (If it is an entry subtask, Intermediate subtasks Export subtask ), in each time slot t, subtask v m,i Uninstall (If unloading occurs at time slot t, Otherwise 0), and subtask v m,i The ID of the drone edge node to which the device is unloaded
[0127] The subtask and the location of the drone refer to the horizontal and vertical coordinates of the task and the drone in the two-dimensional map model, where the height of the drone is set to a fixed value; whether the subtask is unloaded in the current time slot refers to setting the corresponding value to 1 when unloading in the time slot, and setting it to 0 when not unloading in the time slot; the remaining resources of the drone are the total resources of the drone edge node minus the total resources occupied by the subtasks executed on the drone, and the resources considered mainly refer to computing resources;
[0128] The state of the drone includes the set of states of each drone edge node Each element in the collection represents the state of a drone edge node, for example, the element Indicates drone u r The state at time slot t, is the remaining resources of the drone in each time slot, For time slot t drone u r The horizontal and vertical coordinates of the position.
[0129] In this embodiment, if Figure 5 , the action space is set as follows: the action of each time slot is set to According to the scenario constraints, each time slot can offload at most N subtasks, where is the unloading vector corresponding to the subtask in time slot t, Represents the subtask v i Uninstall to u r On, among them represents the offloading result of the Nth subtask. If, in time slot t, only less than N subtasks meet the offloading conditions, the action value corresponding to the subtask that does not need to be offloaded is set to 0; and They represent the flight speed and angle of R UAVs in time slot t respectively.
[0130] In order to further implement the above technical solution, the total system overhead in time slot t is:
[0131]
[0132] in, is the total deployment cost of all UAVs in time slot t, L t is the minimum delay for all subtasks to complete in time slot t, ω 0 and ω 1 Are the corresponding two weight parameters;
[0133] The reward function consists of three parts, specifically:
[0134] (1) The number of drones deployed in time slot t is s. If a drone has been deployed in the previous time slot, the cost of the drone will decrease as the number of deployments increases. That is, the deployment cost of a drone is inversely proportional to the number of times the drone is deployed. The total deployment cost of all drones in time slot t is:
[0135]
[0136] Among them, Q r is the initial deployment cost of each drone, is the number of times each UAV is deployed before time slot t, and s is the number of UAVs deployed in time slot t;
[0137] (2) The total unloading delay of all subtasks unloaded in the current time slot is defined as the unloading delay of the subtask that is completed latest among all subtasks executed in the time slot. Then the completion delay of N subtasks in the time slot can be expressed as:
[0138]
[0139] The minimum delay for all subtasks to complete in time slot t, that is, the total unloading delay:
[0140]
[0141] in, is the completion delay of N subtasks in time slot t;
[0142] (3) If a subtask that needs to be unloaded in time slot t cannot find a suitable drone for unloading, a penalty will be imposed; if a subtask cannot find a suitable drone for unloading for more than n consecutive times, a larger penalty will be imposed.
[0143] In order to further implement the above technical solutions, deep reinforcement learning models include but are not limited to deep Q networks, deep double Q networks and proximal strategy optimization algorithm models.
[0144] In order to further implement the above technical solution, the specific discount reward is to select a discount factor γ, γ∈(0,1). When γ=0, the agent pays more attention to the immediate reward. When γ is close to 1, the agent pays more attention to the long-term reward and considers the contribution of each action to achieving the ultimate goal.
[0145] In this example, the network model is based on a proximal policy network, such as Figure 6 As shown, it contains two Actor networks and one Critic network. The network parameters and structures of the two Actor networks are the same, as shown in Figure 7As shown in the figure, it contains 1 input layer, 2 output layers and 3 Dense fully connected layers. The dimensions of the input layer and the output layer correspond to the dimensions of the state space and the action space respectively. The number of neurons in the fully connected layer is 100, 50 and the dimension of the action space respectively. The activation function is Tanh and Softplus function. The input is the initial state of the subtask and the drone. The continuous action value is output. The most recent state is input into the Critic network. The value corresponding to the state is subtracted from the expected cumulative reward and optimized by the Adam method.
[0146] In order to further implement the above technical solution, Figure 8 ,The specific content of scheduling model training is:
[0147] S1061. Variable assignment: Initialize the variables of deep reinforcement learning, including initializing network parameters, initializing training step size, and setting learning rate, number of network updates and discount factor;
[0148] Specifically, initialize the parameters of the three networks: Actor_old network, Actor_new network and Critic network, set the parameters of Actor_old network to θ, and the parameters of Actor_new network to θ'=θ; initialize the training step size D; set the learning rates of Actor network and Critic network to lr_a and lr_c respectively, the number of network updates s_a and s_l, and the discount factor γ.
[0149] For example, you can set D = 32, learning rate lr_a = 0.0001, lr_c = 0.0002, and discount factor γ = 0.9;
[0150] S1062. Action selection: In the tth time slot, after interacting with the environment, the agent state is input into the policy network, the mean and variance of the action probability density function are output, the normal distribution of the action probability density function is obtained, an action is obtained from the normal distribution, and the log value of the probability density corresponding to the action is calculated;
[0151] Specifically, in the tth time slot, after interacting with the environment, the agent state s t Input into the Actor_old network, output the mean μ of the action probability density function through the Tanh activation function, output the variance σ of the action probability density function through the Softplus activation function, output the normal distribution of the action probability density function dist = Normal(μ,σ), and get an action a from the normal distribution t , calculate the log value of the probability density corresponding to the action;
[0152] For example, l = [-2, 2], where the speed and angle ranges of the drone are [0, 30] and [0, 2π] respectively. The speed and angle ranges are mapped in l to obtain the action value. When the maximum number of drones is 10, the drones selected by the subtask are numbered 0-9. After mapping the range in l and discretizing the obtained values, the specific number values are obtained.
[0153] For example, there are 10 drone numbers, and l = [-2, 2] is divided into 10 continuous intervals. When the action value is in the interval [-2, -1.6], drone No. 0 is selected, and when the action value is in the interval [1.6, 2], drone No. 9 is selected.
[0154] S1063. Environmental interaction: After completing the action selection for the t-th time slot, obtain the drone number selected by the subtask, subtract the resources required for the subtask calculation from the original resources to obtain the drone's computing resources and update the drone resources in the agent state, use the drone's flight speed and angle to update the coordinates of the drone in the t-th time slot in the agent state, select the subtask that meets the unloading conditions in the next time slot and update the agent state;
[0155] S1064. Experience accumulation: In the process of interaction between the agent and the environment, the state of the environment s t , action space a t , the reward r obtained from the environment t And the new state s t+1 Together they form the state of the agent e t =(s t , a t , r t ,s t+1 ), as training sample e=(e 1 , e 2 ..., e T ) is stored in the experience pool. When the number of training samples in the experience pool is an integer multiple of the number of training steps B, a new sample is used to replace an old training sample. In subsequent training, samples are continuously randomly selected from the experience pool and input into the neural network for training to break the correlation between the data.
[0156] S1065. Loss calculation: After all subtasks are unloaded, the training is terminated, and the last state obtained is input into the value network. After obtaining the corresponding agent state value, the discounted reward for each step is calculated; all states are input into the value network to obtain the agent state, the difference between the two agent state values is calculated and the square difference is calculated, and the value network is updated through gradient back propagation; the stored state is input into the new policy network to obtain the ratio of the behavior probabilities of the new and old policy networks, respectively, and the loss function is calculated, back propagation is performed, and the policy network is updated. This step is repeated, and then the old policy network is updated with the weights of the new policy network;
[0157] Specifically, after all subtasks are unloaded, the training is terminated, and the last state s_ is input into the Critic network. After obtaining the corresponding v value, the discounted reward for each step is calculated; all states s are input into the Critic network to obtain the agent state v _ , make a difference between the two values, square the difference, and update the Critic network through gradient backpropagation; on the other hand, input the stored s into the Actor_new network to obtain the corresponding ratio r t , calculate the loss function L, then perform back propagation to update the Actor_new network; repeat this step, and then update the Actor_old network with the Actor_new weights;
[0158] Among them, the ratio r t For Actor_new and Actor_old networks, they are in state s t Take the ratio of behavior probabilities and calculate the ratio of probabilities to control the strategy update amplitude;
[0159] S1066. Strategy export: After a period of training, the intelligent agent network model is obtained. Given a certain input state, the intelligent agent network model outputs the optimal action with the maximum expected reward.
[0160] In order to further implement the above technical solution, the new and old strategy networks are respectively in state s t The ratio of the probability of taking the action is:
[0161]
[0162] Control the strategy update amplitude by calculating the ratio of probabilities;
[0163] Among them, s t = {v t ,u t},s t is the environmental state, v t is the status of all subtasks, u t is the state of the drone, at is the action space, a t ={a_d t , a_c t}, a_d t is the discrete action value, a_c t is the continuous action value;
[0164] The function for scheduling model training is as follows:
[0165]
[0166] Among them, L CLIP (θ) is the training loss function, is the estimate of the advantage function, δ t =r t +γv(s t+1 )-v(s t ), e t =(s t , a t , r t ,s t+1 ) is the state of the agent, and the state of the agent is used as a training sample e = (e 1 , e 2 ..., e T ) is stored in the experience pool, r t is the reward obtained from the environment, s t+1 is the new state generated, γ is the discount factor, ∈ is a small positive number, and the clip function is used to limit the probability ratio r t The range of variation of (θ) prevents the update step from being too large.
[0167] In order to further implement the above technical solution, the specific content of the ground control station scheduling decision is as follows:
[0168] According to the actual deployment scenario, the agent state is abstracted as the input of the agent network model, and the output is transmitted to the mobile device and the drone edge node respectively. The drone edge node to which the subtask is offloaded and the flight angle and speed of the drone are determined based on the model output.
[0169] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, a service function deployment and drone edge node resource allocation method based on deep reinforcement learning is implemented.
[0170] A processing terminal includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and is characterized in that when the processor executes the computer program, a service function deployment based on deep reinforcement learning and a method for allocating resources of unmanned aerial vehicle edge nodes are implemented.
[0171] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0172] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for service function deployment and drone edge node resource allocation based on deep reinforcement learning, characterized in that: The following steps are involved: S101. Deployment problem modeling: Construct a two-dimensional task scheduling model based on ground tasks and UAV deployment scenarios; The deployment scenario is that multiple mobile devices are distributed in the region. Each mobile device has a computationally intensive task DAG. The DAG task is offloaded to the drone edge node for execution. Each drone deploys different types and numbers of service functions SF according to resource capacity to execute the subtasks of the DAG task; The two-dimensional task scheduling model is based on the UAV position, subtask unloading position and unloading delay, and the number of UAVs. In each time slot of the task unloading process, the subtask is unloaded to the edge node of the UAV to execute the subtask of the DAG task; S102. State space setting: According to the deep reinforcement learning framework, the set of all subtasks and drones is regarded as an agent, and the state of the agent is defined; S103. Action space setting: setting the action of the agent to select the drone number for the subtask to be unloaded in each time slot and the action of each drone in each time slot. The action of the drone includes speed and angle; S104. Reward function setting: Based on the weighted sum of reducing the offloading delay of DAG tasks and the deployment cost of drone edge nodes, define the system overhead and design the reward function of the agent action; S105. Training model design: Based on the set agent states, actions, and action rewards, select a deep reinforcement learning model for the multi-agent and configure the model structure; S106. Scheduling model training: Use multi-agent reinforcement learning method to train the scheduling model, maximize the cumulative discount reward, and obtain a trained agent network model; S107. Scheduling model deployment: Deploy the trained agent network model to the ground control station responsible for service function deployment and scheduling in the multi-UAV edge computing system for use in scheduling decisions by the ground control station; In step S104, the total system overhead in time slot t is: in, is the total deployment cost of all drones in time slot t, R is the total number of drone edge nodes, and L t is the minimum delay for all subtasks to be completed in time slot t, ω0 and ω1 are the corresponding two weight parameters; The reward function consists of three parts, specifically: The total deployment cost of all drones at time slot t: Among them, Q r is the initial deployment cost of each drone, is the number of times each UAV is deployed before time slot t, and s is the number of UAVs deployed in time slot t; The minimum delay for all subtasks to complete in time slot t, that is, the total unloading delay: in, is the completion delay of N subtasks in time slot t; If a subtask that needs to be unloaded in time slot t cannot find a suitable drone to unload, a penalty will be given; if more than n subtasks cannot find a suitable drone to unload for consecutive times, a larger penalty will be given; the penalty is the reward value in time slot t minus a fixed value.
2. According to claim 1, a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning is characterized in that: The coordinates of the drone's position are: Where t∈T={1,2,...,T} is the time slots into which the task offloading process is divided, T= is the number of time slots, τ is the length of each time slot, l max In order to limit the flight range of the drone to the side length of the square target area, is the initial coordinate of the drone, H is the flight altitude, is a fixed value, and v t is the flight speed of the UAV at time slot t, θ t is the flight angle, v t ∈[0,v max ],θ t ∈[θ,θ max ],v max is the maximum flight speed, θ max is the maximum turning angle; At time slot t, drone u r With drone u r′ The distance between them is: in, d min is the safe distance between two drones.
3. According to the method of service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: The offloading delay of the mobile device to offload the subtask to the drone includes data transmission time, task execution time and waiting time; The data transmission time includes the data transmission time of the entry subtask and the non-entry subtask: For the entry subtask, in the first time slot, subtask v m,1 Transmitted from the mth mobile device to drone u r The data transfer time is: in, is the data volume of the entry subtask, For the mth mobile device to drone u r The transmission rate of the wireless link between For non-entry subtasks, the data generated by the execution of all predecessor subtasks needs to be transmitted from the drone edge node where the predecessor subtask is unloaded to the drone edge node where the non-entry subtask is unloaded. The transmission time is: in, For drone u r to u r′ The transmission bandwidth between, k represents the kth subtask executed in sequence in time slot t, data i,j For the predecessor task v m,i The result data generated by the execution; The task execution time is: in, Respectively represent the subsequent subtask v m,j The amount of data, the number of CPU cycles required to calculate each bit of data, For the subsequent subtask v m,j Calculate the number of CPU cycles required, obtained by executing the program analyzer; The waiting time is the time from when the input data of the subtask reaches the start time of execution: in, v m,i In drone u r The queuing time on depends on u r Deployed on v m,i The corresponding SF, that is, the queue length of s(m, i), is the execution time of the kth subtask in the queue, β r v m,i The queue position in the queue; Subtask v m,j The total unloading delay is the sum of the transmission time, task execution time and waiting time, expressed as:
4. According to the method of service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: The state of the agent includes the state of all subtasks and the state of the drone in each time slot; The status of each subtask at each time slot, including the data size of the subtask, the number of CPU cycles required for calculation, the horizontal and vertical coordinates of the mobile device position where the subtask is located in the initial state, the horizontal and vertical coordinates of the mobile device position where the subtask is located after training starts, or the horizontal and vertical coordinates of the drone position where the latest completed subtask among all the predecessor subtasks of the subtask is unloaded, the position of the subtask in the DAG task, whether the subtask is unloaded at each time slot, and the number of the drone edge node to which the subtask is unloaded; The state of the drone includes the set of states of each drone edge node, the remaining resources of the drone in each time slot, and the horizontal and vertical coordinates of the drone position.
5. According to a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: Deep reinforcement learning models include but are not limited to deep Q networks, deep double Q networks, and proximal policy optimization algorithm models.
6. According to a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: The discount reward is specifically to select a discount factor γ, γ∈(0,1). When γ=0, the agent pays more attention to the immediate reward. When γ is close to 1, the agent pays more attention to the long-term reward and considers the contribution of each action to achieving the final goal.
7. According to the method of service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: The specific content of scheduling model training is: S1061. Variable assignment: Initialize the variables of deep reinforcement learning, including initializing network parameters, initializing training step size, and setting learning rate, number of network updates and discount factor; S1062. Action selection: In the tth time slot, after interacting with the environment, the agent state is input into the policy network, the mean and variance of the action probability density function are output, the normal distribution of the action probability density function is obtained, an action is obtained from the normal distribution, and the log value of the probability density corresponding to the action is calculated; S1063. Environmental interaction: After completing the action selection for the t-th time slot, obtain the drone number selected by the subtask, subtract the resources required for the subtask calculation from the original resources to obtain the drone's computing resources and update the drone resources in the agent state, use the drone's flight speed and angle to update the coordinates of the drone in the t-th time slot in the agent state, select the subtask that meets the unloading conditions in the next time slot and update the agent state; S1064. Experience accumulation: In the process of interaction between the agent and the environment, the environment state, action space, rewards obtained from the environment, and new states generated together constitute the agent state, which is stored in the experience pool as training samples. When the number of training samples in the experience pool is an integer multiple of the number of training steps, a new sample is used to replace an old training sample. In subsequent training, samples are continuously randomly selected from the experience pool and input into the neural network for training to break the correlation between data; S1065. Loss calculation: After all subtasks are unloaded, the training is terminated, and the last state obtained is input into the value network. After obtaining the corresponding agent state value, the discounted reward for each step is calculated; all states are input into the value network to obtain the agent state, the difference between the two agent state values is calculated and the square difference is calculated, and the value network is updated through gradient back propagation; the stored state is input into the new policy network to obtain the ratio of the behavior probabilities of the new and old policy networks, respectively, and the loss function is calculated, back propagation is performed, and the new policy network is updated. Repeat this step, and then the old policy network is updated with the weights of the new policy network; S1066. Strategy export: After a period of training, the intelligent agent network model is obtained. Given a certain input state, the intelligent agent network model outputs the optimal action with the maximum expected reward.
8. According to claim 7, a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning is characterized in that: The new and old policy networks are respectively in state s t The ratio of the probability of taking the action is: Control the strategy update amplitude by calculating the ratio of probabilities; Among them, s t = {v t ,u t },s t is the environmental state, v t is the status of all subtasks, u t is the state of the drone, a t is the action space, a t ={a_d t , a_c t }, a_d t is the discrete action value, a_c t is the continuous action value; The function for scheduling model training is as follows: Among them, L CLIP (θ) is the training loss function, is the estimate of the advantage function, δ t =r t +γv(s t+1 )-v(s t ), e t =(s t , a t , r t ,s t+1 ) is the state of the agent, and the state of the agent is used as the training sample e = (e1, e2, ..., e T ) is stored in the experience pool, r t is the reward obtained from the environment, s t+1 is the new state generated, γ is the discount factor, ∈ is a small positive number, and the clip function is used to limit the probability ratio r t The range of variation of (θ) prevents the update step from being too large.
9. According to a method for service function deployment and drone edge node resource allocation based on deep reinforcement learning in claim 1, it is characterized in that: The specific contents of the ground control station dispatch decision are: According to the actual deployment scenario, the agent state is abstracted as the input of the agent network model, and the output is transmitted to the mobile device and the drone edge node respectively. The drone edge node to which the subtask is offloaded and the flight angle and speed of the drone are determined based on the model output.
Citation Information
Patent Citations
Deployment and scheduling methods and systems for multi-UAV edge computing services
CN113391647B
Multi-unmanned aerial vehicle edge calculation path optimization and dependent task scheduling optimization method and system
CN116451934A
Flight control and calculation unloading method and system for multi-unmanned aerial vehicle mobile edge calculation
CN115454527A