A method for UAV-assisted multi-node task offloading scheduling

The Markov model is constructed through a model-free reinforcement learning method, combining small learning goals and pre-reward mechanisms, and improving the ε-greedy strategy, solving the problems of user node differences and privacy protection in edge computing scenarios by drones, and maximizing benefits within limited service time.

CN113867934BActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110918758.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-11
Publication Date
2025-07-04
Estimated Expiration
2041-08-11

AI Technical Summary

Technical Problem

In the drone-assisted edge computing scenario, how to obtain the best benefits by choosing flight paths and offloading strategies within a limited service time, while solving user node differences and privacy protection issues.

Method used

Using a model-free reinforcement learning method, a Markov model is constructed, combining small learning goals, pre-rewards and large reward sensitive mechanisms, the ε-greedy strategy is improved, and the Q table is updated through Q-Learning to realize the task offloading and scheduling of the drone within limited service time.

Benefits of technology

It maximizes the benefits of drones within limited service time, meets privacy protection needs, and does not require too much prior knowledge, and has good reusability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113867934B_ABST
    Figure CN113867934B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-node task offloading scheduling method assisted by an unmanned aerial vehicle (UAV). Based on the traditional model-free reinforcement learning method with value function update, the present invention optimizes the assistance scheduling problem in the scenario of UAV-assisted edge computing. On this basis, methods such as small learning objectives, pre-rewards, and large-reward sensitivity are innovatively proposed. Finally, under the constraints such as the time-delay sensitivity of UAV user nodes, the problem of maximizing the benefits of the UAV by selecting a flight path through strategies within a limited service time is achieved. The method of the present invention does not require too much prior knowledge and does not need to deeply understand the in-depth information of each user node, meeting the requirements of privacy protection. Moreover, the present invention has good reusability in similar application scenarios, and the practical value of the invention is relatively strong.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention belongs to the field of edge computing, and particularly relates to a reinforcement learning method for multi-node task offloading scheduling within a drone-assisted tour path. Background Art:

[0002] In some edge computing scenarios where it is inconvenient to directly deploy servers and provide services, drones can play an important coordination role due to their flexibility and convenience. Thus, the application of drone-assisted mobile edge computing task offloading scheduling has emerged. How to obtain the maximum benefit by selecting a flight path and an offloading strategy within the limited service time of the drone has become a new challenge. Among them, the problems of differences between user nodes and privacy protection are both difficult problems that are currently difficult to solve. The existing solutions include dynamic programming methods, convex optimization methods, Lyapunov stability methods, ant colony algorithms, particle swarm algorithms, etc. These methods may have good performance in some specific scenarios, but there is still much room for improvement in terms of the complexity of algorithm design, scalability, and data privacy protection.

[0003] With the development of AI, various reinforcement learning algorithms have been proven to have significant advantages in solving sequential decision-making problems. They are very suitable for dealing with the policy selection problem in the complex search space of edge computing scenarios. They only require less prior knowledge to bring relatively good solutions to problems, and at the same time meet the requirements of privacy protection. Reinforcement learning can be roughly divided into two categories: model-based reinforcement learning and model-free reinforcement learning. Due to the increasing emphasis on data security, it is difficult to obtain prior knowledge related to detailed data of multiple user nodes. Therefore, model-free reinforcement learning is more suitable for solving the task offloading scheduling problem under edge computing. Model-based reinforcement learning can also be further divided into two major categories. One is the policy optimization method, which does not need to maintain a value function model, but directly searches for the optimal policy. It often uses a parameterized policy and maximizes the expected return by updating this parameter. The other is the reinforcement learning method based on value function update, generally referring to the Q-Learning algorithm. Q is a historical experience memory table related to the current state and action selection, which can represent the cumulative expectation of the benefits obtained by taking actions in a certain state at a certain moment. The Q-Learning algorithm constructs an agent representing the algorithm, places it in the Markov model of the problem to be solved, and selects whether to make a new action selection by querying the accumulated learning experience or randomly select an action through a search strategy. The agent will record the learning results of each time and affect the next selection by updating the learning experience. As the number of training times increases, the action selections made by the agent based on the learning experience will become more and more accurate until it approximates the optimal solution to the problem. Since the policy optimization method has a large computational amount and is more complex to implement when the state search space is too large and there are too many parameters, this method proposes a reinforcement learning method for drone-assisted multi-node task offloading scheduling based on value function update. Summary of the Invention:

[0004] The object of the present invention is to solve the problem of maximizing the benefits of drones in the case of limited service time and user prior knowledge in edge computing scenarios.

[0005] The edge computing scenario mainly includes a tour path, several users, and edge servers. Among them, the user nodes have different task arrival streams. Tasks that are not collected by the drone will be stranded locally at the user nodes. Moreover, user tasks are sensitive to latency, and the task value will decay over time. To enable all nodes to obtain the services of the ground server, when achieving the goal of maximizing benefits, it is required that all user nodes participating in the offloading scheduling service have at least one experience of being provided with offloading services by the drone. For this reason, the present invention proposes a reinforcement learning method for drone-assisted multi-node task offloading scheduling. This method only requires very little prior knowledge and can obtain learning experience only by some simple interactions with the environment during the flight process, and then can obtain a near-optimal solution for the policy offloading scheduling path that maximizes the benefit goal.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a reinforcement learning method for drone-assisted multi-node task offloading scheduling, which is characterized by including the following steps:

[0007] Step 1: Refine the key features of the application model for the ground edge server and multiple user nodes on the drone-assisted tour path to collect and offload tasks, and construct a Markov model. In the Markov model constructed by the present invention, the state is represented by S = {loc, remtime, attri, flag}, where loc represents the current position of the drone on the tour path; remtime represents the remaining time for the drone to provide services; attri represents the node attribute currently visited. In the present invention, 0 represents a user node, and 1 represents a server node; flag is a service marking vector for user nodes, used to mark whether multiple user nodes on the tour path have been offloaded. Each row element can take values of 0 or 1. In the present invention, 0 represents that the current task has not been offloaded, and 1 represents that the current task has been offloaded. The action of the Markov model is the behavior of the agent in the environment. In actual decision-making, the drone will make corresponding actions according to the ε-greedy policy in different states. The action space of the Markov model in the present invention is all the nodes on the tour path, including user nodes and ground server nodes;

[0008] Step 2: Initialize the Q-table of the reinforcement learning method. The row attribute of the Q-table is the state in the Markov model, and the column attribute is the action in the Markov model. Each state-action corresponds to a value on the Q-table, and its size is the expected cumulative reward corresponding to the state-action. The initial values in the Q-table are randomly generated numbers after standard normalization, and these random numbers are all close to 0. Set the maximum iteration period and the starting state of the reinforcement learning method;

[0009] Step three, the present invention completes the constraint that all user nodes have been served at least once by setting small learning goals. Initialize the user node service mark vector flag in the starting state and set it all to 0, which means it has not been processed. When the drone arrives at the server node to unload the task, the number corresponding to the flag of the user node to which the unloaded task belongs will be set to 1. When the flag vector is all 1, the small goal is achieved. The reinforcement learning method will determine whether the small goal of the current state is completed by monitoring the flag mark in the state. The setting of small goals is to encourage the intelligent agent of the reinforcement learning method to actively explore the environment and find ways to achieve small goals, but our ultimate goal is to achieve the big goal of maximizing benefits. The reward setting of all small goals cannot affect the reward of the big goal. The present invention sets the reward for the agent to explore an unexplored user node before completing the small goal to 1, and its value is much smaller than the reward for arriving at the same node after completing the big goal. When the agent explores a user node that has already been explored, the reward is 0. When the agent reaches a ground server node, it will offload all the collected tasks to the server and update the flag. If the small goal is not completed, the reward is 0, but the actual reward will be accumulated and stored. When the small goal is completed, it will be given to the agent at one time.

[0010] Step 4. The present invention avoids sparse rewards by using pre-rewards. When the agent completes the small goal, the agent will receive the reward normally. In the actual environment, the benefits contained in the task can only be obtained when the drone arrives at the server to unload the task, and the number of user nodes is much greater than the number of server nodes. This will cause the agent to be in a 0 reward situation in most cases, that is, there will be a problem of sparse rewards. In order to improve the efficiency of training reinforcement learning methods, the present invention proposes a concept of pre-reward for the unloading scheduling scenario of drone-assisted edge computing, and allocates a small part of the sparse rewards that the agent can only obtain on the server to the user node in advance. This idea allows the environment to give the action a pre-reward in advance when the drone flies to a user node to perform task collection work. The size of the reward is related to the size of the reward obtained when the task is unloaded to the server after the total service time delay. The present invention sets the pre-reward size to the following formula through experimental experience:

[0011]

[0012] Where SF is the reduction factor, is the total number of tasks collected by the UAV from the nth user node for the tth time; σ n Indicates the value decay factor; value n Indicates the initial value of the nth node task; Total indicates the total duration;

[0013] The reason for attenuating the value of the total task duration and then reducing it by the SF factor is to ensure that the size of the pre-reward is much smaller than the actual reward obtained when it is offloaded to the server. Otherwise, the agent will abandon the behavior of offloading to the server, which is obviously contrary to the ultimate goal. That is, the following constraint formula needs to be satisfied:

[0014]

[0015] where represents the reward when the drone flies to the server node. Since we give a pre-reward when collecting user tasks, in order to ensure that the cumulative reward sum in the entire model is equal to the maximum remaining task value, we subtract the pre-reward part from the reward set for flying to the server.

[0016] In addition, there may still be unoffloaded user tasks before the end of the service time of the drone, and rewards are given for this part of the tasks when they are collected. Therefore, this part of the rewards needs to be removed additionally. That is, the penalty reward for the last decision when the remaining service time of the drone is 0 is as follows:

[0017]

[0018] Step 5: The present invention makes certain improvements to the ε-greedy strategy to ensure that in the initial stage of training, the agent is more inclined to non-experienced searches, and at the end of the training cycle, it is more inclined to the convergence of the training results. In the unimproved ε-greedy strategy, the action selection of the agent will choose exploratory actions and learning actions according to the comparison between a random number between 0 and 1 and the value of ε. That is, the larger the value of ε, the more inclined the agent is to explore new actions. To ensure the exploration of the agent in the initial stage of training, the value of ε in the ε-greedy strategy should not be too small. It is best to be close to 1 in the initial stage of training. As the number of algorithm iteration cycles increases, to ensure the convergence of the algorithm, the value of ε needs to be close to 0. Therefore, the present invention maps the number of iteration cycles to a value of ε through a negative exponential function, and the formula is shown as follows:

[0019] ε = e -β*episode

[0020] where the β parameter is used to control the growth rate of ε. To ensure that the algorithm can converge as the iteration cycle gets closer to the maximum training cycle number, β satisfies the following formula:

[0021]

[0022] Step 6: The process of the UAV providing services between nodes corresponds to the process of state transition of the Markov model. Each state transition of the Markov model generates a learning unit, including the agent's previous state, the action selected in the previous state, the reward given by the environment for this state transition, and the current state. After obtaining a learning unit, the temporal difference method is used to complete a single-step update of the model-free reinforcement learning algorithm. The present invention uses the update formula of Q-Learning to complete the above single-step update process and saves it in the Q-table. The update formula is as follows:

[0023]

[0024] Among them, Q(s,a) represents the cumulative expected value of taking action a in state s in the current Q-table; α and γ respectively represent the learning rate and reward decay factor of the reinforcement method; r is the reward given by the environment for the current state transition; Q(s',a') represents the maximum Q value in the next state s';

[0025] Step 7: Since the state-action space dimension in the environment where the problem solved by the present invention is located is large, and affected by the constraints of small learning objectives, in the case of limited training cycles, ordinary reinforcement learning methods are prone to fall into local optimal solutions. Therefore, the present invention adds a stack with memory ability to the agent of the reinforcement learning method, which can store the learning path of the current cycle. When the agent obtains a reward in a certain learning unit, the agent will compare the current encountered reward with the maximum reward encountered before. If it is larger, we will backtrack the entire path from the stack and re-learn the learning nodes on the entire path. This move is to enable the agent to perceive the large-reward path encountered accidentally in the next cycle. We do not make the agent sensitive to large rewards at the beginning, but only start to be sensitive to the large rewards encountered next after completing small objectives. The purpose of this is because the main objective in the early stage of the search is to achieve small objectives, and the rewards obtained are not real rewards. Being sensitive to large rewards will not only be sensitive to the immediate large rewards obtained by a certain action, but also be sensitive to the cumulative rewards obtained over an entire training cycle, that is, sensitive to the cumulative remaining value of task offloading obtained by completing our large objective, and will also backtrack the path with a higher cumulative remaining value and re-learn the entire path;

[0026] Step 8: Stop training when the maximum training cycle of the algorithm arrives, output the value of the Q-table, and select the action corresponding to the maximum state-action Q value of the current state using the greedy strategy starting from the start state as the action selection of the UAV in the actual application scenario. Repeat the above operation for the next state until the end state. Finally, an action sequence from the start state to the end state will be obtained, which is the offloading scheduling strategy of task offloading scheduling.

[0027] The beneficial effects of the present invention:

[0028] In view of the characteristics of the scenario of the drone-assisted multi-node task offloading and scheduling, the present invention improves the original Q-Learning algorithm, and innovatively proposes methods such as small learning objectives, pre-rewards, and large-reward sensitivity. Under the constraints of time-sensitivity of drone user nodes, etc., the present invention realizes the goal of maximizing the benefits of the drone by selecting a flight path through strategies within a limited service time. The method of the present invention does not require too much prior knowledge and does not need to deeply understand the in-depth information of each user node, which meets the requirements of privacy protection. Moreover, the present invention has good reusability in similar application scenarios, and the practical value of the invention is relatively strong. Description of the Drawings:

[0029] Figure 1 It is a schematic diagram of the Markov model provided by the embodiment of the present invention;

[0030] Figure 2 It is the flow of the reinforcement learning method provided by the embodiment of the present invention. Detailed Embodiment:

[0031] In order to enable relevant personnel to more clearly understand the technical content of the present invention, the following embodiments will be used for detailed introduction.

[0032] An implementation example of a reinforcement learning method for drone-assisted multi-node task offloading and scheduling includes the following steps:

[0033] Step 1: Initialize the Q-table. The row attributes are the states including environmental features, including four features: the current agent position, the remaining service time of the drone, the attributes of the current node, and the user node service mark vector. The column attributes are different action selections, and the action selections are natural numbers from 1 to the total number of nodes. The initial value of each state-action pair on the Q-table is set to a random number between (-0.1, 0.1);

[0034] Step 2: Initialize the maximum number of cycles, the reinforcement learning parameters, the current maximum target reward, and the maximum learning unit reward. Set the current training cycle number to 0;

[0035] Step 3: Initialize the current state of the agent as the starting state, the drone position is at the starting point of the tour path, the remaining service time is the total service time that the drone can provide, the current node attribute is 0, and the user node service mark vector is all set to 0;

[0036] Step 4: Update the current greedy policy parameter ε = e -β*episode , initialize the memory stack, and set the learning unit memory stack to be empty;

[0037] Step 5: Execute the ε-greedy strategy. Obtain a random number. If the random number is less than ε, randomly select a node from all nodes as the action selection for the current state. If the random number is greater than ε, consult the row in the Q-table corresponding to the current state and select the action with the maximum Q-value as the action selection for the current state;

[0038] Step 6: The agent executes the action to move from the current state to the next state. If the action selection is to reach a user node, the drone will descend from the cruising altitude to the service altitude, collect all the tasks remaining at the current user node. At this time, judge whether the small learning objective has been completed. If it has not been completed and the user node has not been visited, obtain an exploration small reward 1 before the small objective is not completed. If the user node has been visited, there is no reward. If the small objective has been completed, obtain the pre-reward for the tasks uploaded on the current user node, and its size is the value of the time delay decay from the task generation time to the end time of the drone service multiplied by a reduction factor of 0.1. If the action selection is to reach a server node, the drone will also descend from the cruising altitude to the service altitude. The difference is that it will unload all the collected tasks, accumulate the actual rewards obtained from the environment after unloading the tasks, and then set the service mark attribute of the user node where the tasks have been unloaded to 1. At this time, judge whether the small learning objective has been completed. If it has not been completed, give a reward of 0. But if all the values of the service mark vector flag in the state are 1, it means that the small objective is just completed in this task unloading, and the actual rewards accumulated before the task unloading will be given to the reward of the current step. If the small learning objective has been completed before the task unloading, the agent will calculate the reward obtained from the currently unloaded tasks, and at the same time exclude the size of the pre-reward given when the drone arrives at the user node. The difference is the single-step reward for this state transition;

[0039] Step 7: After determining the learning unit reward for this state transition, update the next state and use the update formula of Q-Learning Update the Q-table, where α is the learning rate. α affects the convergence speed of the algorithm. When the learning rate α is too large, the convergence speed is fast but it may cause the model to fall into the local optimal solution prematurely. When it is too small, the convergence speed is too slow. γ is the reward decay coefficient. γ is used to balance the influence of the subsequent rewards on the current immediate reward. The larger γ is, the closer the Q-value of the current state action pair is to the value of the large target, and the smaller γ is, the closer the reinforcement learning method is to the greedy algorithm. Finally, push the complete learning unit into the learning unit memory stack;

[0040] Step 8: If the small objective has been completed, judge whether the size of the reward obtained in Step 6 is greater than the maximum learning unit reward. If it is greater than the maximum learning unit reward, pop the learning unit memory stack in turn, take out the learning unit, and re-use the update of Q-Learning Update the Q-table once and restore the memory stack of the learning unit;

[0041] Step 9: If the agent has not reached the end state, keep repeating Steps 5, 6, 7, and 8 until the UAV service time ends and enters the end state. If there are still tasks to be unloaded when the agent reaches the end state, a penalty reward equal to the sum of all task pre-rewards will be given to the agent. Calculate the cumulative task value gain from the start state to the end state in the current cycle. If it is greater than the current maximum target reward, pop the memory stack of the learning unit in sequence, take out the learning unit, and reuse the update formula of Q-Learning Update the Q-table once;

[0042] Step 10: If the number of training cycles has not reached the maximum number of cycles, repeat Steps 3, 4, 5, 6, 7, 8, and 9 until the maximum training cycle is reached. If the maximum training cycle is reached, stop training using the reinforcement learning method, output the Q-table, and starting from the initial state, select the action with the maximum Q value in the corresponding Q-table according to the greedy algorithm. Repeat the above operations after transferring to the next state until the end state. Record all the action selections to obtain an unloading scheduling decision sequence, and output the result as a solution to the problem of maximizing the benefit in the case of the limited service time of the UAV and the user's prior knowledge in the edge computing scenario.

[0043] It should be known that the parts not elaborated in detail in this specification all belong to the prior art. Those skilled in the relevant art should understand that the above embodiments are only for helping readers understand the principles and implementation methods of the present invention, and the scope protected by the present invention is not limited to such embodiments. All equivalent replacements made on the basis of the present invention are within the protection scope of the rights of the present invention.

Claims

1. A method for multi-node task offloading scheduling assisted by an unmanned aerial vehicle, characterized in that , The implementation process of this method is as follows: Step 1: The drone flies along the tour path, and when necessary, descends to a lower altitude to assist in collecting data from multiple user nodes on the ground at close range and offloads tasks at the edge server. A Markov model is constructed for this application scenario. Step 2: Initialize the Q-table of the reinforcement learning method. The row attribute of the Q-table is the state in the Markov model, and the column attribute is the action in the Markov model. Each state-action corresponds to a state-action value on the Q-table, and its magnitude is the expected cumulative reward corresponding to the state-action. The initial values in the Q-table are randomly generated numbers after standard normalization, and these random numbers are all close to 0. Step 3: Set the constraints in the application scenario as small goals of reinforcement learning, and maximize the remaining value of the tasks obtained after policy scheduling as the large goal. The large goal must be achieved after the small goal. An exploratory small reward is set for the small goal of reinforcement learning, and its role is to enable the agent to complete the small goal normally without being affected by the rewards of the large goal. In order to make the cumulative reward obtained by the agent in the interaction with the environment meet the requirements of the large goal, a storage interval is used to memorize the real rewards obtained from the environment on the path of completing the small goal. When the small goal is completed, the agent will obtain the cumulative real rewards stored in the storage interval at one time. When the small goal is not completed, the exploratory small reward obtained by the agent is less than the real reward obtained during the process of achieving the large goal after completing the small goal. Step 5: Set a pre-reward. The pre-reward is a reward obtained in advance when the drone provides services to user nodes. The magnitude of the pre-reward is set as a part of the reward that should be obtained after the task is offloaded to the server. A penalty reward of the same magnitude will be given to all tasks that are not offloaded to the server at the end state. The pre-reward is set as follows: where SF is the scaling factor, is the total number of tasks collected by the UAV from the nth user node at the tth time; σ n represents the value decay factor; value n represents the initial value of the task of the nth user node; Total represents the total duration; Step 6: At the beginning of a training cycle of the reinforcement learning method, the agent starts from the initial state on the Markov model and selects the next action of the current state for the agent according to the improved ε-greedy policy. Step 7: After the agent makes an action selection, it will reach the next environmental state, and the environmental state will give corresponding rewards according to the current features. Step 8: When the maximum number of training cycles of the algorithm reaches, stop training, output the maximum cumulative reward of training convergence, and according to the values in the Q-table, use the greedy policy to obtain an action sequence from the start state to the end state, which is the action policy for multi-node task offloading scheduling.

2. The method for multi-node task offloading scheduling assisted by a drone according to claim 1, wherein In the Markov model constructed in Step 1, the state is represented by S = {loc, remtime, attri, flag}, where loc represents the current position of the drone on the tour path; remtime represents the remaining time for the drone to provide services; attri represents the attributes of the currently visited node; flag is a user node service marking vector used to mark whether multiple user nodes on the tour path have been offloaded and processed. In the Markov model, the action space is the positions of multiple nodes on the tour path.

3. A method for multi-node task offloading scheduling assisted by a drone according to claim 1 or 2, characterized in that When the agent encounters a user node during training, there are two types of rewards, namely the actual environment interaction reward and the pre-reward. The actual environment reward can only obtain a certain amount of reward when the drone provides an offloading service for the server, and the reward obtained for providing a task collection service for the user node is 0.

4. A method for multi-node task offloading scheduling assisted by a drone, characterized in that Among them, the size of ε in the improved ε-greedy strategy is a negative exponential function related to the training cycle. In the initial stage of training, the agent is more inclined to non-experiential search, while at the end of the training cycle, the agent is more inclined to the convergence of the training results. The improved ε-greedy strategy maps the number of iteration cycles to ε through a negative exponential function, and the formula is shown as follows: ε = e -β*episode ; Among them, the β parameter is used to control the growth rate of ε. To ensure that the algorithm can converge as the iteration cycle gets closer and closer to the maximum training cycle number, β satisfies the following formula:

5. A method for multi-node task offloading scheduling assisted by a drone according to claim 4, characterized in that Step six realizes the process of reaching from one state to another state, and a learning unit for state transition will be generated. This learning unit includes the previous state eigenvalue, the selected action, the obtained unit reward, and the next state. The agent will learn and update the Q-table based on the update formula of Q-Learning. Finally, when the drone service time ends, it will enter the end state, and the Q-table will be inherited to start the next training cycle.

6. The method for multi-node task offloading scheduling assisted by a drone according to claim 5, characterized in that During each training cycle, the agent will have a learning unit stack, a maximum unit reward, and a maximum cumulative reward, which are used to store all the learning units of the agent during this training cycle, the maximum learning unit reward encountered in historical training, and the maximum cumulative reward encountered in history, respectively. If a larger learning unit reward is encountered, the agent will first copy the learning unit stack, and then pop all the learning units in the stack one by one and update the Q-table again. If a larger cumulative reward is encountered, the agent will also pop all the learning units in the stack one by one and update the Q-table again. Each time a new training starts, the stack will be reset to empty.

7. The method for multi-node task offloading scheduling assisted by a drone according to claim 6, wherein The positions of the user node and the ground server node are both on the tour path, and the drone can choose to fly along the path clockwise or counterclockwise. The loop height of the drone is not at the same height as the height for task collection and offloading. When the drone executes tasks, it needs to descend from the cruising height to the service height to provide services. And the tasks generated by the user node arrive evenly, but the arrival times of different tasks are different, and the offloading rates and the values calculated by the tasks are also different. The unprocessed tasks will pile up locally at the user, and the task value is time-sensitive before offloading, and the task value will decrease with the passage of time.

8. A method for multi-node task offloading scheduling assisted by a drone, characterized in that The drone needs to abide by the following restrictions during the offloading scheduling process: a) The nodes selected by the drone for service must be the user nodes and server nodes that have been registered in the environment and need services. b) All registered nodes that need services must have received at least one offloading service. c) The task completion time of the drone must be within the limited service completion time, and the drone stops serving when the service time ends.

Citation Information

Patent Citations

  • Unmanned aerial vehicle task unloading method and system based on reinforcement learning in edge calculation

    CN111787509A

  • Air-ground combined mobile edge computing unloading optimization method

    CN112911648A