Unmanned aerial vehicle task unloading rapid adaptation method based on reinforcement learning

Through a reinforcement learning method, using a reinforcement learning model with decoupling strategies and environmental representation, the flight trajectory and mission offload decision of the drone are optimized, which solves the problem that the drone network is difficult to adapt quickly in a dynamic environment, and achieves efficient network services and resource utilization.

CN119938174AActive Publication Date: 2025-05-06NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510350262.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-06
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

It is difficult for existing drone networks to achieve rapid adaptation and generalization in dynamic environments, resulting in inefficient network coverage and waste of resources.

Method used

Using reinforcement learning methods, we use the reinforcement learning model to establish a decoupled strategy and environmental representation, and use offline experience to train the environment embedding vector and policy embedding vector, and combine the value prediction network and gradient rise algorithm to optimize the flight trajectory and mission offload decision of the drone to achieve rapid adaptation in the new environment.

Benefits of technology

It significantly improves the adaptability efficiency of the drone network in dynamic environments, reduces the generalization cost, avoids the need for repeated training models, and improves the efficiency and quality of network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938174A_ABST
    Figure CN119938174A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle task unloading rapid adaptation method based on reinforcement learning, and relates to the technical field of unmanned aerial vehicle flight control. According to the method, learning is carried out from offline experience collected in environments with different states, online strategy adaptation is executed in an environment with a new state dynamic state, the method is suitable for a real scene with high online interaction cost, and generalization can be achieved on the basis of offline experience data. According to the method, a reinforcement learning model based on a decoupling strategy and environment representation and a strategy adaptation algorithm based on gradient rising are established to obtain the action of the unmanned aerial vehicle under the strategy of optimal unmanned aerial vehicle performance, so that the unmanned aerial vehicle and a base station are combined to provide edge computing service for a computing task generated by ground terminal equipment. According to the method, a small amount of online interaction between the unmanned aerial vehicle and the current environment is needed only in the stage of online adaptation to the test environment, and compared with other mainstream methods needing a large number of online interaction samples, the training cost of the model can be greatly reduced in the actual environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle flight control, and in particular to a method for rapid adaptation of unloading tasks of unmanned aerial vehicles based on reinforcement learning. Background Art

[0002] With the rapid development of IoT technology, the traffic demand of various services has exploded. For example, the widespread application of big data and cloud computing has made it difficult for traditional ground wireless networks to meet the traffic demand of massive applications. In addition, due to the limitations of network capacity and coverage, ground base stations cannot be quickly deployed to provide network services when outdoor disasters such as forest fires occur. Therefore, there is an urgent need for a network architecture that is easy to deploy and has good scalability to cope with complex and harsh environments and meet the traffic needs of many mobile users.

[0003] Given the advantages of low cost, easy scalability and rapid deployment, drones have attracted widespread attention from academia and industry as aerial base stations and have become a potential effective solution. However, unlike traditional ground static base stations, drones have limited energy storage and transmission power, making it difficult to support long-term deployment, and their service range is also limited by time. In addition, the location of users often changes dynamically, and network demand in different areas also fluctuates accordingly. In traditional drone networks, since a single drone can only perceive user information within its own service range and lacks collaboration with other drones, there is often a problem of multiple drones concentrating in the same area, resulting in low network coverage efficiency and waste of drone resources. Therefore, designing a reasonable flight trajectory for drones to maximize service coverage within a limited flight time has become the key to improving the efficiency and quality of drone network services.

[0004] Deep reinforcement learning (DRL) has become an important tool for solving complex decision-making problems. The core goal of reinforcement learning is to learn a strategy through interaction with the environment to maximize the cumulative reward. Recently, DRL has made significant achievements in some high-dimensional state observation and planning problems, such as video games, Go, and robot control. Inspired by this, the DRL framework has gradually become a popular solution for solving traditional optimization problems in drone networks. However, the generalization problem of reinforcement learning is an important challenge in current artificial intelligence research. The generalization of reinforcement learning aims to enable reinforcement learning agents to perform well in unknown environments outside the training environment. Traditional reinforcement learning methods usually rely on repeated interactions and optimizations of a specific environment. However, the limitation of this method is that the trained strategy is often only effective in a specific training environment, and when the dynamics, observations, rewards and other characteristics of the environment change, the performance of the strategy may drop sharply. This lack of generalization ability reflects two key challenges of reinforcement learning in the real world: one is the uncertainty of the environment. The real environment is usually diverse and complex, and cannot be fully simulated by a limited training environment. For example, a robot may need to deal with ground with different materials, and the friction and elasticity of each material will affect the effect of the action. Second, the costly training process. In many reinforcement learning application scenarios (such as medical diagnosis or industrial control), it may not be feasible to conduct a large number of experiments directly in the real environment because it is both expensive and risky. Therefore, generalization must be achieved based on limited training data or offline data.

[0005] Many methods have been proposed in the literature on generalization in reinforcement learning, such as data augmentation methods proposed by Yarats et al. in "Deep variational information bottleneck" and domain randomization methods proposed by Peng et al. in "Learning dexterous in-hand manipulation" to increase the similarity between the training environment and the test environment (i.e., the environment for generalization or adaptation). These methods generally require knowledge of the variation of the environment and the ability to generate the environment. In contrast, some other works aim to achieve fast adaptation in the test environment without such access and knowledge. For example, Finn et al. proposed gradient-based meta-RL in "Outperforming the atari human benchmark", aiming to learn a meta-policy that can adapt to the test environment within a few policy gradient steps. Also in this branch, Rakelly et al. proposed context-based meta-RL in "A contrastive log-ratio upper bound of mutual information", which exploited a context-conditional policy that allowed adaptation through generalization between contexts. In "Fast reinforcement learning via slow reinforcement learning", Fu et al. proposed to learn useful contextual representations in the training environment to capture changes; the context of the test environment can be inferred from a few exploration interaction experiences. Current solutions to this problem have improved generalization capabilities to a certain extent, but they usually rely on knowledge of environmental changes and require manual design of randomization strategies, which is expensive. Meta-reinforcement learning solutions allow agents to interact with the training environment arbitrarily within the interaction budget. However, in real-world problems, online interactions are usually expensive, while offline experiences are often available and relatively abundant. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a method for rapid adaptation of drone task offloading based on reinforcement learning in response to the deficiencies of the above-mentioned prior art. The method learns from offline experience collected in environments with different states, and performs online strategy adaptation in an environment with new state dynamics. The method is suitable for real-life scenarios with high online interaction costs, and can achieve generalization based on offline experience data, thereby reducing the generalization cost.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: The present invention provides a method for rapid adaptation of unmanned aerial vehicle task offloading based on reinforcement learning, involving an unmanned aerial vehicle, a base station and a ground terminal device. The unmanned aerial vehicle collects computing tasks generated by the ground terminal device, and maximizes the number of collected computing tasks by optimizing the flight trajectory of the unmanned aerial vehicle. By optimizing the task offloading decision of the unmanned aerial vehicle, the computing service efficiency provided by the unmanned aerial vehicle and the base station for the computing tasks generated by the ground terminal device is maximized, including the following steps: Step 1: Establish a reinforcement learning model based on decoupling strategy and environment representation, and initialize the parameters of the reinforcement learning model based on decoupling strategy and environment representation; The reinforcement learning model based on the decoupling strategy and environment representation includes several drones and an environment, wherein the several drones interact with the environment, learn iteratively, and optimize the actions of the several drones in the environment; the environment includes several ground terminal devices and several base stations; the actions of the several drones in the environment include collecting computing tasks generated by several ground terminal devices in the environment, making task offloading decisions, and jointly providing computing services for several ground terminal devices in the environment with several base stations in the environment, wherein the task offloading decision is a decision to offload the computing tasks collected by the drones to several base stations in the environment; Initialize the parameters of the reinforcement learning model based on the decoupling strategy and environment representation, including the state space S of the environment, the reward function R, and the action space A of the drone. The specific method is: Divide the environment into training environment and test environment, and set the training environment set to , the test environment set is , the training strategy set is ,in, is the number of training environments, is the number of test environments, Create an offline experience buffer for the number of training strategies As training samples, For the Training Environment Use the Training strategies Generated offline experience; Assume that the drone starts from the initial state and continuously interacts with several ground terminal devices and several base stations in the environment. The trajectory of the drone is , used to record the state transition of the drone and the rewards obtained during the state transition, where: is the environmental state, For the drone's movements, Rewards for drones; set in training environment A set of drone trajectories generated in Contextual information for the environment , and set context information Environmental status in and the actions of drones For drone behavior ; Step 2: Use training samples, different training environments, and different training strategies to train a reinforcement learning model based on decoupled strategy and environment representation, obtain environment embedding vectors of different training environments and strategy embedding vectors of different training strategies, and predict the value prediction network reward value obtained by using different training strategies in different training environments; Step 2.1: Build a context encoder By comparing the context information of the drone in different training environments, the environment embedding vectors of different training environments are extracted; For any training environment in the training environment set , using the training environment Offline experience Contextual information sampled in As anchor sample and positive samples , from another training environment Offline experience The context information sampled in is used as negative samples , based on anchor samples and positive samples Generate similar sample pairs based on anchor samples and negative samples Generate different sample pairs; Building a context encoder , used to convert the training environment Offline experience Contextual information sampled in Mapping to environment embedding vector , context encoder The loss function for: (1) in, For the context encoder The updated parameter matrix, is the mathematical expectation, Anchor sample The environment embedding vector, , The context encoder Extracted positive samples , negative samples The environment embedding vector, Anchor sample The environment embedding vector The transposed vector of ; Use training samples to train context encoder Train and get the trained context encoder , and based on the context encoder Extract training environment Contextual information The environment embedding vector ;

[0008] Step 2.2: Build a policy encoder and policy decoder , through the policy encoder Extracting training strategies The policy embedding vector , and then through the strategy decoder Prediction Training Strategy The action of dismounting the drone; Building a policy encoder , for the training strategy The behavior of drones , using the training environment Offline experience The behavior of the drones sampled in The embedding vector of The policy embedding vector ; Building a policy decoder , the environmental state and training strategies The policy embedding vector As a policy decoder Input, prediction training strategy The action of dismounting the drone; By minimizing l 2 Loss Function Update Strategy Encoder and policy decoder , as shown in the following formula: (2) Get the trained policy encoder and policy decoder , through the trained policy encoder Extracting training strategies The policy embedding vector , and then through the trained strategy decoder Prediction Training Strategy The action of dismounting the drone; Step 2.3, build parameters are Value prediction network , using the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value ; Value Prediction Network It is a multi-layer nonlinear neural network. , based on the environmental status , Training strategy The policy embedding vector , Training Environment The environment embedding vector As a value prediction network The Monte Carlo method is used to train the value prediction network. , value prediction network The loss function As shown in the following formula: (3) in, , For long-term incentive benefits, γ is the discount factor, is the time step t The reward function is, T is the total number of time steps, let This is the Monte Carlo regression. is the initial environment state; Use the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value ; Step 3: Based on the reinforcement learning model of decoupling strategy and environment representation, a strategy adaptation algorithm based on gradient ascent is established to optimize the strategy embedding vector of the drone taking different strategies in the test environment, and the action of the drone under the strategy with the best performance is obtained to interact with the current environment; Step 3.1: From the training strategy set Sampling any training strategy , the drone uses this training strategy With test environment Interact and collect test environment Contextual information ; Step 3.2: Predict network reward value based on value Establish a gradient ascent-based strategy adaptation algorithm and use the gradient ascent-based strategy adaptation algorithm to optimize the test environment Training strategies for drones The strategy embedding vector is used to obtain the optimized training strategy The policy embedding vector ; The gradient-based strategy adaptation algorithm sets the number of iterations , each iteration uses the training strategy Sampling a UAV trajectory , For test environment The environmental status, For training strategy The action of the drone, Action for the drone The new environment obtained The environmental state of the drone As an environmental encoder Input to get the test environment The environment embedding vector ; The drone trajectory As a policy encoder Input, get the training strategy The policy embedding vector ; Use the gradient ascent algorithm to follow the value prediction network Predicted value Predicted network reward value With the learning rate η Added direction optimization training strategy The policy embedding vector , as shown in the following formula: (4) Until completion Iterations, reaching the preset number of iterations After that, the optimized training strategy is obtained. The policy embedding vector ; Step 3.3: Use the policy decoder Based on the optimized training strategy The policy embedding vector Get the drone's actions under the best performance strategy and interact with the current environment; Step 4: Implement drone-assisted mobile edge computing based on the drone's actions under the best drone performance strategy, so that the drone and the base station jointly provide edge computing services for the computing tasks generated by the ground terminal equipment; Step 4.1, generating an optimized UAV flight trajectory and task offloading decision according to the UAV action under the strategy with the best UAV performance; Step 4.2: The drone collects and stores the computing tasks of the ground terminal equipment. According to the optimized flight trajectory of the drone, the number of tasks generated by the ground terminal equipment collected by the drone is maximized and stored in the drone's task queue, completing the collection of computing tasks of the ground terminal equipment. Step 4.3: Determine the allocation result of the computing tasks of the ground terminal device according to the task offloading decision of the UAV, including the number of computing tasks allocated to the UAV and the number of computing tasks offloaded to the base station; Step 4.4, executing the computing tasks according to the computing task allocation result; the UAV performs the computing of the tasks that do not need to be offloaded, and transmits the tasks that need to be offloaded from the UAV to the base station, and performs the computing on the base station; Step 4.5: After the task calculation in the base station and the UAV is completed, the task calculation results in the base station and the UAV are returned to the ground terminal device.

[0009] The beneficial effects of adopting the above technical solution are: the invention provides a method for rapid adaptation of unmanned aerial vehicle task offloading based on reinforcement learning, which separates the policy representation from the representation of environmental dynamic features, trains the two representation spaces separately through a contrastive learning mechanism, and uses mutual information optimization to ensure decoupling while maintaining the integrity of the representation. The value prediction network reward value is introduced to evaluate the value of different strategies and dynamic combinations of the environment, and offline experience learning is used to infer the environmental representation through a small amount of online data in the new environment to achieve rapid adjustment of the strategy. A gradient-ascending-based policy adaptation algorithm is established to optimize the policy embedding vectors of drones taking different strategies in the test environment, train the environment and policy representation on offline experience data of multiple dynamic features, and efficiently optimize the strategy in the test environment, avoiding retraining the entire model and significantly improving the adaptation efficiency. The training data of the invention is mainly offline samples, and only a small amount of online interaction between the drone and the current environment is required in the test environment. Compared with other mainstream methods that require a large number of online interaction samples, it can greatly reduce the training cost of the model in the actual environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 A schematic diagram of a reinforcement learning framework provided by an embodiment of the present invention; Figure 2 A schematic diagram of a strategy adaptation algorithm based on gradient ascent provided in an embodiment of the present invention; Figure 3 A schematic diagram of an experimental environment for a reinforcement learning rapid adaptation method based on a decoupling strategy and environment representation provided by an embodiment of the present invention; Figure 4The comparative test results provided for the embodiments of the present invention, wherein (a) is a schematic diagram of the trajectory of the drone without generalization, (b) is a schematic diagram of the trajectory of the drone after generalization, and (c) is a graph showing the experimental results of the reward value comparison between this embodiment and other methods. DETAILED DESCRIPTION

[0011] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0012] The present embodiment is a method for rapid adaptation of unmanned aerial vehicle task offloading based on reinforcement learning, involving unmanned aerial vehicles, base stations and ground terminal equipment. The unmanned aerial vehicle collects computing tasks generated by the ground terminal equipment, and maximizes the number of collected computing tasks by optimizing the flight trajectory of the unmanned aerial vehicle. By optimizing the task offloading decision of the unmanned aerial vehicle, the computing service efficiency provided by the unmanned aerial vehicle and the base station for the computing tasks generated by the ground terminal equipment is maximized. Figure 1 As shown, the following steps are included: Step 1: Establish a reinforcement learning model based on decoupling strategy and environment representation, and initialize the parameters of the reinforcement learning model based on decoupling strategy and environment representation; The reinforcement learning model based on the decoupling strategy and the environment representation includes a number of drones and an environment, wherein the drones interact with the environment, learn iteratively, and optimize the actions of the drones in the environment; the environment includes a number of ground terminal devices and a number of base stations; the actions of the drones in the environment include collecting computing tasks generated by a number of ground terminal devices in the environment, making task offloading decisions, and jointly providing computing services for a number of ground terminal devices in the environment with a number of base stations in the environment, wherein the task offloading decision is a decision to offload the computing tasks collected by the drones to a number of base stations in the environment; In the reinforcement learning model based on decoupling strategy and environment representation, the interaction between the environment and the UAV is modeled as a Markov decision process (MDP), which is used to describe the state transition of the environment, the reward mechanism and the action selection of the UAV during the interaction process, including: state space S, action space A, transition probability P, reward function R, discount factor γ; the state space S is the set of all possible environmental states, which is used to describe the specific situation of the environment at each time point; the action space A is the set of all possible actions of the UAV, and each action of the UAV will have an impact on the state of the environment; the transition probability P is the probability that the current environmental state is transferred to another environmental state after the UAV takes a certain action in a certain environmental state, and the transition probability P is a probability model based on environmental dynamics; the reward function R is the immediate reward obtained by the UAV after performing a certain action in a given environmental state, and the reward signal is the basis for evaluating the quality of the UAV's behavior; the discount factor γ is the discount coefficient for calculating future rewards, which reflects the importance of the UAV to future rewards. The value of the discount factor γ is between 0 and 1. The closer the value of the discount factor γ is to 1, the more the UAV values ​​future rewards; Initialize the parameters of the reinforcement learning model based on the decoupling strategy and environment representation, including the state space S of the environment, the reward function R, and the action space A of the drone. The specific method is: Initialize the state space S of the environment based on the observation values ​​of the drone. The observation values ​​of the drone include: the location of the drone and the locations of the three ground terminal devices closest to the drone; initialize the action space A of the drone, including: the movement direction, movement speed, and task offloading decision of the drone; initialize the reward function R, which is set to the sum of the three target reward functions of the drone's total energy consumption Etotal, task delay Ttotal, and task collection number Ntotal; Divide the environment into training environment and test environment, and set the training environment set to , the test environment set is , the training strategy set is ,in, is the number of training environments, is the number of test environments, Create an offline experience buffer for the number of training strategies As training samples, For the Training Environment Use the Training strategies Generated offline experience; Assume that the drone starts from the initial state and continuously interacts with several ground terminal devices and several base stations in the environment. The trajectory of the drone is , used to record the state transition of the drone and the rewards obtained during the state transition, where: is the environmental state, For the drone's movements, Rewards for drones; set in training environment A set of drone trajectories generated in Contextual information for the environment , and set context information Environmental status in and the actions of drones For drone behavior .

[0013] Step 2: Use training samples and different training environments to train a reinforcement learning model based on decoupled strategies and environment representation, obtain environment embedding vectors of different training environments and strategy embedding vectors of different training strategies, and predict the value prediction network reward value that can be obtained by using different training strategies in different training environments; Step 2.1: Build a context encoder By comparing the context information of the drone in different training environments, the environment embedding vectors of different training environments are extracted; For any training environment in the training environment set , using the training environment Offline experience Contextual information sampled in As anchor sample and positive samples , from another training environment Offline experience The context information sampled in is used as negative samples , based on anchor samples and positive samples Generate similar sample pairs based on anchor samples and negative samples Generate different sample pairs; Building a context encoder , used to convert the training environment Offline experience Contextual information sampled in Mapping to environment embedding vector , context encoder The loss function for: (1) in, For the context encoder The updated parameter matrix, is the mathematical expectation, Anchor sample The environment embedding vector, , The context encoder Extracted positive samples , negative samples The environment embedding vector, Anchor sample The environment embedding vector The transposed vector of ; Use training samples to train context encoder Train and get the trained context encoder , and based on the context encoder Extract training environment Contextual information The environment embedding vector ; Step 2.2: Build a policy encoder and policy decoder , through the policy encoder Extracting training strategies The policy embedding vector , and then through the strategy decoder Prediction Training Strategy The action of dismounting the drone; Building a policy encoder , for the training strategy The behavior of drones , using the training environment Offline experience The behavior of the drones sampled in The embedding vector of The policy embedding vector ; Building a policy decoder , the environmental state and training strategies The policy embedding vector As a policy decoder Input, prediction training strategy The action of dismounting the drone; By minimizing l 2 Loss Function Update Strategy Encoder and policy decoder , as shown in the following formula: (2) Get the trained policy encoder and policy decoder , through the trained policy encoder Extracting training strategies The policy embedding vector , and then through the trained strategy decoder Prediction Training Strategy The action of dismounting the drone; Step 2.3, build parameters are Value prediction network , using the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value ; Value Prediction Network It is a multi-layer nonlinear neural network. , based on the environmental status , Training strategy The policy embedding vector , Training Environment The environment embedding vector As a value prediction network The Monte Carlo method is used to train the value prediction network. , value prediction network The loss function As shown in the following formula: (3) in, , For long-term incentive benefits, γ is the discount factor, is the time step t The reward function is, T is the total number of time steps, let This is the Monte Carlo regression. is the initial environment state; Use the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value ,The value prediction network algorithm is shown in Table 1; Table 1 Value prediction network algorithm:

[0014] Step 3: Based on the reinforcement learning model of decoupling strategy and environment representation, a strategy adaptation algorithm based on gradient ascent is established to optimize the strategy embedding vector of the drone taking different strategies in the test environment, and the action of the drone under the strategy with the best performance is obtained to interact with the current environment; Step 3.1: From the training strategy set Sampling any training strategy , the drone uses this training strategy With test environment Interact and collect test environment Contextual information ; Step 3.2: Predict network reward value based on value Establish a gradient ascent-based strategy adaptation algorithm and use the gradient ascent-based strategy adaptation algorithm to optimize the test environment Training strategies for drones The strategy embedding vector is used to obtain the optimized training strategy The policy embedding vector ; Setting the number of iterations for the gradient ascent-based policy adaptation algorithm , each iteration uses the training strategy Sampling a UAV trajectory , For test environment The environmental status, For training strategy The action of the drone, Action for the drone The new environment obtained The environmental state of the drone As an environmental encoder Input to get the test environment The environment embedding vector ; The drone trajectory As a policy encoder Input, get the training strategy The policy embedding vector ; Use the gradient ascent algorithm to follow the value prediction network Predicted value Predicted network reward value With the learning rate η Added direction optimization training strategy The policy embedding vector , as shown in the following formula: (4) Until completion Iterations, reaching the preset number of iterations After that, the optimized training strategy is obtained. The policy embedding vector ,The strategy adaptation algorithm based on the representation gradient ascent is shown in Table 2; Table 2 shows the strategy adaptation algorithm based on gradient ascent:

[0015] Using the Policy Decoder Based on the optimized training strategy The policy embedding vector Get the drone's actions under the best performance strategy and interact with the current environment; In this embodiment, the best performing strategy encountered during the training process is used as the adjusted strategy, and the strategy and the current environment are used as the input of the strategy decoder to obtain the action that the current drone should take, such as Figure 2 shown.

[0016] Step 4: Implement drone-assisted mobile edge computing based on the drone's actions under the best drone performance strategy, so that the drone and the base station jointly provide edge computing services for the computing tasks generated by the ground terminal equipment; In the UAV-assisted mobile edge computing model established based on steps 2 and 3, the generalized model is trained in the collected offline data set to allow the UAV to quickly adapt to the real environment. By optimizing the trajectory and task offloading decision of the UAV, the UAV and the edge cloud can jointly provide edge computing services for the computing tasks generated by the ground terminal devices. UAV-assisted mobile edge computing includes: Step 4.1, generating an optimized UAV flight trajectory and task offloading decision according to the UAV action under the strategy with the best UAV performance; Step 4.2: The drone collects and stores the computing tasks of the ground terminal equipment. According to the optimized flight trajectory of the drone, the number of tasks generated by the ground terminal equipment collected by the drone is maximized and stored in the drone's task queue, completing the collection of computing tasks of the ground terminal equipment. Step 4.3: Determine the allocation result of the computing tasks of the ground terminal device according to the task offloading decision of the UAV, including the number of computing tasks allocated to the UAV and the number of computing tasks offloaded to the base station; Step 4.4, executing the computing tasks according to the computing task allocation result; the UAV performs the computing of the tasks that do not need to be offloaded, and transmits the tasks that need to be offloaded from the UAV to the base station, and performs the computing on the base station; Step 4.5: After the task calculation in the base station and the UAV is completed, the task calculation results in the base station and the UAV are returned to the terminal device.

[0017] In this embodiment, an experimental environment for a reinforcement learning rapid adaptation method based on a decoupling strategy and environment representation is constructed, such as Figure 3 Shown, including N drones, Mground terminal devices and 1 base station, and use coordinates to represent their locations. The ground terminal devices have no computing power, while the computing and communication resources that drones can provide are limited. Unlike drones, base stations are composed of edge servers, which are small-scale cloud data centers that provide real-time data processing, analysis and decision-making. Each ground terminal device needs to process computing-intensive tasks, and drones provide edge computing services for ground terminal devices. Since ground terminal devices have no computing power, the computing tasks are completely offloaded to drones. However, the computing power of drones is limited. This embodiment uses a joint method of drones and edge clouds to provide offloading services for tasks generated by ground terminal devices. In this embodiment, the trajectory of the drone is planned to save energy, reduce latency, collect as many tasks as possible, and avoid collisions between drones.

[0018] In order to demonstrate the effectiveness of the present invention, this embodiment creates many environments with different characteristics. Then, the environment set is split into training and test subsets so that when testing, the agent must find a strategy that performs well in an environment that has not been seen before. Finally, the average reward value obtained by the generalization method proposed in the present invention in the test set is calculated and compared with the average value of the rewards obtained by the models trained in each test environment using the PPO algorithm, and the average value of the rewards obtained by the models trained in each training set in each test set. At the same time, the two current mainstream reinforcement learning generalization algorithms, PDVF and MAML, are compared. The experimental results are shown in Figure 2. Figure 4 As shown, (a) is a schematic diagram of the drone trajectory without generalization, (b) is a schematic diagram of the drone trajectory after generalization, and (c) is a graph showing the experimental results of the reward value comparison between this embodiment and other methods.

[0019] The experimental results show that the average reward obtained by the method proposed in this invention (ThisMethod) is significantly higher than the average reward obtained by the PPO model without generalization (PPOInTrain). It can also be seen from the trajectories of the four drones that the collisions of drones are significantly reduced. The performance of this method is higher than that of MAML and PDVF, and slightly lower than that of the PPO model trained in the test environment (PPOInTest). The experimental results fully demonstrate the effectiveness of the generalization method proposed in this invention.

[0020] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for rapid adaptation of unloading UAV tasks based on reinforcement learning, involving UAVs, base stations and ground terminal equipment, characterized in that: The drone collects computing tasks generated by the ground terminal equipment, maximizes the number of collected computing tasks by optimizing the flight trajectory of the drone, and maximizes the computing service efficiency provided by the drone and the base station for the computing tasks generated by the ground terminal equipment by optimizing the task offloading decision of the drone, including the following steps: Step 1: Establish a reinforcement learning model based on decoupling strategy and environment representation, and initialize the parameters of the reinforcement learning model based on decoupling strategy and environment representation; Step 2: Use training samples, different training environments, and different training strategies to train a reinforcement learning model based on decoupled strategy and environment representation, obtain environment embedding vectors of different training environments and strategy embedding vectors of different training strategies, and predict the value prediction network reward value obtained by using different training strategies in different training environments; Step 3: Based on the reinforcement learning model of decoupling strategy and environment representation, a strategy adaptation algorithm based on gradient ascent is established to optimize the strategy embedding vector of the drone taking different strategies in the test environment, and the action of the drone under the strategy with the best performance is obtained to interact with the current environment; Step 4: Implement drone-assisted mobile edge computing based on the drone's actions under the best drone performance strategy, so that the drone and the base station can jointly provide edge computing services for the computing tasks generated by the ground terminal equipment.

2. According to claim 1, a method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning is characterized in that: The reinforcement learning model based on the decoupling strategy and environment representation described in step 1 includes several drones and an environment. The several drones interact with the environment and learn iteratively to optimize the actions of the several drones in the environment. The environment includes several ground terminal devices and several base stations. The actions of the several drones in the environment include collecting computing tasks generated by several ground terminal devices in the environment, task offloading decisions, and jointly providing computing services for several ground terminal devices in the environment with several base stations in the environment. The task offloading decision is a decision to offload the computing tasks collected by the drones to several base stations in the environment.

3. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 2 is characterized in that: Step 1 initializes the parameters of the reinforcement learning model based on the decoupling strategy and environment representation, including the state space S of the environment, the reward function R, and the action space A of the drone. The specific method is: Initialize the state space S of the environment based on the observation values ​​of the drone. The observation values ​​of the drone include: the location of the drone and the locations of the three ground terminal devices closest to the drone; initialize the action space A of the drone, including: the movement direction, movement speed and task unloading decision of the drone; initialize the reward function R as the sum of the three target reward functions: the total energy consumption Etotal of the drone, the task delay Ttotal and the number of tasks collected Ntotal; Divide the environment into training environment and test environment, and set the training environment set to , the test environment set is , the training strategy set is ,in, is the number of training environments, is the number of test environments, Create an offline experience buffer for the number of training strategies As training samples, For the Training Environment Use the Training strategies Generated offline experience; Assume that the drone starts from the initial state and continuously interacts with several ground terminal devices and several base stations in the environment. The trajectory of the drone is , used to record the state transition of the drone and the rewards obtained during the state transition, where: is the environmental state, For the drone's movements, Rewards for drones; set in training environment A set of drone trajectories generated in Contextual information for the environment , and set context information Environmental status in and the actions of drones For drone behavior .

4. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 3 is characterized in that: The step 2 comprises: Step 2.1: Build a context encoder By comparing the context information of the drone in different training environments, the environment embedding vectors of different training environments are extracted; Step 2.2: Build a policy encoder and policy decoder , through the policy encoder Extracting training strategies The policy embedding vector , and then through the strategy decoder Prediction Training Strategy The action of dismounting the drone; Step 2.3, build parameters are Value prediction network , using the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value .

5. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 4 is characterized in that: The specific method of step 2.1 is: For any training environment in the training environment set , using the training environment Offline experience Contextual information sampled in As anchor sample and positive samples , from another training environment Offline experience The context information sampled in is used as negative samples , based on anchor samples and positive samples Generate similar sample pairs based on anchor samples and negative samples Generate different sample pairs; Building a context encoder , used to convert the training environment Offline experience Contextual information sampled in Mapping to environment embedding vector , context encoder The loss function for: (1) in, For the context encoder The updated parameter matrix, is the mathematical expectation, Anchor sample The environment embedding vector, , The context encoder Extracted positive samples , negative samples The environment embedding vector, Anchor sample The environment embedding vector The transposed vector of ; Use training samples to train context encoder Train and get the trained context encoder , and based on the context encoder Extract training environment Contextual information The environment embedding vector .

6. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 5 is characterized in that: The specific method of step 2.2 is: Building a policy encoder , for the training strategy The behavior of drones , using the training environment Offline experience The behavior of the drones sampled in The embedding vector of The policy embedding vector ; Building a policy decoder , the environmental state and training strategies The policy embedding vector As a policy decoder Input, prediction training strategy The action of dismounting the drone; By minimizing the l2 loss function Update Strategy Encoder and policy decoder , as shown in the following formula: (2) Get the trained policy encoder and policy decoder , through the trained policy encoder Extracting training strategies The policy embedding vector , and then through the trained strategy decoder Prediction Training Strategy The action of the drone.

7. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 6 is characterized in that: The specific method of step 2.3 is: Value Prediction Network It is a multi-layer nonlinear neural network. , based on the environmental status , Training strategy The policy embedding vector , Training Environment The environment embedding vector As a value prediction network The Monte Carlo method is used to train the value prediction network. , value prediction network The loss function As shown in the following formula: (3) in, , For long-term incentive benefits, γ is the discount factor, is the time step t The reward function is, T is the total number of time steps, let This is the Monte Carlo regression. is the initial environment state; Use the trained value prediction network Predictive training environment Training strategy The value that can be obtained predicts the network reward value .

8. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 7 is characterized in that: The step 3 comprises: Step 3.1: From the training strategy set Sampling any training strategy , the drone uses this training strategy With test environment Interact and collect test environment Contextual information ; Step 3.2: Predict network reward value based on value Establish a gradient ascent-based strategy adaptation algorithm and use the gradient ascent-based strategy adaptation algorithm to optimize the test environment Training strategies for drones The strategy embedding vector is used to obtain the optimized training strategy The policy embedding vector ; Step 3.3: Use the policy decoder Based on the optimized training strategy The policy embedding vector Get the drone's actions under the best strategy for drone performance and interact with the current environment.

9. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 8, characterized in that: The specific method of step 3.2 is: The gradient-based strategy adaptation algorithm sets the number of iterations , each iteration uses the training strategy Sampling a UAV trajectory , For test environment The environmental status, For training strategy The action of the drone, Action for the drone The new environment obtained The environmental state of the drone As an environmental encoder Input to get the test environment The environment embedding vector ; The drone trajectory As a policy encoder Input, get the training strategy The policy embedding vector ; Use the gradient ascent algorithm to follow the value prediction network Predicted value Predicted network reward value With the learning rate η Added direction optimization training strategy The policy embedding vector , as shown in the following formula: (4) Until completion Iterations, reaching the preset number of iterations After that, the optimized training strategy is obtained. The policy embedding vector .

10. The method for rapid adaptation of unloading unmanned aerial vehicle tasks based on reinforcement learning according to claim 9, characterized in that: The step 4 comprises: Step 4.1, generating an optimized UAV flight trajectory and task offloading decision according to the UAV action under the strategy with the best UAV performance; Step 4.2: The drone collects and stores the computing tasks of the ground terminal equipment. According to the optimized flight trajectory of the drone, the number of tasks generated by the ground terminal equipment collected by the drone is maximized and stored in the drone's task queue, completing the collection of computing tasks of the ground terminal equipment. Step 4.3: Determine the allocation result of the computing tasks of the ground terminal device according to the task offloading decision of the UAV, including the number of computing tasks allocated to the UAV and the number of computing tasks offloaded to the base station; Step 4.4, executing the computing tasks according to the computing task allocation result; the UAV performs the computing of the tasks that do not need to be offloaded, and transmits the tasks that need to be offloaded from the UAV to the base station, and performs the computing on the base station; Step 4.5: After the task calculation in the base station and the UAV is completed, the task calculation results in the base station and the UAV are returned to the ground terminal device.

Citation Information

Patent Citations

  • Metareinforcement learning method based on comparative learning and mutual information

    CN114139681A

  • Unmanned aerial vehicle flight decision-making method based on meta-reinforcement learning parallel training algorithm

    CN114895697A

  • Unmanned aerial vehicle auxiliary collaborative task unloading method based on deep reinforcement learning

    CN115580900A

  • Unmanned aerial system path optimization based on machine learning

    US20240427346A1

Cited By

  • Unmanned aerial vehicle communication resource and trajectory optimization method and device based on meta reinforcement learning, equipment and medium

    CN121613917A