Multi-unmanned aerial vehicle cooperation edge calculation method based on reinforcement learning

By adopting reinforcement learning methods in multi-drone collaborative edge computing systems, using genetic algorithms and multi-agent deep deterministic strategy gradient algorithms, the challenges of drone collaboration and resource management are solved, and more efficient energy use and response speed are achieved.

CN120183256APending Publication Date: 2025-06-20BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337880.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Multi-UAV collaborative edge computing methods have challenges in collaborative work, task allocation, path planning and resource management among drones, and require optimization of communication links between drones and ground MEC servers to ensure the stability and security of data transmission.

Method used

Using reinforcement learning-based methods, the learning rate of Actor and Critic modules is pre-trained through genetic algorithms, and the multi-agent depth deterministic strategy gradient algorithm is used to determine the computing task offload of ground users and the flight trajectory of the drone.

Benefits of technology

It improves the energy efficiency and resource management capabilities of the UAV collaborative edge computing system, optimizes the UAV flight trajectory and mission offloading, and enhances the system's flexibility and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183256A_ABST
    Figure CN120183256A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-unmanned aerial vehicle cooperation edge calculation method based on reinforcement learning. Wherein the multiple unmanned aerial vehicles serve as mobile base stations and provide calculation services for ground users; presetting the period frequency and the storage size of a central processing unit for providing calculation service on each unmanned aerial vehicle; taking task calculation service delay, energy consumption and unmanned aerial vehicle track weighted sum minimization as optimization targets; carrying out pre-training by adopting a genetic algorithm, and taking a fitness function of the genetic algorithm as a reward function of a multi-agent depth deterministic strategy gradient algorithm; determining an unloading strategy and a track of each unmanned aerial vehicle by adopting the algorithm; and a multi-intelligent depth deterministic strategy gradient algorithm is adopted to determine a delivery unmanned aerial vehicle of each task calculation result, and the result is fed back to a user. According to the method, the flight path of each unmanned aerial vehicle, the user service quality and the energy efficiency of the whole system are fully considered, timely delivery of tasks is ensured, and the whole energy consumption of the system is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi - UAV collaborative edge computing method based on Reinforcement Learning (RL), belonging to the field of edge computing. Background Art

[0002] With the proliferation of Internet of Things (IoT) devices and the popularization of 5G technology, we are stepping into an era of data explosion. The massive data generated by these devices requires powerful computing capabilities to process, especially for those computationally intensive applications such as autonomous driving, intelligent monitoring, and augmented reality. However, due to its centralized architecture, the traditional cloud computing model faces problems such as high network latency, large bandwidth consumption, and slow response speed, which limit the quality of service and user experience.

[0003] In such a technical background, Mobile Edge Computing (MEC) emerges as the times require. It pushes computing capabilities from the cloud to the network edge, closer to the data source and users. By deploying servers at the network edge, MEC reduces data transmission latency, improves service response speed, and alleviates the burden on the central cloud server. In addition, MEC can also provide lower energy consumption and higher data security.

[0004] However, the MEC servers deployed on the ground have the problem of insufficient flexibility, especially in scenarios that require quick response and dynamic deployment. To solve this problem, Unmanned Aerial Vehicle (UAV) technology is introduced into MEC, forming an edge computing method based on multi - UAV collaboration. UAVs have the capabilities of high mobility and rapid deployment, and can move flexibly in the air to provide instant computing and network services for specific areas.

[0005] The multi - UAV collaborative edge computing method combines the computing advantages of MEC and the flexibility of UAVs, and can provide on - demand computing resources for mobile users and temporary events. For example, in the event of a natural disaster, UAVs can quickly fly to the affected area to provide communication restoration and data processing services. In large - scale events, UAVs can act as temporary MEC nodes to provide additional network capacity and computing support.

[0006] In terms of technical challenges, multi - UAV collaborative edge computing needs to solve the problem of collaborative work among UAVs, including task allocation, path planning, and resource management. At the same time, issues such as the energy limitation of UAVs, flight safety, and compliance with regulations also need to be considered. In addition, the communication link between UAVs and ground MEC servers also needs to be optimized to ensure the stability and security of data transmission.

[0007] In summary, the edge computing method based on multi-UAV collaboration provides a new idea for solving the latency and bandwidth problems in mobile computing, and its development will have a profound impact on fields such as the Internet of Things, 5G communication, and smart cities. With the continuous progress and improvement of related technologies, we are expected to see more intelligent, flexible, and efficient computing services in the future. Summary of the Invention

[0008] Aiming at the deficiencies of the above existing technologies, the present invention provides a multi-UAV collaborative edge computing method based on reinforcement learning, including: pre-training the learning rates of the Actor module and the Critic module by a genetic algorithm; determining the computing task offloading of ground users and the flight trajectories of UAVs by a multi-agent deep deterministic policy gradient algorithm.

[0009] The object of the present invention is achieved through the following technical solutions:

[0010] A multi-UAV collaborative edge computing method based on reinforcement learning, the method includes the following steps:

[0011] 1) Construct a multi-UAV collaborative edge computing environment;

[0012] 2) Initialize the genetic algorithm model and pre-train it;

[0013] 3) On the basis of 2), set the reward function and learning rate of the multi-agent deep deterministic gradient policy algorithm;

[0014] 4) On the basis of 3), sample the multi-agent deep deterministic policy gradient algorithm according to the positions and computing tasks of ground users to determine the flight trajectories and offloading tasks of each UAV.

[0015] 5) On the basis of 4), each ground user offloads the computing task to the corresponding UAV and starts the computing service. Description of the Drawings

[0016] Figure 1 Schematic diagram of a multi-UAV collaborative edge computing method based on reinforcement learning;

[0017] Figure 2 UAV-assisted edge computing architecture diagram;

[0018] Figure 3 Schematic diagram of a multi-agent deep deterministic gradient policy algorithm model integrating a genetic algorithm. Detailed Embodiment

[0019] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. The following description covers many specific details in order to provide a comprehensive understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention can be implemented without some of these specific details. The following description of the embodiments is only to provide a clearer understanding of the present invention by showing examples of the present invention. The present invention is in no way limited to any specific configuration and algorithm proposed below, but covers any modifications, substitutions, and improvements of related elements, components, and algorithms without departing from the spirit of the present invention.

[0020] The following will refer to the attached Figure 1 The specific steps of an embodiment of a multi-UAV collaborative edge computing method based on reinforcement learning according to the present invention are as follows:

[0021] The first step is to construct a multi-UAV collaborative edge computing environment.

[0022] As Figure 2 shown, according to information such as the number of ground users, the number of UAVs, the computing power of UAVs, the UAV channel bandwidth, the maximum transmission power of ground users, and the number of tasks, a multi-UAV collaborative edge computing environment is constructed. During the subsequent training process, the agent and the environment continuously interact and update the network parameters.

[0023] The second step is to initialize the genetic algorithm model and pre-train it, and set the reward function and learning rate of the multi-agent deep deterministic gradient policy algorithm.

[0024] First, set the fitness function of the genetic algorithm as the reward function of the MADDPG algorithm. Then initialize the chromosomes and evaluate them using the fitness function. Next, update the chromosomes through selection, crossover, and mutation operations, retain the chromosomes with higher fitness, and continuously iterate and train until the learning rate with the highest fitness value is obtained. The specific steps are as follows:

[0025] 1. Initialize the number of chromosomes, crossover probability, mutation probability, and number of iterations iter max ;

[0026] 2. Initialize the environment, experience replay buffer, system parameters, number of training episodes ep max , and network data of ground users, etc.;

[0027] 3. Set the fitness function fitness = Cost, evaluate each individual in the population, and determine its fitness;

[0028] 4. Selection, according to the fitness of the individuals, select the individuals with higher fitness to enter the next generation;

[0029] 5. Cross, randomly pair the selected individuals, and generate new offspring through the crossover operation;

[0030] 6. Mutation, randomly change some genes of certain offspring individuals with a certain mutation probability to increase the population diversity;

[0031] 7. Generate a new generation of population, ensuring that the best individuals directly enter the next generation;

[0032] 8. Check whether the termination conditions are met, reaching the maximum number of iterations and the quality of the solution meeting the preset threshold. If the termination conditions are not met, return to step 4. If the conditions are met, output the optimal learning rates LR_A and LR_C.

[0033] The third step, the multi-agent deep deterministic policy gradient algorithm determines the flight trajectories and offloading tasks of each UAV.

[0034] As Figure 3 shown, a multi-agent deep deterministic policy gradient algorithm that incorporates a genetic algorithm is provided to determine the flight trajectories and offloading tasks of each UAV. The design of the algorithm model includes the following steps:

[0035] 1. State space design. The system state space reflects the situation observed from the constructed system environment, including the states of ground users and the available resources of UAVs. Therefore, the state space s(t) at time slot t can be expressed by the formula:

[0036] s(t) = {c1(t), c2(t), ···, c d (t), d1(t), d2(t) ··· d d (t), f1(t), f2(t) ···, f n (t), l1(t), l2(t) ··· l n (t)}. Among them, c d (t) represents the number of CPU cycles required to complete the request task of the d-th ground user at time t, d d (t) represents the data volume of the request task at time t, f n (t) represents the available resources of the UAV base station at time t, l n (t) represents the location of the UAV at time t, which can be further expressed as l n (t) = (x t , y t ).

[0037] 2. Action Space Design. Based on the observed environmental state space, the agent decides which drone the ground user should choose for offloading, how much computing resource and channel resource should be allocated to the ground user, as well as the next speed and direction of the drone. Therefore, the action space a(t) of the agent at time slot t can be expressed by the formula:

[0038] a(t) = {x1(t), x2(t), ···, x d (t), b1(t), b2(t) ··· b f (t), e1(t), e2(t) ···, e f (t)}

[0039] where, x d (t) represents the offloading strategy of ground user d at time t, b f (t) represents the available resources of drone f at time t, and e f (t) represents the available bandwidth of drone f at time t.

[0040] 3. Reward Function Design. The reward function comes from the fitness of the genetic algorithm. In order to minimize the total energy consumption of computing offloading, resource allocation, and drone flight, the design of the reward function should be able to guide each drone agent to conduct autonomous learning and minimize the cost, so as to optimize resource allocation and drone flight trajectory. The reward R is expressed as:

[0041]

[0042] where, represents the delay and quality of service of the computing task transmitted from the ground user to the drone base station, represents the energy consumed by the flight of the drone.

[0043] After the algorithm model design is completed, this paper uses the multi-agent deep deterministic policy gradient algorithm integrated with the genetic algorithm to determine the flight trajectories and offloading tasks of each drone:

[0044] 1. First, define the reward function of the MADDPG algorithm as the fitness evaluation criterion of the genetic algorithm. Second, initialize the chromosomes in the population and evaluate them according to the fitness function. And optimize the chromosomes by implementing genetic operations such as selection, crossover, and mutation, while retaining those chromosomes with higher fitness. This process will be repeated continuously, and through iterative training, until the parameters with the optimal fitness value, that is, the learning rate, are found.

[0045] 2. Set the learning rate parameter of the MADDPG algorithm to the result pre-trained by the genetic algorithm, and start initializing the multi-UAV collaborative edge computing environment, the experience replay buffer, and the parameters related to the network. At the same time, initialize the Actor and Critic models to prepare for subsequent deep reinforcement learning training.

[0046] 3. Each agent continuously obtains information from the multi-UAV collaborative edge computing environment to learn the optimal policy. This information covers the current state of the ground users and the computing and bandwidth resources provided by the UAV base stations. Based on these learned policies, each agent converts the current environmental state into corresponding actions, and then makes decisions on resource allocation and task offloading. In the early stage of training, to enhance the exploration ability of different actions, Gaussian noise is introduced when the agent selects actions. The selection of actions can be expressed by the following formula:

[0047] a(t) = μ(s t ) + ε(t)

[0048] where μ(s t ) is the policy action taken by the agent in state s t , and ε(t) is the Gaussian noise added at this moment.

[0049] 4. After selecting action a(t), each agent can calculate the current reward, send it into the multi-UAV collaborative edge computing environment, and return the next state s(t + 1). In each training cycle, each agent obtains the quadruple data, namely {s(t), a(t), R(t), s(t + 1)}, from the experience replay pool as the training set for the Actor and Critic modules. To ensure the real-time performance of the network model, the Critic module continuously updates the parameters by reducing the loss function, and the loss function can be expressed as the formula:

[0050] Loss = E[(p(t) - Q(s(t), a(t)|θ)) 2

[0051] where Q(s(t), a(t)|θ) represents the action value function, which is used to evaluate the value of the policy. The action value function needs to be calculated for each action through s(t) and s(t + 1). p(t) represents the target value of the current action, which is obtained through the parameters of the Critic module. According to Q(s(t), a(t)|θ) and the experience replay pool, the Actor module updates θ μ .

[0052] ​5. The above update only involves the Actor module and the Critic module, and their corresponding target networks have not been updated. The target network is updated by using a soft update strategy to update the network parameters in the Actor and Critic modules. Finally, the best network parameters are obtained, so as to give the best flight trajectories and offloading tasks of each drone under the corresponding conditions. The multi-agent deep deterministic policy gradient algorithm integrated with the genetic algorithm can be used to optimize the flight trajectories and offloading tasks of drones in a drone-assisted mobile edge computing system to improve the energy efficiency of the system.

[0053] The present invention relates to a multi-UAV collaborative edge computing method based on reinforcement learning proposed above. It should be understood that the detailed description of the technical solutions of the present invention with the aid of the preferred embodiments is illustrative rather than restrictive. Those of ordinary skill in the art can modify the technical solutions recorded in each embodiment or perform equivalent replacements on some of the technical features based on reading the specification of the present invention; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A multi-UAV collaborative edge computing method based on reinforcement learning, characterized in that: The method comprises the following steps: 1) Build a multi-UAV collaborative edge computing environment; 2) Establish a communication link between the drone base station and the ground user; 3) Initialize the CPU cycle frequency and storage of the drone computing server; 4) Initialize the genetic algorithm model and pre-train; 5) Setting the reward function and learning rate of the multi-agent deep deterministic gradient strategy algorithm based on genetic algorithm pre-training; 6) A multi-agent deep deterministic policy gradient algorithm is used to determine the flight trajectory and offload tasks of each UAV based on the location and computing tasks of the ground user; 7) Each ground user offloads the computing task to the corresponding UAV and starts computing service; 8) Each UAV delivers the results of the computing task to ground users.

2. The method according to claim 1, characterized in that The genetic algorithm model pre-training includes: setting the fitness function of the genetic algorithm to the reward function of the multi-agent deep deterministic gradient strategy algorithm, initializing the chromosomes and updating the fitness of each chromosome, updating the chromosomes through selection, crossover and mutation operations, retaining chromosomes with higher fitness and iterating, and finally obtaining a learning rate with higher fitness.

3. The method according to claim 1, characterized in that The multi-agent deep deterministic gradient strategy algorithm includes: initializing a multi-UAV collaborative edge computing environment, an experience playback area, and an Actor and Critic network model in the multi-agent.

4. The method according to claim 1 or claim 3, characterized in that: The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is used to determine the UAV flight trajectory and offloading tasks, including: the agent learns strategies through continuous interaction with the multi-UAV collaborative edge computing environment, where the feedback information of the environment includes the status of ground users and the available computing resources of the UAV base station, and then formulates resource allocation plans, offloading strategies and UAV flight trajectories based on the learned strategies.

5. The method according to claim 4, characterized in that According to the multi-agent deep deterministic policy gradient algorithm, it includes: each agent has four network structures, two of which are evaluation networks and two are target networks; each agent has its own Actor network, which is used to determine the action a that should be taken under a given state s; each agent has its own Critic network, which is used to evaluate the value of actions under the current strategy; the MADDPG algorithm uses a target network, namely the target network of the Actor and the Critic; these target networks are delayed replicas of the original networks, and they are used to calculate the target value function to reduce the variance during training.

6. The method according to claim 4, characterized in that The multi-agent deep deterministic policy gradient algorithm includes: the state space of the deep deterministic policy gradient algorithm is the state of the ground user and the available resources of the drone base station; based on the observed environmental state space, the agent decides which drone the ground user should choose for unloading and how much computing resources and channel resources should be allocated to the ground user.

7. The method according to claim 4, characterized in that The multi-agent deep deterministic policy gradient algorithm includes: the goal of the reward function is to minimize the total system cost of computational offloading and resource allocation; therefore, the design of the reward function should be able to guide the agent to learn autonomously and take actions that minimize the total system cost.