Multi-unmanned aerial vehicle cooperative computing unloading and resource allocation method and system for sudden disaster emergency search and rescue
Through the collaborative computing and offloading and resource allocation method of multiple drones, combined with the deep reinforcement learning model and IMATD3 algorithm, the drone trajectory planning and search and rescue equipment task offloading decisions are optimized, and the problems of limited energy consumption and resource competition in emergencies and emergency disaster scenarios are solved, and efficient and fair search and rescue services and load balancing are achieved.
Patent Information
- Application Number
- CN202510661851.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In emergencies and emergency disaster scenarios, the onboard energy consumption of drones is limited, making it difficult to provide computing services for ground search and rescue equipment for a long-term and efficient manner. At the same time, under the competition for multiple search and rescue equipment resources, service fairness and drone load balancing are difficult to ensure.
The collaborative computing offloading and resource allocation method of multiple drones is adopted. By building a multi-drone-assisted three-dimensional computing offloading network architecture for emergencies, combining deep reinforcement learning models and improved multi-agent dual-delay deep deterministic strategy gradient algorithm (IMATD3), the drone trajectory planning and search and rescue equipment task offloading decisions are optimized, and the system utility function is maximized.
It effectively improves the execution efficiency of the collaborative rescue mission of multiple drones, ensures the fairness of search and rescue equipment services, and optimizes the energy consumption and load balancing of drones.
Smart Images

Figure CN120179322A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computing offloading technology, and in particular to an unmanned aerial vehicle (UAV)-assisted computing offloading and resource allocation. More specifically, it relates to a method and system for multi-UAV collaborative computing offloading and resource allocation for emergency disaster rescue scenarios Background Art
[0002] With the rapid development of wireless communication technology and mobile computing, UAVs have been increasingly widely used in fields such as emergency communication, broadband connection, and environmental monitoring. Especially in areas without Internet infrastructure or where the communication network is damaged, the mobile edge computing (MEC) system carried by UAVs can provide computing offloading services for ground search and rescue equipment, thus meeting the requirements of real-time and low latency. UAVs can be equipped with MEC servers to build a flexible mobile sensing platform. This configuration can not only ensure that search and rescue equipment obtains a high-quality service experience, but also flexibly adapt to various actual application scenarios. In complex scenarios with multiple search and rescue devices, the task requirements of search and rescue devices change dynamically, and the allocation of computing resources and task offloading decisions become particularly critical.
[0003] Currently, the MEC system carried by UAVs has been widely used for broadband connection and emergency communication in areas without Internet. However, the on-board energy consumption of UAVs is limited. How to provide long-term and efficient computing services for search and rescue equipment under limited energy consumption constraints has become a hot research issue. In addition, the task requirements of search and rescue equipment change dynamically. How to ensure the fairness of search and rescue equipment services and the load balance of UAVs under the competition of multiple search and rescue equipment resources are also current research difficulties. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to propose a method and system for multi-UAV collaborative computing offloading and resource allocation for emergency disaster rescue. For the scenario of multi-UAV collaborative search and rescue in sudden natural disasters (such as earthquakes), the optimization goal is to maximize the system utility function that takes into account the fairness of offloading for search and rescue equipment, the load balance of UAVs, and energy consumption, achieving a good balance between complexity and performance in the face of disaster environments. UAVs can dynamically adjust resource allocation according to the requirements of ground rescue tasks to provide the optimal quality of service.
[0005] Technical Solution: To achieve the above object of the invention, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for multi-UAV collaborative computing offloading and resource allocation for emergency disaster rescue, including the following steps: Build a multi-UAV assisted three-dimensional computing offloading network architecture for emergency disaster search and rescue scenarios, where the UAVs provide computing services for search and rescue equipment, and the search and rescue equipment moves within the disaster area and generates computing tasks; According to the task offloading decision variables of the search and rescue equipment for each UAV, determine the offloading ratio parameters of the search and rescue equipment in each time slot, and then measure the task offloading fairness among the search and rescue equipment; according to the resource allocation decision variables of the UAVs for each search and rescue equipment, determine the resource allocation ratio parameters of the UAVs, and then measure the load balancing degree among the UAVs; taking into account the offloading fairness of the search and rescue equipment, the load balancing of the UAVs and reducing the energy consumption of the UAVs, establish a system utility function, and establish a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that decides task offloading; Train the deep reinforcement learning model based on the improved Multi-Agent Twin Delayed Deep Deterministic Policy Gradients (IMATD3), introduce the Ornstein-Uhlenbeck (OU) noise and the priority cloning learning mechanism, and learn and optimize the trajectory planning of the UAVs and the task offloading decisions of the search and rescue equipment.
[0006] Furthermore, the system includes multiple rescue devices and multiple UAVs. The running time of the system includes multiple consecutive time slots, and the duration of each time slot can be dynamically adjusted according to the urgency of the rescue task; the movement of the UAVs in the system cannot exceed the boundary of the target disaster area, the flight altitude does not exceed the upper limit, the distance flown by the UAVs between adjacent time slots does not exceed the maximum distance, and the distance between the UAVs is not less than the safety distance.
[0007] Furthermore, the task offloading fairness among the search and rescue equipment is expressed as follows: ; where is the offloading ratio parameter of the search and rescue equipment n in the time slot t , N is the number of rescue devices, U is the number of UAVs, is the time slot t the search and rescue equipment n for the UAV u task offloading decision variable; the load balancing degree among the UAVs is expressed as follows: ; where is the resource allocation ratio parameter of the UAV u in the time slot t , Time slot t Unmanned aerial vehicle (UAV) u For search and rescue equipment n Resource allocation decision variable
[0008] Furthermore, in the deep reinforcement learning model, each UAV represents an agent. The observation space of the UAV includes the coordinates of all search and rescue equipment and all UAVs, the coordinates of the target point, the remaining energy consumption of the UAV, as well as the task offloading situation of the search and rescue equipment and the load situation of the UAV; the action space of the UAV includes the flight angle, flight distance, range coverage, and resource allocation
[0009] Furthermore, in the deep reinforcement learning model, the rewards include remaining energy consumption and fairness rewards, arrival rewards, and wrong action penalties; in time slot t The remaining energy consumption and fairness rewards Are defined as Where Is the task offloading fairness among search and rescue equipment Is the load balancing degree among UAVs And Are weights Is the remaining energy consumption of the UAV; the arrival reward Is expressed as Where And Are the UAV and the target position respectively Represents the Euclidean distance M 1 is a constant greater than 0 M 2 is a constant used to prevent the denominator from being 0; when the UAV flies out of the target disaster area, fails to complete the task within the time limit, runs out of energy, or collides, a wrong action penalty is obtained
[0010] Furthermore, when training the deep reinforcement learning model based on the improved multi-agent double delayed deep deterministic policy gradient algorithm, for each UAV, its partial observation state is obtained. While selecting actions based on the greedy policy and the policy network, the OU noise mechanism is used to add noise perturbations to the action selection. The wrong action penalty is calculated according to the selected action and the environmental state, and it is judged whether the UAV collides or flies out of the target disaster area; if a collision or flying out of the target disaster area occurs, the UAV returns to the initial position; the partial observation state of the next moment is obtained. For each search and rescue equipment, it is selected based on the minimum energy consumption greedy policy. For each UAV, the reward value is calculated; the global state of the next moment is obtained, and the experience of the current moment is stored in the experience pool
[0011] Furthermore, when training the deep reinforcement learning model based on the improved multi-agent double-delayed deep deterministic policy gradient algorithm, for each drone, a batch of samples is taken from the experience pool according to the priority, and the loss functions of the two value networks are calculated. Then, the value network parameters are updated, the policy network parameters are updated based on the deterministic policy gradient ascent method, and the target network parameters are adjusted. Among them, the cloning learning mechanism is adopted to clone some parameters of the current policy network to the target network.
[0012] In a second aspect, the present invention provides a multi-drone collaborative computing offloading and resource allocation system for emergency search and rescue in the face of sudden disasters, including: A scenario construction module for building a multi-drone assisted three-dimensional computing offloading network architecture for emergency search and rescue scenarios in the face of sudden disasters, where the drones provide computing services for the search and rescue equipment, and the search and rescue equipment moves in the disaster area and generates computing tasks. A problem modeling module for determining the offloading ratio parameters of the search and rescue equipment in each time slot according to the task offloading decision variables of the search and rescue equipment for each drone, thereby measuring the fairness of task offloading among the search and rescue equipment; determining the resource allocation ratio parameters of the drones according to the resource allocation decision variables of the drones for each search and rescue equipment, thereby measuring the load balancing degree among the drones; taking into account the offloading fairness of the search and rescue equipment, the load balancing of the drones, and reducing the energy consumption of the drones, establishing a system utility function, and establishing a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that determines task offloading. A reinforcement learning module for training the deep reinforcement learning model based on the improved multi-agent double-delayed deep deterministic policy gradient algorithm, introducing OU noise and the priority cloning learning mechanism, and learning and optimizing the trajectory planning of the drones and the task offloading decisions of the search and rescue equipment.
[0013] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters are implemented.
[0014] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters are implemented.
[0015] Beneficial effects: In view of the problems that the airborne energy consumption of drones is limited in sudden emergency disaster scenarios and it is difficult to ensure the fairness of services and the load balance of drones under the competition of multiple search and rescue equipment resources, a multi-drone collaborative computing offloading and resource allocation method is proposed. Considering the influence of factors such as the flight direction, altitude, and distance of drones, the system utility function is maximized by controlling the trajectories of drones and the task offloading decisions of search and rescue equipment. Among them, the system utility function takes into account the offloading fairness of search and rescue equipment, the load balance of drones, and the energy consumption of drones. An algorithm based on IMATD3 is used to solve the path planning problem of drones. The optimal strategy is dynamically learned through the interaction between the policy network and the environment, and efficient decision-making is carried out in the continuous action space. An OU noise mechanism is introduced into the IMATD3 algorithm to enhance the exploration ability of agents in the dynamic disaster environment, helping drones to better adapt and make decisions in the complex and changeable disaster environment. At the same time, combined with the priority cloning learning mechanism, high-value experience samples are learned. By cloning some policy network parameters to the target network, the learning process is accelerated and the stability and convergence speed of the algorithm are improved.
[0016] In summary, the present invention can effectively improve the execution efficiency of multi-drone collaborative rescue tasks, ensure the service fairness of search and rescue equipment, and at the same time optimize the energy consumption and load balance of drones. Brief Description of the Drawings
[0017] Figure 1 Scenario schematic diagram of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters in the embodiments of the present invention.
[0018] Figure 2 Detailed flowchart of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters in the embodiments of the present invention.
[0019] Figure 3 Reward result graph of the IMATD3 algorithm and the comparison algorithm in the embodiments of the present invention under different task amounts.
[0020] Figure 4 Reward result graph of the IMATD3 algorithm and the comparison algorithm in the embodiments of the present invention under different maximum computing resources of drones.
[0021] Figure 5 Fairness comparison result graph of search and rescue equipment between the IMATD3 algorithm and the comparison algorithm in the embodiments of the present invention under different time slots.
[0022] Figure 6 Load balance comparison result graph of drones between the IMATD3 algorithm and the comparison algorithm in the embodiments of the present invention under different time slots. Detailed Embodiments
[0023] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] In Figure 1 , a scenario of multi-UAV collaborative computing offloading and resource allocation for emergency search and rescue in the face of sudden disasters is described. It can be seen that multiple UAVs act as MEC servers to provide services for search and rescue equipment within their coverage areas. Combining Figure 1 , an embodiment of the present invention discloses a method for multi-UAV collaborative computing offloading and resource allocation for emergency search and rescue in the face of sudden disasters. First, a multi-UAV assisted three-dimensional computing offloading network architecture for emergency search and rescue scenarios is built, where the UAVs provide computing services for search and rescue equipment, and the search and rescue equipment moves within the disaster area and generates computing tasks. Then, according to the task offloading decision variables of the search and rescue equipment for each UAV, the offloading ratio parameters of the search and rescue equipment in each time slot are determined, so as to measure the fairness of task offloading among the search and rescue equipment. According to the resource allocation decision variables of the UAVs for each search and rescue equipment, the resource allocation ratio parameters of the UAVs are determined, so as to measure the load balancing degree among the UAVs. Taking into account the offloading fairness of the search and rescue equipment, the load balancing of the UAVs and the reduction of UAV energy consumption, a system utility function is established, and a deep reinforcement learning model is established for the search and rescue equipment that determines task offloading with the goal of maximizing the utility function. Finally, based on the improved multi-agent double delayed deep deterministic policy gradient algorithm, the deep reinforcement learning model is trained, and the OU noise and priority cloning learning mechanism are introduced to learn and optimize the trajectory planning of the UAVs and the task offloading decisions of the search and rescue equipment.
[0025] The embodiment of the present invention aims at the problem of collaborative computing offloading of UAVs as MEC servers in emergency search and rescue scenarios for sudden disasters, and solves the problem of how to provide long-term computing services for ground search and rescue equipment under limited onboard energy consumption by optimizing the flight trajectories of UAVs and the task offloading decisions of search and rescue equipment. The embodiment of the present invention takes into account the offloading fairness of the search and rescue equipment, the load balancing of the UAVs and the energy consumption optimization to maximize the system utility function. The improved multi-agent double delayed deep deterministic policy gradient (IMATD3) algorithm is adopted, and through the mode of centralized training and distributed execution, it dynamically adapts to task requirements and resource competition. The OU noise mechanism is used to enhance the exploration ability of the agents of the UAVs in the complex and changeable disaster environment and avoid local optimal solutions. At the same time, combined with the priority cloning learning mechanism, high-value experience samples are learned, and through the cloning of some policy network parameters, the learning process is accelerated and the stability and convergence speed of the algorithm are improved. This method significantly improves the service fairness and load balancing of the search and rescue equipment, and at the same time effectively reduces the energy consumption, and can better meet the complex requirements in the dynamic scenario of multiple search and rescue equipment.
[0026] InFigure 2 Among them, a detailed process of a multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters is described. Through interaction with the environment, it dynamically learns to control the path planning of UAVs and the task offloading decisions of search and rescue equipment.
[0027] Next, in combination with Figure 2 A multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters will be further described in detail. Specifically, the multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters includes the following steps: Step (1), build a multi-UAV assisted three-dimensional computing offloading network architecture for emergency search and rescue scenarios in the face of sudden emergencies. Multiple UAVs act as MEC servers to provide services for search and rescue equipment within their coverage areas. The system can adopt directional antenna technology and OFDM technology to effectively overcome the communication challenges brought by the complex electromagnetic environment in the disaster area. The UAVs deployed in the system are all equipped with computing modules and can provide computing services for search and rescue equipment through the MEC servers carried. The search and rescue equipment moves in the disaster area and generates computing tasks. Part of the computing tasks can be processed locally and part can be offloaded to the UAV edge servers.
[0028] Step (2), establish communication, delay, and energy consumption models including N search and rescue equipment and U UAVs, and then establish a joint computing migration and resource allocation model, specifically including: (2a), establish a communication model for three-dimensional computing offloading of multi-UAV assisted search and rescue. The system includes N key rescue equipment such as U emergency communication terminals and UAVs. The search and rescue equipment and UAVs are represented by sets and respectively. The running time of the system includes T consecutive time slots, which can be represented by the set . The duration of each time slot is , which can be dynamically adjusted according to the urgency of the rescue task. The position of each UAV can be represented by , where D represents the boundary range of the UAV on the and coordinates when the UAV moves in the target disaster area, ensuring that the movement of the UAV does not exceed the boundary of the target disaster area. In addition, , H max represents the maximum value of the UAV flight altitude, ensuring that the flight altitude of the UAV is within the upper limit. The maximum distance that the UAV flies between adjacent time slots is d max. To avoid collisions between drones, the drones u and the drone satisfy the distance between , d min represents the safety distance between drones.
[0029] (2b), the communication channel between the search and rescue equipment and the drone can be modeled as a combination related to large-scale fading and small-scale fading. The search and rescue equipment n and the drone u The channel gain can be expressed as ; among them, represents the distance between the search and rescue equipment n and the drone u, represents the channel gain when the reference distance is 1m; l is the path loss exponent, representing the rate at which the signal decays with distance during propagation; represents the small-scale fading coefficient, and the small-scale fading is usually modeled as Rice fading. The search and rescue equipment n and the drone u The link signal-to-noise ratio (SNR) can be expressed as ; among them, represents the search and rescue equipment n The transmit power of the signal transmitted to the drone u , represents the background noise of the communication link, which is the Gaussian noise power. For the search and rescue equipment, the actual data transmission rate of its data can be expressed as ; among them, B represents the channel bandwidth of the link. The average transmission rate can be expressed as ; among them, , representing the ratio of the channel gain to the noise power; is the fading factor obtained by approximating with logistic regression.
[0030] (2c), establish a delay model, and implement a partial offloading mechanism for tasks. Part of them can be executed on the search and rescue equipment, and the other part can be offloaded to the drone for calculation, and the two parts are processed in parallel. The computing nodes are represented by the set , and each offloading node is represented by o . When o = 0, it means that the task is executed locally. When , it means that the task is offloaded to the MEC server for processing. The decision variable of the search and rescue equipment can be expressed as , 。Each search and rescue device has a corresponding rescue task that needs to be processed and calculated, which is represented by wherein, represents the time slot t search and rescue device n data volume size, represents the number of CPU rotations required to calculate this task. When and it means that the task is partially offloaded; and it means that the task is completely offloaded to the drone for processing; and o =0 means that the task is processed locally; otherwise, .
[0031] When the task of the search and rescue device is offloaded to the corresponding drone, the transmission delay represents the time for the search and rescue device to transmit part of the task to the drone, and the calculation delay represents the time required for the drone to process the task of the search and rescue device, which can be expressed as ; wherein, represents the time slot t search and rescue device n to the drone u decision variable, represents the time slot t drone u to the search and rescue device n resource allocation decision variable.
[0032] Each task offloading calculation needs to be completed within a time slot , so the time of the offloaded part can be expressed as ; The local calculation time of the task can be expressed as ; wherein, represents the local computing resources of the search and rescue device. Similarly, the local computing part also needs to be completed within a time slot, that is .
[0033] Therefore, the completion delay of each search and rescue device task needs to satisfy .
[0034] (2d), establish an energy consumption model, and the computing resources provided by the drone for the search and rescue device can be expressed as ; wherein, , f max represents the maximum computing resources provided by the UAV for search and rescue equipment. The computing energy consumption of the UAV can be expressed as ; where is a coefficient related to the UAV hardware, . The propulsion power of the UAV during flight can be expressed as ; where represents the flight speed of the UAV, represents the tip speed of the UAV rotor, K from 1 to K 5 are constants related to the aerodynamics and weight of the UAV. The flight energy consumption of the UAV can be expressed as .
[0035] The total energy consumption of each UAV in each time slot can be expressed as .
[0036] The remaining energy consumption of all UAVs in time slot t can be expressed as ; where represents the remaining energy consumption of the UAV in time slot t, represents the initial energy consumption of each UAV.
[0037] In order to enable each search and rescue equipment to enjoy the offloading service fairly and avoid some search and rescue equipment being in the task offloading state for a long time while other search and rescue equipment cannot be served, the offloading ratio parameter of the search and rescue equipment in each time slot is defined as , which can be expressed as .
[0038] Construct the offloading balance function of each search and rescue equipment, which is used to measure the balance degree of offloading tasks between equipment and can be expressed as .
[0039] Define the resource allocation ratio parameter of the UAV, which is used to measure the share of computing resources allocated by each UAV for search and rescue equipment during the whole task process and is used as a dynamically adjustable index during the task allocation process and can be expressed as .
[0040] To measure the load balance between UAVs, construct the load balance function of each UAV, which can be expressed as .
[0041] In order to balance the offloading fairness of search and rescue equipment, the load balance of UAVs and reduce the energy consumption of UAVs, establish the system utility function, which can be expressed as ; among which, and represent the weight parameters of the unloading fairness of the search and rescue equipment and the load balance of the UAVs, satisfying .
[0042] (2e), in summary, the following objective function and constraint conditions can be established: ; Among which, represents the movement of the UAV, represents the task unloading decision. The objective function is the defined system utility function. The constraint condition C1 ensures that the flight range of the UAV is limited by the three-dimensional space boundary of the target disaster area H min represents the minimum safe altitude of the UAV flight; the constraint condition C2 ensures that no collision occurs between the UAVs; the constraint conditions C3 and C4 are the coverage constraints between the UAVs and the search and rescue equipment, and the position of each search and rescue equipment on the ground can be represented by represented, among which, , represents the angle between the horizontal plane where the UAV is located and the search and rescue equipment; the constraint condition C5 represents the signal-to-noise ratio constraint, ensuring the communication quality between the search and rescue equipment and the UAV, represents the minimum SNR to ensure the transmission of key data such as vital signs; the constraint condition C6 stipulates that the tasks of the search and rescue equipment need to be completed within a time slot range; the constraint condition C7 limits the computing resources of the UAV; the constraint condition C8 represents that the remaining energy consumption of the UAV needs to be less than the original energy consumption.
[0043] Step (3), each search and rescue equipment obtains the UAV position, computing resource occupancy, and task information, as the basis for establishing the system utility function.
[0044] Step (4), considering the load balance between the UAVs, taking into account the unloading fairness of the search and rescue equipment, the load balance of the UAVs, and reducing the energy consumption of the UAVs, establish a system utility function, and establish a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that determines task unloading, including the following specific steps: (4a), based on the decision-making of the UAV trajectory planning and the task unloading of the search and rescue equipment in mobile edge computing, construct a multi-agent Markov decision process, and each UAV represents an agent. The partial observation space of each UAV consists of the coordinates of all search and rescue equipment, UAVs, and target points in the system, as well as the remaining energy consumption of the UAV, and the observation space of the UAV can be expressed as , among which, represents the target point position, represents the remaining energy consumption of the UAV from the start to time slot t, .
[0045] (4b), incorporate the task offloading situation of the search and rescue equipment and the load situation of the UAV into the state space of the system , which can be expressed as .
[0046] The action space of each UAV includes flight angle, flight distance, coverage, and resource allocation. Among them, the flight angle includes the horizontal direction angle and the vertical direction angle. Normalize each action, and the action space can be expressed as ; where represents the normalized parameter of the vertical direction angle of the UAV; represents the normalized parameter of the horizontal direction angle; represents the normalized parameter of the flight distance; represents the normalized coverage parameter, which can be calculated as ; where represents the normalized resource allocation parameter, which can be calculated as .
[0047] (4c), the reward function is divided into three parts, namely the remaining energy consumption and fairness reward, arrival reward, and wrong action penalty. For a single UAV, more attention is paid to the remaining energy consumption in each time slot, which helps to observe the state of the UAV in each time slot. The remaining energy consumption and fairness reward is defined as .
[0048] Designing a suitable arrival reward for each UAV can ensure that the UAV can reach the specified target area. Since the UAV may run continuously for dozens to hundreds of time slots during flight, and in actual operation, the UAV can only receive signals intermittently and judge whether it reaches the destination. Therefore, its arrival reward is usually very sparse. And because the initial strategy is randomly generated, the probability of reaching the target is close to zero. To solve the problem of sparse rewards, design the arrival reward , which can be expressed as ; where M 1 is a constant greater than 0, M 2 is a constant used to prevent the denominator from being 0.
[0049] When the UAV flies out of the target disaster area, fails to complete the task within the time limit, runs out of energy, or collides, the system will give corresponding penalties according to different violation degrees, and the penalty factor is set to .
[0050] (4d), considering the remaining energy consumption, fairness rewards, arrival rewards, and wrong action penalties of the drones, the reward function is designed as .
[0051] The sum of the rewards of all drones is denoted as .
[0052] Step (5), based on IMATD3 to train the deep reinforcement learning model, adding an OU noise process to enhance the cruising ability of the drones in the dynamic disaster environment through time-correlated random perturbations, and adopting a priority cloning learning mechanism to intelligently copy and replay high-value experience samples to accelerate policy convergence. The specific steps are as follows: (5a), Start the environment simulator, initialize the environment and the network parameters of each agent, including the policy network, value network, and their corresponding target networks. Obtain the global environmental state.
[0053] (5b), Set the training period m and the maximum training period M, and loop through the following steps until the training is completed. At each time step, obtain the global environmental state at the current moment, obtain the partial observation state of each drone, and based on the greedy policy and the policy network to select an action . At the same time, use the OU noise mechanism to add noise perturbations to the action selection, so that the drones can better explore the environment during the decision-making process, avoid falling into local optimal solutions, and enhance the exploration ability of the agents in the dynamic disaster environment. Calculate the wrong action penalty according to the selected action and the environmental state , and determine whether the drone collides or flies out of the target disaster area; if a collision occurs or the drone flies out of the target disaster area, return to the initial position.
[0054] (5c), Obtain the partial observation state at the next moment . For each search and rescue device, select based on the minimum energy consumption greedy policy. For each drone, calculate the reward value .
[0055] (5d), Obtain the global state at the next moment , and store the experience at the current moment in the experience pool.
[0056] (5e) For each drone, a small batch of samples is taken from the experience pool according to the priority, and the loss functions of the two value networks are calculated, and then the value network parameters are updated. Based on the deterministic policy gradient ascent method, the policy network parameters are updated, and the target network parameters are adjusted. A priority-based experience cloning learning mechanism is added. Through this priority mechanism, drones can pay more attention to those experience samples with larger errors during the training process, thus accelerating the learning process. At the same time, in order to further improve the stability of the algorithm, a cloning learning mechanism is adopted to clone some parameters of the current policy network to the target network, which not only preserves the stability of the target network but also ensures the continuous optimization of the policy network through the update of some parameters.
[0057] (5f) In the network update stage of each time step, a small batch of samples is taken from the experience pool according to the sample priority. For each sample, the Q-values of the two value networks are calculated. The target Q-value is calculated, that is, the minimum value of the Q-values output by the two target value networks.
[0058] (5g) The value network parameters are optimized by minimizing the error between the Q-value output by the value network and the target Q-value.
[0059] (5h) Based on the output of the first online value network, the policy network parameters are updated using the deterministic policy gradient ascent method.
[0060] (5i) The target network parameters are gradually adjusted through a soft update mechanism.
[0061] Step (6), in the execution stage, each drone obtains the corresponding action based on the trained action network according to the environmental information it observes. After the task offloading is completed, the system assigns rewards according to the task completion situation and updates the observation space of the drone in real time, thereby further promoting the execution of subsequent actions and the dynamic adjustment of the state. The specific steps are as follows: (6a) Obtain the global environmental state at the current moment . Set the execution period w and the maximum execution period W. For each drone, obtain its partial observation state , and output the action based on the trained policy network . During the execution process, the OU noise mechanism is still retained to cope with the uncertainty in case of emergencies.
[0062] (6b) Execute the action , complete the search and rescue equipment computing offloading task, and calculate the reward value . Obtain the observation state at the next moment , and repeat the above process until the task is completed.
[0063] (6c) After the execution is completed, the cumulative reward value is statistically output 。
[0064] In Figure 3 , the reward result graphs of different algorithms provided by the embodiments of the present invention under different task volumes are described. Three other algorithms are simulated and compared, namely: the multi-agent SAC algorithm (Multi-Agent Soft Actor-Critic, MASAC), a deep reinforcement learning algorithm based on a stochastic policy and a maximum entropy mechanism; the traditional convex optimization algorithm (TRAMD), which uses the traditional convex optimization algorithm to plan the flight of the unmanned aerial vehicle from the starting point to the ending point; and the random algorithm (Random), in which in each time slot, each unmanned aerial vehicle randomly selects the flight direction and flight distance. It can be seen from the results that the IMATD3 algorithm demonstrates the best performance and adaptability. Taking the task of 15Kb as an example, the reward of the IMATD3 algorithm is 88% higher than that of MASAC, 194% higher than that of TRAMD, and 545% higher than that of Random, fully verifying the superiority and stability of the IMATD3 algorithm in high-task-load scenarios and its ability to effectively balance the task allocation and the energy consumption of unmanned aerial vehicles when the data volume increases.
[0065] In Figure 4 , the reward result graphs of different algorithms provided by the embodiments of the present invention under the maximum computing resources of different unmanned aerial vehicles are described. As the computing resources increase, the advantages of the IMATD3 algorithm gradually become prominent. Even in the case of the minimum computing resources, the IMATD3 still performs optimally. Taking the maximum computing resource of the unmanned aerial vehicle as 6GHz as an example, the reward of the HMATD3 algorithm is 61% higher than that of MASAC, 132% higher than that of TRAMD, and 447% higher than that of Random, verifying the ability of the IMATD3 algorithm to maximize the utilization of computing resources.
[0066] In Figure 5 , the comparison result graphs of the search and rescue equipment fairness indicators of different algorithms in different time slots are described. The IMATD3 algorithm always maintains the highest fairness of the search and rescue equipment in all time slots. Taking the critical rescue period of time slot 30 as an example, the fairness of the search and rescue equipment of the IMATD3 algorithm is 21% higher than that of MASAC, 48% higher than that of TRAMD, and 149% higher than that of Random, fully demonstrating the superiority of the IMATD3 algorithm in the resource scheduling of search and rescue equipment.
[0067] In Figure 6It describes the comparison result graph of the load balancing metrics of different algorithms for UAVs in different time slots provided by the embodiments of the present invention. In different time slots, through the exploration strategy enhanced by OU noise, the IAMTD3 algorithm makes the load balancing metrics always higher than other algorithms. Taking time slot 20 as an example, the load balancing metrics of the IAMTD3 algorithm are 29% higher than those of MASAC, 101% higher than those of TRAMD, and 136% higher than those of Random, which fully demonstrates the excellent performance of the IMATD3 algorithm in terms of load balancing and can achieve efficient resource scheduling in multi-UAV collaborative tasks.
[0068] According to the description of the present invention, those skilled in the art should not find it difficult to see that the multi-UAV collaborative computing offloading and resource allocation method proposed by the present invention for emergency search and rescue in the face of sudden disasters can effectively improve the service fairness of search and rescue equipment, achieve load balancing of UAVs and reduce energy consumption.
[0069] Based on the same inventive concept, the embodiments of the present invention also disclose a multi-UAV collaborative computing offloading and resource allocation system for emergency search and rescue in the face of sudden disasters, including: a scenario construction module for building a multi-UAV assisted three-dimensional computing offloading network architecture for emergency search and rescue scenarios, where UAVs provide computing services for search and rescue equipment, and the search and rescue equipment moves in the disaster area and generates computing tasks; a problem modeling module for determining the offloading ratio parameters of the search and rescue equipment in each time slot according to the task offloading decision variables of the search and rescue equipment for each UAV, and then measuring the task offloading fairness among the search and rescue equipment; determining the resource allocation ratio parameters of the UAVs according to the resource allocation decision variables of the UAVs for each search and rescue equipment, and then measuring the load balancing degree among the UAVs; taking into account the offloading fairness of the search and rescue equipment, the load balancing of the UAVs, and reducing the energy consumption of the UAVs, establishing a system utility function, and establishing a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that decides task offloading; a reinforcement learning module for training the deep reinforcement learning model based on the improved multi-agent double delayed deep deterministic policy gradient algorithm, introducing the OU noise and the priority cloning learning mechanism, and learning and optimizing the trajectory planning of the UAVs and the task offloading decisions of the search and rescue equipment.
[0070] The embodiments of the present invention also disclose a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters are implemented.
[0071] The embodiments of the present invention also disclose a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters are implemented.
[0072] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the steps of the method of the present invention are implemented. The program codes can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server. Where the present invention is not described in detail, it is all well-known technology to those skilled in the art.
Claims
1. A multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters, characterized in that, It includes the following steps: Build a multi-UAV assisted three-dimensional computing offloading network architecture for the search and rescue scenario of sudden emergency disasters, where the UAVs provide computing services for the search and rescue equipment, and the search and rescue equipment moves within the disaster area and generates computing tasks; According to the task offloading decision variables of the search and rescue equipment for each UAV, determine the offloading ratio parameters of the search and rescue equipment in each time slot, and then measure the fairness of task offloading among the search and rescue equipment; according to the resource allocation decision variables of the UAVs for each search and rescue equipment, determine the resource allocation ratio parameters of the UAVs, and then measure the load balancing degree among the UAVs; taking into account the offloading fairness of the search and rescue equipment, the load balancing of the UAVs and reducing the energy consumption of the UAVs, establish a system utility function, and establish a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that determines task offloading; Train the deep reinforcement learning model based on the improved multi-agent double-delayed deep deterministic policy gradient algorithm, introduce the Ornstein-Uhlenbeck noise and the priority cloning learning mechanism, and learn and optimize the trajectory planning of the UAVs and the task offloading decisions of the search and rescue equipment.
2. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 1, characterized in that, The system includes multiple rescue devices and multiple UAVs. The running time of the system includes multiple consecutive time slots, and the duration of each time slot can be dynamically adjusted according to the urgency of the rescue task; the movement of the UAVs in the system cannot exceed the boundary of the target disaster area, the flight altitude does not exceed the upper limit, the distance flown by the UAVs between adjacent time slots does not exceed the maximum distance, and the distance between the UAVs is not less than the safety distance.
3. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 1, characterized in that, The fairness of task offloading among the search and rescue equipment is expressed as follows: ; Among them is the search and rescue equipment n In the time slot t is the unloading ratio parameter N is the number of rescue equipment U is the number of drones is the time slot t search and rescue equipment n for the drone u task unloading decision variable; The load balancing degree among drones is expressed as follows: ; wherein is the resource allocation ratio parameter of the UAV u in time slot t ; and is the resource allocation decision variable of the UAV t for the search and rescue equipment u in time slot n .
4. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 1, characterized in that, In the deep reinforcement learning model, each UAV represents an agent. The observation space of the UAV includes the coordinates of all search and rescue equipment and all UAVs, the coordinates of the target point, the remaining energy consumption of the UAVs, as well as the task offloading situation of the search and rescue equipment and the load situation of the UAVs; the action space of the UAVs includes the flight angle, the flight distance, the range coverage, and the resource allocation.
5. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 1, characterized in that, In the deep reinforcement learning model, the rewards include the remaining energy consumption, fairness reward, arrival reward, and penalty for wrong actions; in a time slot t , the remaining energy consumption and fairness reward are defined as , where is the task offloading fairness among search and rescue devices, is the load balancing degree among UAVs, and are weights, is the remaining energy consumption of the UAV; the arrival reward is expressed as , where , are the UAV and the target location respectively, represents the Euclidean distance, M 1 is a constant greater than 0, M 2 is a constant used to prevent the denominator from being 0; when the UAV flies out of the target disaster area, fails to complete the task within the time limit, runs out of energy, or collides, a penalty for wrong actions is obtained.
6. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 5, characterized in that, When training the deep reinforcement learning model based on the improved multi-agent double-delayed deep deterministic policy gradient algorithm, for each UAV, obtain its partial observation state, and while selecting actions based on the greedy policy and the policy network, use the Ornstein-Uhlenbeck noise mechanism to add noise perturbations to the action selection, calculate the wrong action penalty according to the selected action and the environmental state, and determine whether the UAV collides or flies out of the target disaster area; if a collision or flying out of the target disaster area occurs, return to the initial position; obtain the partial observation state at the next moment, for each search and rescue equipment, make a selection based on the minimum energy consumption greedy policy, and for each UAV, calculate the reward value; Obtain the global state at the next moment and store the experience at the current moment in the experience pool.
7. The multi-UAV collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to claim 1, characterized in that, When training a deep reinforcement learning model based on an improved multi-agent double-delayed deep deterministic policy gradient algorithm, for each drone, a batch of samples is taken from the experience pool according to the priority, and the loss functions of the two value networks are calculated. Then, the value network parameters are updated, the policy network parameters are updated based on the deterministic policy gradient ascent method, and the target network parameters are adjusted. Among them, a cloning learning mechanism is adopted to clone some parameters of the current policy network to the target network.
8. A multi-UAV collaborative computing offloading and resource allocation system for emergency search and rescue in the face of sudden disasters, characterized in that, It includes: A scenario construction module for building a multi-drone assisted three-dimensional computing offloading network architecture for emergency disaster search and rescue scenarios. Among them, the drones provide computing services for the search and rescue equipment, and the search and rescue equipment moves in the disaster area and generates computing tasks. A problem modeling module for determining the offloading ratio parameters of the search and rescue equipment in each time slot according to the task offloading decision variables of the search and rescue equipment for each drone, and then measuring the task offloading fairness among the search and rescue equipment; determining the resource allocation ratio parameters of the drones according to the resource allocation decision variables of the drones for each search and rescue equipment, and then measuring the load balancing degree among the drones; taking into account the offloading fairness of the search and rescue equipment, the load balancing of the drones, and reducing the energy consumption of the drones, establishing a system utility function, and establishing a deep reinforcement learning model with the goal of maximizing the utility function for the search and rescue equipment that determines task offloading. A reinforcement learning module for training a deep reinforcement learning model based on an improved multi-agent double-delayed deep deterministic policy gradient algorithm, introducing Ornstein-Uhlenbeck noise and a priority cloning learning mechanism to learn and optimize the trajectory planning of the drones and the task offloading decisions of the search and rescue equipment.
9. A computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to any one of claims 1-7.
10. A computer program product, including a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-drone collaborative computing offloading and resource allocation method for emergency search and rescue in the face of sudden disasters according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle fair cooperation and task unloading optimization method and system
CN116887355A
Method for realizing high-energy-efficiency calculation unloading through strategy gradient algorithm in multi-unmanned aerial vehicle assisted mobile edge calculation
CN117499867A
Unmanned aerial vehicle auxiliary edge unloading method based on multi-agent deep reinforcement learning
CN118265041A
Task unloading method for emergency rescue scene under space-air-ground integrated network
CN119676767A
Unmanned aerial vehicle mobile edge computing resource allocation method based on deep reinforcement learning
CN119922617A
Cited By
Task allocation method
CN120996525A
Method for judging cargo unloading accessibility in disaster scene based on multi-source data fusion
CN122414949A