Proximity information associated multi-unmanned aerial vehicle dynamic task allocation method in emergency scene
By constructing the multi-UAV system model and dynamic task model in emergency scenarios in emergency scenarios, using the multi-UAV depth deterministic strategy gradient algorithm associated with neighboring information, the problems of dynamic tasks increase and computational complexity improvement in multi-UAV task allocation are solved, and higher task coverage and efficiency are achieved.
Patent Information
- Application Number
- CN202510462377.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In emergency scenarios, the allocation of multi-UAV missions faces the problems of increasing dynamic tasks and increasing computational complexity, and it is difficult for existing technologies to maximize task coverage and efficiency at disaster sites.
The dynamic task allocation method of multi-UAV based on proximity information association is adopted to represent the problem of maximizing task coverage as part of the observable Markov decision-making process. The proximity information association module and the deep deterministic strategy gradient algorithm are used to optimize task allocation to improve coverage through information exchange between drones.
The task completion rate is significantly improved, the coverage time is better than traditional methods, and the task completion rate is increased by about 5%-20%, which significantly reduces the communication and computing complexity in large-scale tasks, and has the ability to handle high-dimensional and large-scale task scheduling.
Smart Images

Figure CN120371012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for multi-UAV dynamic task allocation with adjacent information association in an emergency scenario, which provides reasonable task planning for multi-UAVs in an emergency scenario. Specifically, in post-disaster emergency rescue, it performs dynamic task allocation for multi-UAVs, enabling the UAVs to complete more tasks within a limited time, and belongs to the technical field of multi-UAV task allocation. Background Technique
[0002] Multi-UAV task allocation refers to the process of effectively allocating and arranging each UAV to complete specific tasks when multiple UAVs (Unmanned Aerial Vehicles) work together. It is an important issue in multiple UAV systems and is widely applied in fields such as military, search and rescue, agricultural monitoring, environmental protection, and logistics distribution.
[0003] In the field of emergency rescue, after a disaster occurs, traditional rescue methods are often affected by factors such as geographical environment, traffic blockage, and time constraints, resulting in low rescue efficiency or inability to respond quickly. The emergence of UAV technology provides a fast, efficient, and safe solution for emergency rescue. UAVs can quickly reach the disaster area to perform tasks such as real-time monitoring, searching for trapped people, and delivering rescue supplies, thus greatly improving the speed and accuracy of rescue. For example, one group of UAVs is responsible for high-altitude photography, and another group conducts ground searches to jointly complete a comprehensive assessment of the disaster area. After an earthquake, the UAVs quickly fly to the disaster area to collect high-resolution images of the affected area and generate 3D topographic maps to help rescue personnel evaluate the disaster situation and plan rescue routes. In areas flooded by floods, UAVs equipped with infrared imaging devices penetrate smoke and low-light conditions at night to search for the positions of trapped people and transmit real-time images back to the command center to guide rescue operations. At the scene of large-scale fires such as forest fires, UAVs conduct real-time monitoring to obtain the spread of the fire and assist the command center in formulating fire-fighting strategies. For the collaborative task allocation of multiple UAVs, it can ensure the efficient execution of different tasks in the disaster area through reasonable scheduling and collaboration, avoid conflicts between UAVs, and maximize the utilization rate of resources. Through intelligent task allocation, multiple UAVs can play a collaborative role in complex disaster scenarios and complete more complex tasks. Therefore, through reasonable planning of multi-UAV task allocation, we can collect more information in emergency situations and improve the rescue efficiency.
[0004] In the process of multi - UAV task allocation in emergency scenarios, due to the uncertainty of emergency scenarios, a large number of new tasks will appear during the flight of UAVs. The increase in new tasks makes multi - UAV task allocation more difficult. At the same time, because the number of tasks is huge, the number of UAVs cannot completely cover all tasks. However, existing work mostly focuses on multi - UAV task allocation in static situations, and the task allocation goals are all established on the premise that UAVs can cover all tasks. Most studies have placed the optimization goal on minimizing the task completion time, which does not conform to the situation where multi - UAVs cannot cover large - scale tasks in emergency scenarios. Summary of the Invention
[0005] Object of the Invention: To solve the problems existing in the prior art, it is necessary to consider the difficulties brought by the dynamically increasing tasks in emergency scenarios to multi - UAV task allocation. Aiming at the problem of increased computational complexity caused by a large - scale increase in the number of tasks, in the multi - UAV task allocation scenario in emergency scenarios of the present invention, a multi - UAV dynamic task allocation method and device based on proximity information association for emergency scenarios are proposed to solve the above problems of dynamic task increase and computational complexity improvement. Considering that the UAV swarm in emergency scenarios cannot completely cover all tasks, maximizing the UAV task coverage rate is taken as the optimization goal. This problem is expressed as a partially observable Markov decision process, and a multi - UAV dynamic task allocation scheme based on proximity information association for emergency scenarios is proposed, enabling the UAV swarm to maximize the task coverage rate and ensuring the maximum grasp of the latest situation in the disaster area.
[0006] Technical Solution: A multi - UAV dynamic task allocation method based on proximity information association in emergency scenarios, for a multi - UAV system facing emergency scenarios, includes the following steps:
[0007] (1) Construct a multi - UAV system model in emergency scenarios;
[0008] (2) Construct a dynamic task model;
[0009] (3) Construct the multi - UAV dynamic task allocation problem as a problem of maximizing multi - UAV task coverage optimization and construct a maximum task coverage objective function;
[0010] (4) Express maximizing the UAV task coverage as a partially observable Markov decision process;
[0011] (5) Regarding multi - UAVs as different UAVs, using the multi - UAV deep deterministic policy gradient algorithm based on proximity information association, through information exchange between different UAVs (including the current UAV position information, current uncompleted task information, and newly added dynamic task information), as well as the UAV's own current battery state and position, form a system state, and find a task allocation scheme that maximally covers the multi - UAV task points in the emergency scenario within a limited time.
[0012] (6) Randomly initialize the parameters of the Q-network and the policy network, and use the system state information as the input for each UAV.
[0013] (7) During training, each UAV obtains the current reward r and the state o' at the next moment according to the action it takes. Then the UAV will obtain the current set of neighboring UAVs G = (G1, G2,..., G N ), calculate the association information with each neighboring UAV, and associate these neighboring information through the NIR module to generate a unified input vector of neighboring information state φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G and the table of neighboring UAV actions φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ;
[0014] (8) The data of the current state, action, neighboring association information, and reward (o, a, φ(o), φ(a), r, o') are stored in the experience replay buffer. During each training, a small batch of data is randomly sampled from the experience replay buffer to update the policy network and the value network;
[0015] (9) Update the parameters of the target network;
[0016] (10) Repeat steps (6)-(9) until the iteration process ends, and find the task allocation scheme with the largest task coverage rate.
[0017] Further, in step (1), a multi-UAV system model in an emergency scenario is constructed. The set of UAVs consists of a total of N UAVs, which together form a multi-UAV system, represented by a graph structure . Among them, ε represents the set of communication edges between multi-UAVs, representing direct communication and perception between UAVs. The energy consumption of communication between UAVs is ignored.
[0018] indicates the neighboring set of UAV i, defined as The neighbor set is dynamic and updated as the UAVs move or the environment changes.
[0019] The neighbor of UAV i is determined by its communication radius R com , satisfying: where is the Euclidean distance between UAV i and UAV i' at time t, represents the position of UAV i at time t.
[0020] In step (2), a dynamic task model is constructed, and the task set is defined as Task set dynamically increases over time. The task j at the current time point t is represented by j t Each task includes the task coordinates, the priority of the task, and the task status.
[0021] Among them indicates the status of the task at the current time t, pos = (x j , y j ) indicates the task coordinates, prio ∈ [0, 1] represents the priority level of the task, 0 is an ordinary task, and 1 represents an urgent task.
[0022] New tasks are discovered by UAVs during flight. UAV i discovers a new task j at time t new , and shares the new task information with neighboring UAVs within the communication range.
[0023] In step (3), the multi-UAV dynamic task allocation problem is formulated as a problem of maximizing the multi-UAV task coverage optimization, and a maximizing task coverage objective function is proposed. For multi-UAV task allocation in an emergency rescue scenario, there are M tasks, represented as a set UAVs are arbitrarily distributed at various known positions in the emergency scenario. Assume that the UAVs and tasks under discussion are of the same type, and each task requires only one UAV to complete. Based on the above conditions, the problem of allocating M tasks to N UAVs can be expressed as the following optimization formula:
[0024]
[0025]
[0026] Among them, ξ ij represents the situation where the UAV completes the corresponding task. If UAV i completes task j, then ξ ij = 1, otherwise ξ ij = 0. Equation (1a) ensures that the tasks executed by the UAV do not exceed the maximum number of executable tasks. Equation (1b) ensures that each task is either not completed or completed by only one UAV. Equation (1c) is the maximum flight time constraint, and the time for the UAV to execute all tasks cannot exceed the maximum flight time of the UAV itself. Equation (1d) is the safety distance constraint that must be satisfied between UAVs. Equation (1e) is the constraint to ensure that the UAV cannot fly over the obstacle area.
[0027] In step (4), maximizing the UAV mission coverage is represented as a partially observable Markov decision process, and the state space, action space, and reward function in the decision process are specifically represented as follows:
[0028] A. State space: The system space observed by the UAV at time t in the state space should include the position information of the current UAV battery information the set of neighboring UAVs The set of neighboring UAVs includes the position set battery mission information and the current mission list carried Therefore, the state space of the UAV at time t is represented as The state space set of the UAV is
[0029] B. Action space: The action taken by the UAV at time t should include the mission point, flight direction, and flight speed. The formula is as follows:
[0030] C. Reward function: The goal of the reward function is to encourage the UAV group to improve the mission completion efficiency, avoid collisions, avoid obstacles, and reduce the mission completion time when completing the mission. Therefore, the reward function can be divided into mission completion reward, collision reward, and maximum mission completion time reward;
[0031] The mission completion reward is including the mission point reward R close,i (t), and and the reward for distinguishing ordinary missions from emergency missions;
[0032] The collision reward includes the collision reward between UAVs and the collision reward with obstacles, which are respectively and Among them, R collide,ij is the collision reward between UAVs, and R obstacle,i is the obstacle collision reward between the UAV and the obstacle. represents the distance between the UAV and the obstacle. λ1 and λ2 are collision penalty coefficients, and d collide , d obstacle is the obstacle collision threshold.
[0033] Maximum mission completion time reward: Among them, T minis the preset shortest time, κ1 is a constant that controls the influence intensity of time on the reward and determines the decay rate of the reward, κ2 is a constant that controls the decay speed, and MaxCompletionTime represents the maximum time to complete all tasks.
[0034] The total reward function is as follows:
[0035] In step (7), the method constructs a neighboring information association module. In the neighboring information association module, each drone only needs to focus on the set of neighboring information that is directly associated with itself and does not need to focus on global information. It contains two parts: represents the set of drones and task points within the set range from the current drone i; represents the set of actions within the set range from the current drone i. For drone i, the associated neighboring information is calculated through the following formula, observation value aggregation: where is the observation value of neighboring related drone k.
[0036] Action set: where is the action of neighboring related robot k. The reciprocal of the distance between drones is selected as the weight: where based on the distance between drones the weighted coefficient is calculated and the parameter β is added to adjust the weight. Such a design assigns weights according to proximity priority and is especially suitable for emergency scenarios.
[0037] Through the neighboring information association module, the φ function can effectively integrate the collective knowledge and actions of nearby drones. This information association module φ function not only ensures the consistency of the input dimensions but also captures the complexity of the environmental dynamics. By integrating the information from the observations and actions of neighboring related drones, the φ function allows the drones to make more informed decisions under a wider range of environmental conditions and reflects the immediate dynamics of the environment and the synchronous behavior among surrounding robots.
[0038] The method designs a centralized value function G. The corresponding shared value network G takes the neighboring observations and actions of drone i as well as the neighboring association information and as inputs, and the value G of drone i i is calculated as follows:
[0039] The formula for the target value network G i ′ is as follows:
[0040] By evaluating the expectation of the Q value, the shared value network G is trained by minimizing the generalized TD error: where the parameters of the shared policy network μ are optimized by maximizing the expected reward of the UAV, and its gradient formula is:
[0041] Using the soft update method θ′ g ← ηθ g + (1 - η)θ′ g , θ′ a ← ηθ a + (1 - η)θ′ a Update the weights of the target network until the training terminates to obtain a task allocation scheme that maximizes the task coverage rate.
[0042] A multi-UAV dynamic task allocation device for emergency scenarios based on proximity information association, comprising:
[0043] The first module constructs a multi-UAV system model and a dynamic task model for emergency scenarios;
[0044] The second module constructs a target function for maximizing task coverage;
[0045] The third module represents maximizing UAV task coverage as a partially observable Markov decision process;
[0046] The fourth module treats multiple UAVs as different UAVs, uses the multi-UAV deep deterministic policy gradient algorithm based on proximity information association, and through information exchange between different UAVs (including the current UAV position information, the current uncompleted task information, and the newly added dynamic task information), as well as the UAV's own current battery state and position, to form a system state, and finds a task allocation scheme that maximally covers the task points of multiple UAVs in the emergency scenario within a limited time;
[0047] Randomly initialize the parameters of the Q network and the policy network, and use the system state information as the input for each UAV;
[0048] The UAV will obtain the current set of neighboring UAVs G = (G1, G2,..., G N ), calculate the association information with each neighboring UAV, and associate these neighboring information through the NIR module to generate a unified input vector φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G , φ(a) = (φ1(a k), φ2(a k ),..., φ N (a k )) k∈G . During training, each drone obtains the current reward r and the next state o' according to the actions it takes.
[0049] Next, the drone will obtain the current set of neighboring drones G = (G1, G2,..., G N ), calculate the association information with each neighboring drone, associate these neighboring information through the NIR module and generate a unified input vector φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G , φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ; The data of the current state, action, neighboring association information, and reward (o, a, φ(o), φ(a), r, o') are stored in the experience replay buffer. During each training, a small batch of data is randomly sampled from the experience replay buffer to update the policy network and the value network; the target network parameters are updated; until the iteration process ends, the task allocation plan with the largest task coverage rate is found.
[0050] The implementation process of the device is the same as the method and will not be elaborated here.
[0051] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the multi - drone dynamic task allocation method based on neighboring information association in the emergency scenario as described above.
[0052] A computer - readable storage medium stores a computer program that executes the multi - drone dynamic task allocation method based on neighboring information association in the emergency scenario as described above.
[0053] In the post-disaster emergency scenario, due to the unpredictability of the scenario, a large number of new tasks will be brought. At the same time, the number of drones cannot cover all tasks. If the traditional method is still used, it is difficult to meet the rationality and effectiveness of multi-drone task allocation in the emergency scenario. Aiming at the problem of multi-drone dynamic task allocation in the post-disaster emergency scenario, a multi-drone dynamic task allocation method based on proximity information association for emergency scenarios is proposed. First, the problem of maximizing the drone task coverage rate is expressed as a partially observable Markov decision process (POMDP). Secondly, by using the proximity information association module to associate the state information and task information of other drones, the task coverage rate of the drone swarm is improved. Finally, the optimal task allocation strategy is solved by the algorithm of multi-drone deep deterministic policy gradient based on proximity information association (NIR_MADDPG).
[0054] Beneficial effects: The present invention has the following advantages compared with the prior art:
[0055] The present invention is a multi-drone dynamic task allocation method based on proximity information association for emergency scenarios. Considering that in the emergency scenario where there are a large number of new tasks, the emergency command center needs to know as much disaster information as possible. Taking maximizing the drone task coverage rate as the optimization goal, the problem is expressed as a partially observable Markov decision process, and a multi-drone dynamic task allocation method based on proximity information association for emergency scenarios is proposed. The present invention has significant advantages in multi-drone collaborative task allocation, and the task completion rate is increased by about 5%-20% compared with the traditional methods (Greedy, MAAC, MAPPO). When the task scale increases significantly, NIR_MADDPG (the method of the present invention) is also significantly better than the exponential or higher-order growth trend of the baseline method in terms of task coverage time, which indicates that this method reduces the communication and computational complexity of multi-drone collaboration through proximity information aggregation and has the ability to handle high-dimensional and large-scale task scheduling. Description of the Drawings
[0056] Figure 1 It is the flowchart of multi-drone dynamic task allocation in the emergency scenario of the embodiment of the present invention. Specific Embodiments
[0057] The following combines specific embodiments to further clarify the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, those skilled in the art's various equivalent forms of modification of the present invention fall within the scope defined by the appended claims of this application.
[0058] In the post-disaster emergency scenario, the drone swarm initially flies according to the preset task allocation plan. After discovering new tasks, the present invention proposes a multi-drone dynamic task allocation method with proximity information association in the emergency scenario, havingFigure 1 The process and method implementation include the following steps:
[0059] (1) Construct a multi-UAV system model in an emergency scenario;
[0060] (2) Construct a dynamic task model;
[0061] (3) Construct the multi-UAV dynamic task allocation problem as a problem of maximizing multi-UAV task coverage optimization and construct a maximization task coverage objective function;
[0062] (4) Represent the maximization of UAV task coverage as a partially observable Markov decision process;
[0063] (5) Treat the multi-UAVs as different UAVs and use the multi-UAV deep deterministic policy gradient algorithm based on proximity information association. Through information exchange between different UAVs (including the current UAV position information, the current uncompleted task information, and the newly added dynamic task information), as well as the UAV's own current battery state and position, form a system state, and find a task allocation plan that maximally covers the multi-UAV task points in the emergency scenario within a limited time;
[0064] (6) Randomly initialize the parameters of the Q-network and the policy network, and use the system state information as the input for each UAV;
[0065] (7) In training, each UAV obtains the current reward r and the state o' at the next moment according to the action it takes. Then the UAV will obtain the current set of neighboring UAVs G = (G1, G2,..., G N ), calculate the association information with each neighboring UAV, and associate these neighboring information through the NIR module to generate a unified input vector φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G , φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ;
[0066] (8) Store the data of the current state, action, neighboring association information, and reward (o, a, φ(o), φ(a), r, o') in the experience replay buffer. During each training, randomly extract a small batch of data from the experience replay buffer to update the policy network and the value network;
[0067] (9) Update the target network parameters;
[0068] (10) Repeat steps 6 - 10 until the iteration process ends to find the task allocation plan with the maximum task coverage rate;
[0069] In step (1), a multi - UAV system model in the emergency scenario is constructed. The UAV set consists of a total of N UAVs, which jointly form a multi - UAV system and is represented by a graph structure where ε represents the communication edge set between multi - UAVs, representing the direct communication and perception between UAVs. In this chapter, the communication energy consumption between UAVs is ignored because in the high - speed movement scenario of post - disaster emergency rescue, the communication energy consumption only accounts for 1% or less of the total energy consumption.
[0070] denotes the neighboring set of UAV i, defined as The neighbor set is dynamic and is updated as the UAV moves or the environment changes.
[0071] The neighbor of UAV i is determined by its communication radius R com and satisfies: where is the Euclidean distance between UAV i and UAV i′ at time t
[0072] In step (2), a dynamic task model is constructed. The task set is defined as The task set increases dynamically with time. Each task includes the task coordinates, the priority of the task, and the task status.
[0073] where indicates the status of task j at the current time t, pos=(x j , y j ) indicates the task coordinates, prio ∈ [0, 1] represents the priority degree of the task, 0 is an ordinary task, and 1 represents an urgent task.
[0074] New tasks are discovered by UAVs during flight. UAV i discovers a new task j at time t new and shares the new task information with neighboring UAVs within its communication range.
[0075] In step (3), the multi - UAV dynamic task allocation problem is constructed as a problem of maximizing the multi - UAV task coverage optimization, and a task coverage objective function for maximization is proposed. The multi - UAV task allocation in the emergency rescue scenario includes M tasks, represented as a set The drones are arbitrarily distributed at various known positions in the emergency scenario. In this section, it is assumed that the drones and tasks under discussion are of the same type of drones and the same type of tasks, and only one drone is required to complete the task. Based on the above conditions, the problem of allocating M tasks to N drones can be expressed as the following optimization formula:
[0076]
[0077] Among them, ξ ij represents the situation where the drone completes the corresponding task. If drone i completes task j, then ξ ij = 1, otherwise ξ ij = 0. Formula (1a) ensures that the tasks executed by the drones do not exceed the maximum number of tasks that can be executed. Formula (1b) ensures that each task is either not completed or is completed by only one drone; Formula (1c) is the maximum flight time constraint, and the time for the drone to execute all tasks cannot exceed the maximum flight time of the drone itself; Formula (1d) is the safety distance constraint that must be satisfied between drones; Formula (1e) is the constraint to ensure that the drone cannot fly over the obstacle area.
[0078] In step (4), maximizing the drone task coverage is represented as a partially observable Markov decision process. The state space, action space, and reward function in the decision process are specifically represented as follows:
[0079] A. State space: The system space observed by the drone at time t in the state space should include the position information battery information set of neighboring drones which includes the position set battery level task information and the current task list carried So the state space of the drone at time t is represented as The set of the drone's state space is
[0080] B. Action space: The action taken by the drone at time i should include the task point, flight direction, and flight speed. The formula is as follows:
[0081] C. Reward function: The goal of the reward function is to encourage the drone group to improve the task completion efficiency, avoid collisions, avoid obstacles, and reduce the task completion time when completing tasks. Therefore, the reward function can be divided into the following parts: task completion reward including task point rewards R close,i (t) and R normal,i and R emergency,iDistinguish the rewards for ordinary tasks and emergency tasks, collision rewards (including collision rewards between drones and obstacle collision rewards, which are and ), respectively, and the maximum task completion time reward (R time ). Among them, λ1 and λ2 are collision penalty coefficients, d collide , d obstacle is the obstacle collision threshold.
[0082] Maximum task completion time reward: Among them, T min is the preset shortest time, κ1 is a constant that controls the influence intensity of time on the reward and determines the rate of reward attenuation, and κ2 is a constant that controls the attenuation speed.
[0083] The total reward function is:
[0084] In step (7), this method constructs a neighboring information association module. In the neighboring information association module, each drone only needs to focus on the set of neighboring information that has a direct association with itself and does not need to focus on global information. It contains two parts: is the set of drones and task points that are relatively close to the current drone i; is the set of actions that are relatively close to the current drone i. For drone i, the associated neighboring information is calculated through the following formula, observation value aggregation: where is the observation value of the neighboring relevant drone k.
[0085] Action set: where is the action of the neighboring relevant robot k. In this chapter, the reciprocal of the distance between drones is selected as the weight: where the weighted coefficient is calculated based on the distance between drones and the parameter β is added to adjust the weight. Such a design assigns weights according to proximity priority and is particularly suitable for the research scenario of this chapter.
[0086] Through the neighboring information association module, the φ function can effectively integrate the collective knowledge and actions of nearby robots. This aggregation function not only ensures the consistency of the input dimensions but also captures the complexity of the environmental dynamics. By integrating the information from the observations and actions of neighboring relevant robots, the φ function allows the robot to make more informed decisions under a wider range of environmental conditions and reflects the immediate dynamics of the environment and the synchronous behavior between surrounding robots.
[0087] The method designs a centralized value function G. The corresponding shared value network G takes the neighboring observations and actions of drone i as well as neighboring association information and as inputs, and the value G of drone i i is calculated by the following formula:
[0088] The formula for the target value network G i ′ is as follows:
[0089] By evaluating the expectation of the Q value, the shared value network G is trained by minimizing the generalized TD error: where The parameters of the shared policy network μ are optimized by maximizing the expected reward of the drone, and its gradient formula is:
[0090] Using the soft update method θ′ g ←ηθ g +(1 - η)θ′ g , θ′ a ←ηθ a +(1 - η)θ′ a to update the weights of the target network. Finally, the above training process will loop continuously until the training terminates to obtain a task allocation scheme that maximizes the task coverage rate.
[0091] Obviously, those skilled in the art should understand that each step of the above multi - drone dynamic task allocation method for neighboring information association in the emergency scenario of the present invention embodiment or each module of the multi - drone dynamic task allocation system for neighboring information association in the emergency scenario can be implemented by a general - purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program code executable by the computing device, so that they can be stored in the storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario, characterized in that, Multi - UAV system for emergency scenarios, including the following steps: (1) Construct a multi - UAV system model for emergency scenarios; (2) Construct a dynamic task model; (3) Construct the multi - UAV dynamic task allocation problem as a problem of maximizing multi - UAV task coverage optimization and construct a maximizing task coverage objective function; (4) Represent the maximization of UAV task coverage as a partially observable Markov decision process; (5) Treat multi - UAVs as different UAVs, and use the multi - UAV deep deterministic policy gradient algorithm based on proximity information association. Through information exchange between different UAVs and the UAV's own current battery state and position, form a system state, and find a task allocation plan that maximally covers the multi - UAV task points in the emergency scenario within a limited time; (6) Randomly initialize the parameters of the Q - network and the policy network, and use the system state information as the input for each UAV; (7) During training, each drone obtains the current reward r and the state at the next moment according to the actions it takes; then the drone will obtain the current set of neighboring drones G = (G1, G2,..., G N ), calculate the association information with each neighboring drone, and associate these neighboring information through the NIR module to generate a unified input vector neighboring information state φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G and the table neighboring drone actions φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ; (8) Store the data of the current state, action, proximity association information, and reward (o,a,φ(o),φ(a),r,o′) in the experience replay buffer. During each training, randomly sample data from the experience replay buffer to update the policy network and the value network; (9) Update the target network parameters; (10) Repeat steps (6) - (9) until the iteration process ends, and find a task allocation plan with the maximum task coverage rate.
2. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, wherein In step (1), a multi - UAV system model is constructed for the emergency scenario. The UAV set consists of a total of N UAVs, which together form a multi - UAV system and is represented by a graph structure where ε represents the set of communication edges between multi - UAVs, representing direct communication and perception between UAVs; denotes the neighboring set of UAV i, defined as The neighbor set is dynamic and updated with the movement of UAVs or environmental changes; Neighbors of UAV i Determined by its communication radius R com Satisfying: Where Is the Euclidean distance between UAV i and UAV i' at time t, Represents the position of UAV i at time t.
3. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, wherein In step (2), a dynamic task model is constructed, and the task set is defined as Task set which increases dynamically over time. The task j at the current time point t is represented by j t and each task includes the task coordinates, the priority of the task, and the task status; where indicates the status of the task at the current time t, pos = (x j , y j ) indicates the task coordinates, prio ∈ [0, 1] represents the priority level of the task, 0 is an ordinary task, and 1 represents an urgent task; New tasks are discovered by drones during flight. Drone i discovers new task j at time t new and shares the new task information with neighboring drones within its communication range.
4. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, wherein The step (3) constructs the multi-UAV dynamic task allocation problem into a problem of maximizing the multi-UAV task coverage optimization, and proposes a maximizing task coverage objective function. The multi-UAV task allocation in the emergency rescue scenario includes M tasks, which are represented as a set The UAVs are arbitrarily distributed at various known positions in the emergency scenario; it is assumed that the UAVs and tasks under discussion are of the same type of UAV and the same type of task, and each task can be completed by only one UAV. The problem of allocating M tasks to N UAVs can be expressed as the following optimization formula: S.t. where ξ ij represents the situation where the UAV completes the corresponding task. If UAV i completes task j, then ξ ij = 1; otherwise, ξ ij = 0. Formula (1a) ensures that the tasks performed by the UAV do not exceed the maximum number of tasks that can be performed. Formula (1b) means that each task is either not completed or is completed by only one UAV. Formula (1c) is the maximum flight time constraint, and the time for the UAV to perform all tasks cannot exceed the maximum flight time of the UAV itself. Formula (1d) is the safety distance constraint that must be satisfied between UAVs. Formula (1e) is the constraint to ensure that the UAV cannot fly over the obstacle area.
5. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, wherein In step (4), representing the maximization of UAV task coverage as a partially observable Markov decision process, the state space, action space, and reward function in the decision process are specifically represented as: A. State Space: The system space observed by the UAV at the t-th moment in the state space should include the position information of the current UAV Battery information Set of neighboring UAVs The set of neighboring UAVs includes the position set Battery level Task information And the current task list carried Therefore, the state space of the UAV at the t-th moment is expressed as The set of the state space of the UAV is B. Action Space: The action taken by the UAV at the i-th moment should include the mission point, flight direction, and flight speed; C. Reward function: The goal of the reward function is to encourage the UAV group to improve task - completion efficiency, avoid collisions, avoid obstacles, and reduce task - completion time when completing tasks; the reward function is divided into task - completion reward, collision reward, and maximum task - completion time reward; The task completion reward is including the task point reward R close,i (t), and R normal,i and R emergency,i distinguishing the rewards for ordinary tasks and emergency tasks; The collision rewards include the collision rewards between drones and the collision rewards with obstacles, which are respectively and where R collide,ij is the collision reward between drones, and R obstacle,i is the collision reward with obstacles between the drone and the obstacle. represents the distance between the drone and the obstacle, and λ1, λ2 are the collision penalty coefficients, d collide , d obstacle is the obstacle collision threshold. Maximum task completion time reward: where T min is the preset shortest time, κ1 is a constant that controls the influence strength of time on the reward and determines the reward decay rate, κ2 is a constant that controls the decay speed, and MaxCompletionTime represents the maximum time to complete all tasks; The total reward function is as follows:
6. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, wherein In step (7), a neighboring information association module is constructed. In the neighboring information association module, each drone only needs to focus on the set of neighboring information directly associated with itself and does not need to focus on global information; It includes two parts: represents the set of drones and task points within the set range of the current drone i; represents the set of actions within the set range of the current drone i; for drone i, the associated neighboring information is calculated by the following formula Observation value aggregation: where is the observation value of the neighboring relevant drone k; Set of actions: Among them is the action of the robot k adjacent to the relevant one; The reciprocal of the distance between the UAVs is selected as the weight: Among them, based on the distance between the UAVs the weighted coefficient is calculated and the parameter β is added to adjust the weight.
7. The multi-UAV dynamic task allocation method for adjacent information association in an emergency scenario according to claim 1, characterized in that A centralized value function G is designed; the corresponding shared value network t takes the neighboring observations and actions of drone i as well as neighboring association information and as inputs, and the value G of drone i i is given by the following formula: Target value network G i The formula for ' By to evaluate the expectation of the Q value, the shared value network G is trained by minimizing the generalized TD error: where the parameters of the shared policy network μ are optimized by maximizing the expected reward of the UAV, and its gradient formula is: Use the soft update method θ ′g ← ηθ g + (1 - η)θ ′g , θ ′a ← ηθ a + (1 - η)θ ′a Update the weights of the target network until the training terminates to obtain a task allocation scheme that maximizes the task coverage rate.
8. A multi-UAV dynamic task allocation device based on proximity information association for emergency scenarios, characterized in that Including: The first module constructs a multi - UAV system model and a dynamic task model for emergency scenarios; The second module constructs a maximizing task coverage objective function; The third module represents the maximization of UAV task coverage as a partially observable Markov decision process; The fourth module treats multi - UAVs as different UAVs, and uses the multi - UAV deep deterministic policy gradient algorithm based on proximity information association. Through information exchange between different UAVs and the UAV's own current battery state and position, form a system state, and find a task allocation plan that maximally covers the multi - UAV task points in the emergency scenario within a limited time; Randomly initialize the parameters of the Q - network and the policy network, and use the system state information as the input for each UAV; The drone will obtain the current set of neighboring drones G = (G1, G2,..., G N ), calculate the association information with each neighboring drone, associate these neighboring information through the NIR module and generate a unified input vector φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G , φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ; During training, each drone obtains the current reward r and the state o' at the next moment according to the action it takes; Next, the drone will obtain the current set of neighboring drones G = (G1, G2,..., G N ), calculate the association information with each neighboring drone, associate these neighboring information through the NIR module, and generate a unified input vector φ(o) = (φ1(o k ), φ2(o k ),..., φ N (o k )) k∈G , φ(a) = (φ1(a k ), φ2(a k ),..., φ N (a k )) k∈G ; The data of the current state, action, neighboring association information, and reward (o, a, φ(o), φ(a), r, o′) are stored in the experience replay buffer. During each training, data is randomly sampled from the experience replay buffer to update the policy network and the value network; the target network parameters are updated; Until the iteration process ends, find a task allocation plan with the maximum task coverage rate.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the multi - UAV dynamic task allocation method based on proximity information association for emergency scenarios as described in any one of claims 1 - 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for executing the multi-UAV dynamic task allocation method based on proximity information association in an emergency scenario as described in any one of claims 1-7.
Citation Information
Cited By
Multi-UAV (unmanned aerial vehicle) task allocation method, system and equipment for multi-layer cooperation graph, and medium
CN122086105A