A multi-unmanned aerial vehicle task allocation method in a fuzzy dynamic environment
By using cloud models and deep reinforcement learning algorithms, a UAV task allocation model was constructed, which solved the task allocation problem in dynamic and fuzzy environments and achieved efficient, flexible and stable decision-making for UAV task allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-29
Smart Images

Figure CN122114496A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology and relates to a method for multi-UAV task allocation in a fuzzy dynamic environment. Background Technology
[0002] Patent CN117035435A provides a method for multi-UAV task allocation and trajectory planning optimization in dynamic environments, belonging to the field of UAV technology. It employs a multi-agent reinforcement learning algorithm (MA SAC) based on the traditional deep reinforcement learning (SAC) algorithm, integrating the traditional SAC algorithm into a multi-agent network structure. Using a centralized training and distributed execution approach, it enables interaction and learning among agents, achieving higher reward values and task completion rates in a shorter time. This improves the timeliness of task planning strategy decisions, reduces the number of iterations in the intelligent optimization algorithm, and enhances the method's efficiency. By constructing a reward function based on a policy set, it improves the training efficiency and stability of reinforcement learning, solves the problem of reward sparsity in dynamic environments, and increases the convergence speed of the multi-agent reinforcement learning algorithm. However, the reinforcement learning algorithm used in this method is centrally trained, requiring retraining when the number of UAVs changes, leading to increased costs.
[0003] Patent CN116663803A discloses a method and system for multi-UAV task allocation considering uncertainty information, solving the problem of multi-UAV and multi-target allocation in battlefield environments. This patent first considers the uncertainty information of the battlefield environment, mathematically representing the uncertainty variables based on triangular fuzzy number theory. Combining UAV information and target information, it models the multi-UAV task allocation problem under uncertainty as a partially observable Markov decision model. Then, based on the established task allocation model, it improves the reinforcement learning algorithm by designing a deep Q-network algorithm based on improved reward values and introduces a priority experience replay mechanism to accelerate the algorithm's convergence speed. Although this method considers the uncertainty information of the battlefield environment, it does not consider the dynamic nature of the battlefield environment and struggles to handle environments where the number of tasks and UAVs changes, requiring more complex methods to improve performance.
[0004] Patent CN118780523A provides a real-scene modeling UAV system and its control method, relating to the field of UAV modeling technology. The method includes: using multimodal sensors onboard the UAV to perform a preliminary scan of the real-scene environment to form a first environmental perception map; using an improved ant colony algorithm to divide the task area and allocate UAV tasks, generating an initial flight path; using a reinforcement learning algorithm to pre-train the acquired historical flight data to obtain a real-time environmental feedback report; using an edge computing module to analyze the data in the real-time environmental feedback report to generate a second environmental perception map; and using a data fusion algorithm to fuse data from different sensors to generate a high-precision 3D environmental model. However, this method has high requirements for the environment and only uses deterministic input, making it difficult to handle UAV task allocation problems in fuzzy and uncertain environments.
[0005] In summary, current solutions for drone task allocation have the following drawbacks: 1. They are too demanding in terms of environmental requirements, making it difficult to handle task allocation in dynamic environments; 2. Some patents do not consider uncertain information and are only suitable for environments with certain information. They perform poorly in environments with a large amount of fuzzy information; 3. They commonly use the popular deep reinforcement learning method. However, since deep reinforcement learning mostly uses centralized training, it is difficult to continue execution when the number of drones changes, often requiring retraining, which increases the cost of model implementation. Summary of the Invention
[0006] To address the aforementioned problems in the prior art, this invention employs a multi-UAV task allocation method under a fuzzy dynamic environment, comprising:
[0007] S1. Construct a multi-UAV task allocation model; the multi-UAV task allocation model includes UAVs and tasks; each UAV and task includes multiple attributes.
[0008] S2. Use a cloud model to construct cloud representations of each drone attribute and each mission attribute;
[0009] S3. Construct an optimization objective function for multi-drone task allocation based on the cloud representation of each drone attribute and each task attribute;
[0010] S4. The objective function for multi-UAV task allocation is modeled as a Markov decision process to obtain the state space, action space, and reward function based on cloud model multi-attribute decision for each UAV.
[0011] S5. Based on the state space, action space, and reward function of the cloud model multi-attribute decision for each UAV, a deep reinforcement learning algorithm is used to obtain the optimal task allocation decision.
[0012] Beneficial effects:
[0013] 1. Unlike traditional precise modeling or fuzzy number-based methods, this invention introduces cloud model theory into the field of multi-UAV task allocation. It describes and transforms fuzzy linguistic variables in UAV and task attributes, simultaneously expressing the fuzziness and randomness of uncertain information. It can more accurately and robustly quantify and process fuzzy linguistic information in the environment, making environmental modeling closer to reality and thus improving the effectiveness of UAV task allocation. 2. Addressing the issues of reward sparsity and design difficulties in reinforcement learning, this invention constructs a cloud decision matrix, automatically optimizes attribute weights, and calculates the relative cloud distance between the task and the ideal task (positive ideal cloud and negative ideal cloud). This provides UAVs with a dense and information-rich reward signal, effectively solving the reward sparsity problem and improving the effectiveness of UAV task allocation. 3. This invention introduces policy entropy into the loss function of the Actor network to promote exploration and value function pruning into the loss function of the Critic network to stabilize training. 4. This invention supports fully decentralized training and execution, flexibly adapting to the dynamic increase or decrease in the number of UAVs or tasks in the environment without retraining the entire model, demonstrating better stability and adaptability than centralized training algorithms. Attached Figure Description
[0014] Figure 1 A flowchart illustrating a multi-UAV task allocation method in a fuzzy dynamic environment, as provided in this embodiment of the invention;
[0015] Figure 2 The cloud model ablation experiment diagram provided in the embodiment of the present invention;
[0016] Figure 3 A schematic diagram of the reward curve provided in an embodiment of the present invention;
[0017] Figure 4 This is a schematic diagram of the task completion rate curve provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] like Figure 1 As shown, this embodiment of the invention employs a multi-UAV task allocation method in a fuzzy dynamic environment, including:
[0020] S1. Construct a multi-UAV task allocation model; the multi-UAV task allocation model includes multiple UAVs and multiple tasks; each UAV and task includes multiple attributes.
[0021] Task attributes may include fixed value, importance, execution difficulty, and location; drone attributes may include movement speed, mission capacity, execution capability, and location. These attributes are mostly linguistic variables such as "high," "medium," and "low."
[0022] S2. Use a cloud model to construct cloud representations of each drone attribute and each mission attribute;
[0023] For any drone attribute or mission attribute, constructing its cloud representation includes:
[0024] S21, Valid domain of attributes provided by experts , expand the scope of the attribute Divided into N sub-semantics To obtain the sub-semantic set (e.g., very poor, poor, average, good, very good); among them, These are the minimum and maximum values, respectively. Index for sub-semantics;
[0025] The intermediate sub-semantic definition is Then the remaining sub-semantics , It can be represented as:
[0026]
[0027] S22, Calculate each sub-semantic Expectations (Reflecting the central value of cloud droplet distribution in the domain space), entropy (A measure of uncertainty reflecting qualitative concepts) and hyperentropy (Reflecting the uncertainty of entropy), each sub-semantic is obtained. triples This means that the cloud representation enables the model to better express the fuzziness and randomness of attributes;
[0028] S23, Combining all sub-semantics The cloud indicates The cloud representation of the obtained attributes .
[0029] In one embodiment, the golden ratio method can be used to define the effective domain of the attribute. It is divided into 5 sub-semantics, and each sub-semantic is calculated. From the expectation, entropy, and hyperentropy, we obtain a cloud representation of the five sub-semantics of the attribute. :
[0030]
[0031] in, It is the golden ratio. It is the complement of the golden ratio.
[0032] S3. Construct an optimization objective function for multi-drone task allocation based on the cloud representation of each drone attribute and each task attribute;
[0033] Calculate each drone Execute each task payoff function :
[0034]
[0035] in, , , , As weight, , Tasks value Cloud representation of corresponding sub-semantics, tasks Importance The cloud representation of the corresponding sub-semantics, For drones With the task Manhattan distance between them For drones Location, For the task Location, For drones Execute the task Capability utilization value drones execution capability Cloud representation and tasks corresponding to sub-semantics execution difficulty The cloud representation of the corresponding sub-semantics, Represents the distance function.
[0036] Preferably, the distance function ;in, Represented by two clouds, Separate clouds represent Expectation, entropy, and hyperentropy Separate clouds represent Expectation, entropy, and hyperentropy.
[0037] Based on the profit function, construct an optimization objective function for multi-task allocation with the goal of maximizing the total profit of all drones performing all tasks: max ;in, For the number of drones, The initial number of tasks. The number of tasks that are dynamically added during execution. It is a binary variable. Indicates drone Successfully executed the mission When drones Reach the task location through a series of movement actions (front, back, left, right). When the coordinates are located, the environment determination task Completed, at this point =1.
[0038] S4. Model the optimization objective function of multi-UAV task allocation as a Markov decision process to obtain the state space, action space and reward function of each UAV based on cloud model multi-attribute decision-making.
[0039] The state space of a drone includes: the drone's own attributes and mission attributes; the action space includes: front, back, left, and right.
[0040] It should be noted that although optimizing the objective function is solving for the assigned variables... However, in the Markov decision-making process, The determination is achieved by moving the drone to the mission coordinates, and the agent learns the optimal movement strategy to maximize the accumulated benefits.
[0041] Constructing a reward function based on multi-attribute decision-making in cloud models includes:
[0042] S41. Cloud representation based on all attributes of all tasks. Building a cloud decision matrix Where i and j are the indexes of the task and task attributes, respectively, and m and n are the number of tasks and the number of task attributes, respectively. For the i-th task The cloud representation of the sub-semantic corresponding to the j-th attribute, These are the expectation, entropy, and hyperentropy of the cloud representation of the sub-semantic corresponding to the j-th attribute of the i-th task, respectively.
[0043] S42, Cloud-based decision matrix An optimization function is constructed with the objectives of maximizing expectation and minimizing entropy and hyperentropy. The Lagrange multiplier method is used to solve the optimization function to obtain the optimal weights for each task attribute. ;
[0044] Optimization function: Its constraints are and , It is the weight of the j-th task attribute.
[0045] S43. Based on the optimal weight of task attribute j Cloud decision matrix Each task The task attribute j is aggregated to obtain the weighted comprehensive cloud representation matrix. ; among them, each task Weighted composite cloud representation , For the i-th task The weighted composite cloud representation of expectation, entropy, and hyperentropy;
[0046] Expected aggregation Entropy aggregation Hyperentropy aggregation .
[0047] S44, Cloud-based decision matrix Select the optimal cloud representation for each task attribute respectively. And worst cloud representation The optimal cloud representation of each task attribute is determined based on the optimal weight of the task attribute. And worst cloud representation We perform weighted aggregation separately to obtain the positive ideal cloud. and Negative Ideal Cloud ;
[0048] Optimal cloud representation Worst cloud representation .
[0049] Zhenglixiangyun Negative Ideal Cloud ;in, The optimal weight for task attribute j.
[0050] A better cloud model should have a larger expected value (representing a higher central value of revenue) and smaller entropy and hyperentropy (representing lower uncertainty and higher stability). In order to ensure that the positive ideal cloud can represent the theoretically optimal solution and avoid being limited by the current candidate task set, this invention adopts the method of constructing a virtual ideal solution (i.e., aggregating the optimal values of each attribute).
[0051] S45, according to each task Weighted composite cloud representation With Zhenglixiang Cloud Negative Ideal Cloud Distance calculation for each task Relative closeness ;
[0052] Task Relative closeness :
[0053]
[0054] in, For the task Weighted composite cloud representation With Zhenglixiang Cloud distance, For the task Weighted composite cloud representation With negative ideal cloud distance, The smaller the value, the closer the task is to the optimal choice.
[0055] S46. Select the minimum relative proximity. Based on the minimum relative proximity and their corresponding tasks Build each drone reward function ;in, The index of the task with the lowest relative proximity. For drones The task with the lowest relative similarity The distance between them For drones The distance between it and the task furthest away.
[0056] This function guides the drone to the task that yields the best reward. The reward function is dense, which can provide effective guidance signals for the reinforcement learning agent at each step, avoiding the problem of low learning efficiency caused by traditional sparse rewards.
[0057] S5. Based on the state space, action space, and reward function of the cloud model multi-attribute decision for each UAV, a deep reinforcement learning algorithm is used to obtain the optimal task allocation decision.
[0058] This invention employs a fully decentralized Independent Proximal Policy Optimization (IPPO) algorithm, with each UAV independently running a separate PPO algorithm for training. This ensures good scalability of the system to changes in the number of UAVs. Specifically, the optimal multi-task allocation decision is obtained using a deep reinforcement learning algorithm, including:
[0059] S51, Each Intelligent Agent The drone deploys an Actor network (policy network) and a Critic network (value network) locally.
[0060] S52, Each Intelligent Agent By interacting with the local environment through its own sensors, it generates experience samples. and empirical samples Cache it in the local experience pool; among which, respectively intelligent agents The state at time t, the action at time t, the state at time t+1, and the reward at time t;
[0061] The specific process includes:
[0062] intelligent agent Collects the current local status using its own sensors. (It only includes its own observable features, without a global map or the complete states of other agents).
[0063] intelligent agent The policy network is based on the current state Output action distribution, sample to obtain the current action ;
[0064] intelligent agent Execute the current action The environment returns to the next local state. and local instant rewards ;
[0065] empirical samples Store in the local experience pool.
[0066] S53, Each Intelligent Agent Experience samples are sampled from the local experience pool. Based on the sampled experience samples, the PPO algorithm is used to train the respective Actor network and Critic network, resulting in the trained Actor network and Critic network.
[0067] To encourage drones to explore a wider action space and avoid getting trapped in local optima, this invention incorporates intelligent agents. The policy entropy term was added to the loss function of the Actor network. , For entropy, Represents intelligent agents The Actor network in state The probability distribution of actions is given below. A higher policy entropy indicates a more uniform probability of the drone choosing different actions, and stronger exploratory behavior; the final loss function... ;in, It is the entropy coefficient, used to adjust the weight of policy entropy in the total loss. Its weight is dynamically adjusted during training. It is the importance sampling ratio. It is an advantage estimate. It is the cutting factor. These are the parameters of the Actor network.
[0068] To stabilize the training process of the Critic network, for each agent... The value function is clipped by adding a clipping term to the original mean squared error loss function. This restricts the current value function. The update range cannot deviate from the value function of the previous time step. Too far, final loss function ;in, For the expectation of time step t, For value function, For the parameters of the Critic network, The objective value of the value function, old parameters The value function under, For the sake of advantage estimation, For the valuation of the cropped product, For the cropping operation, This is the cutting factor.
[0069] S54. Each agent obtains its optimal task allocation decision based on its own trained Actor network and Critic network.
[0070] The verification of this invention was conducted in a 10×10 two-dimensional mesh simulation environment built on OpenAI Gym. The environment contains two types of entities: drones and tasks. The initial number and location of drones and tasks can be randomly generated. Tasks can also be dynamically added at random locations during the drones' exploration of the environment.
[0071] like Figure 2As shown, replacing the cloud model (ours) with triangular fuzzy numbers (ours-A(fuzzy)) or using a regular sparse reward function (ours-A(null)) results in a significantly lower completion rate at each time step compared to using the cloud model, demonstrating the effectiveness of the cloud model and the designed reward function.
[0072] like Figure 3 , Figure 4 As shown, compared with centralized training algorithms (CTDE) such as MAPPO, VDN, and MADDPG, the curves of the Episode Rewards and Completion Rate of this invention exhibit faster convergence speed, higher stability, and better final performance.
[0073] The above-described embodiments further illustrate the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for multi-UAV task allocation in a fuzzy dynamic environment, characterized in that, include: S1. Construct a multi-UAV task allocation model; the multi-UAV task allocation model includes UAVs and tasks; each UAV and task includes multiple attributes. S2. Use a cloud model to construct cloud representations of each drone attribute and each mission attribute; S3. Construct an optimization objective function for multi-drone task allocation based on the cloud representation of each drone attribute and each task attribute; S4. The objective function for multi-UAV task allocation is modeled as a Markov decision process to obtain the state space, action space, and reward function based on cloud model multi-attribute decision for each UAV. S5. Based on the state space, action space, and reward function of the cloud model multi-attribute decision for each UAV, a deep reinforcement learning algorithm is used to obtain the optimal task allocation decision.
2. The method for multi-UAV task allocation in a fuzzy dynamic environment according to claim 1, characterized in that, For any drone attribute or mission attribute, constructing its cloud representation includes: Obtain the valid domain of the attribute and divide the valid domain of the attribute into N sub-semantics. Calculate each sub-semantic Expectations ,entropy and hyperentropy To obtain each sub-semantic The cloud indicates Combining all sub-semantics The cloud indicates The cloud representation of the obtained attributes ;in, This is an index for sub-semantics.
3. The method for multi-UAV task allocation in a fuzzy dynamic environment according to claim 2, characterized in that, The objective function for constructing multi-UAV task allocation includes: Calculate each drone Execute each task The profit function : ; in, , , , As weight, , Tasks value Cloud representation of corresponding sub-semantics, tasks Importance The cloud representation of the corresponding sub-semantics, For drones With the task The distance between them For drones Execute the task Capability utilization value drones execution capability Cloud representation and tasks corresponding to sub-semantics execution difficulty The cloud representation of the corresponding sub-semantics, Represents the distance function; Construct an optimization objective function for multi-UAV task allocation based on the profit function: max ;in, For the number of drones, The initial number of tasks. For dynamically added tasks, It is a binary variable. Indicates drone Successfully executed the mission .
4. The multi-UAV task allocation method in a fuzzy dynamic environment according to claim 3, characterized in that, The process of constructing the reward function based on multi-attribute decision-making in cloud models includes: S41. Cloud representation based on all attributes of all tasks. Building a cloud decision matrix ;in, For the i-th task The cloud representation of the sub-semantic corresponding to the j-th attribute, For the i-th task The expectation, entropy, and hyperentropy of the cloud representation of the sub-semantic corresponding to the j-th attribute, i and j are the index of the task and the index of the task attribute, respectively, and m and n are the number of tasks and the number of task attributes, respectively. S42, Cloud-based decision matrix An optimization function is constructed with the goal of maximizing expectation and minimizing entropy and hyperentropy. The optimization function is solved using the Lagrange multiplier method to obtain the optimal weights of each task attribute. S43. Adjust the cloud decision matrix according to the optimal weights of the task attributes. Each task The task attributes are aggregated to obtain a weighted comprehensive cloud representation matrix. ;in, For the task The weighted average cloud representation; S44, Cloud-based decision matrix Select the optimal cloud representation for each task attribute respectively. And worst cloud representation The optimal cloud representation of each task attribute is determined based on the optimal weight of the task attribute. And worst cloud representation We perform weighted aggregation separately to obtain the positive ideal cloud. and Negative Ideal Cloud ; S45, according to each task Weighted composite cloud representation With Zhenglixiang Cloud Negative Ideal Cloud Distance calculation for each task Relative closeness ; S46. Select the minimum relative proximity. Based on the minimum relative proximity and their corresponding tasks Build each drone The reward function.
5. The multi-UAV task allocation method in a fuzzy dynamic environment according to claim 4, characterized in that, Optimization function: The constraints are and ;in, Let be the weight of the j-th task attribute.
6. The multi-UAV task allocation method in a fuzzy dynamic environment according to claim 4, characterized in that, Task Relative closeness : ; in, For the task Weighted composite cloud representation With Zhenglixiang Cloud distance, For the task Weighted composite cloud representation With negative ideal cloud The distance.
7. The multi-UAV task allocation method in a fuzzy dynamic environment according to claim 4, characterized in that, drones reward function ;in, For drones The task with the lowest relative similarity The distance between them For drones The distance between it and the task furthest away.
8. The method for multi-UAV task allocation in a fuzzy dynamic environment according to claim 1, characterized in that, The optimal task allocation decision obtained using deep reinforcement learning algorithms includes: Each agent Deploy an Actor network and a Critic network locally; the agents are drones. Each agent Experience samples are generated through interaction with the local environment. and empirical samples Cache it in the local experience pool; among which, respectively intelligent agents The state at time t, the action at time t, the state at time t+1, and the reward at time t; Each agent Experience samples are sampled from the local experience pool, and the respective Actor network and Critic network are trained based on the sampled experience samples to obtain the respective trained Actor network and Critic network. Each agent Each network, based on its trained Actor and Critic networks, yields its optimal task allocation decision.
9. A method for multi-UAV task allocation in a fuzzy dynamic environment according to claim 8, characterized in that, intelligent agent When training its own Actor network, a policy entropy term is added to the loss function of the Actor network. ;in, For entropy, Indicates the state of the Actor network. The probability distribution of the next action.
10. A method for multi-UAV task allocation in a fuzzy dynamic environment according to claim 8, characterized in that, intelligent agent When training its own Critic network, a pruning term is added to the loss function of the Critic network. The final loss function of the Critic network is obtained. ;in, For the expectation of time step t, For value function, For the parameters of the Critic network, The objective value of the value function, For the cropping operation, This is the cutting factor.