Heterogeneous aircraft cluster multi-scale cross-domain autonomous confrontation decision-making method
Through the distributed expansion consensus auction algorithm and multi-agent near-end strategy optimization algorithm, the problem of low task allocation efficiency of heterogeneous aircraft groups is solved, and efficient collaborative operation in complex dynamic environments is achieved.
Patent Information
- Application Number
- CN202510795538.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-16
AI Technical Summary
In the prior art, the task allocation efficiency of heterogeneous aircraft groups is low, and the communication burden is heavy, making it difficult to adapt to the dynamic environment. The performance differences of heterogeneous aircraft are not fully considered, resulting in inefficient synergistic strategies.
The distributed expansion consensus auction algorithm and multi-agent near-end strategy optimization algorithm are adopted to achieve reinforcement learning of heterogeneous aircraft clusters by establishing an adversarial task set, determining opponent goals, and using distributed task allocation modules.
It reduces the communication burden, improves the task allocation efficiency, and supports the coordinated operation of heterogeneous aircraft in complex dynamic environments.
Smart Images

Figure CN120353241A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent systems and unmanned aerial vehicles, and specifically relates to a multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters. Background Art
[0002] With the rapid innovation of unmanned aerial vehicle technology, the application boundaries of aircraft clusters in scenarios continue to expand. However, the existing technologies have the following deficiencies in the cooperative task allocation and control of heterogeneous aircraft: 1) Most of the task allocation algorithms are based on centralized computing, with heavy communication burdens and difficult to adapt to dynamic environments; 2) The performance differences of heterogeneous aircraft are not fully considered, resulting in low allocation efficiency.
[0003] Therefore, there is an urgent need for a distributed and highly robust autonomous decision-making method. Summary of the Invention
[0004] The present invention provides a multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, aiming to solve the problems of low task allocation efficiency of heterogeneous aircraft groups in the existing technology, and the problem that the multi-task cooperation strategy of heterogeneous aircraft can be further optimized.
[0005] The present invention is realized through the following technical solutions: A multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, the method comprising the following steps: Step 1: Establish an adversarial task set for the heterogeneous aircraft cluster; Step 2: Determine the opponent's targets based on the adversarial task set in Step 1; Step 3: Based on the opponent's targets in Step 2, obtain the allocation of task-target pairs using a distributed task allocation module; Step 4: Based on the task-target pairs in Step 3, through cooperative strategy optimization, realize the reinforcement learning of the heterogeneous aircraft cluster.
[0006] Further, the heterogeneous aircraft cluster in Step 1 specifically includes multiple types of aircraft, and the performance parameters of each type of aircraft include reconnaissance ability, endurance ability, and interference ability; The adversarial task set in Step 1 includes search tasks, encirclement tasks, and tracking tasks, and each task has a defined location, ability requirements, and priority.
[0007] Further, the opponent's targets in Step 2 are specifically that some positions are known, and the positions and threat levels are obtained through dynamic updates.
[0008] Further, the distributed task allocation module in step 3 is specifically as follows: The extended consensus auction algorithm is adopted, comprehensively considering task adaptability, aircraft performance, target characteristics, and environmental factors, and calculating the task-target pairs through the value function in the auction phase and the allocation in the consensus phase.
[0009] Further, the extended consensus auction algorithm specifically includes the following steps: Step 3.1: In the auction phase, each aircraft traverses the task-target pairs and generates a bid list based on the value function, where the value function is comprehensive task adaptability, target matching, distance efficiency, and priority. Step 3.2: In the consensus phase, obtain the number of opponent targets, obtain the task-target allocation scale through rule-driven, the aircraft broadcasts the bid list through the local communication network, receives the bid information of neighboring aircraft, updates the consensus table, retains the group of aircraft with the highest bid, and iterates multiple times to achieve task allocation consistency.
[0010] Further, step 3.1 in the auction phase is specifically as follows: Step 3.1.1: Each aircraft defines a matching matrix according to task adaptability and calculates task adaptability in combination with aircraft attributes. Step 3.1.2: Determine the minimum sensor range and endurance required for the task according to target matching. Step 3.1.3: Calculate the distance efficiency between the aircraft and the target. Step 3.1.4: Calculate the priority according to the mobility and importance of the target. Step 3.1.5: Generate a bid list and select the task-target pair that maximizes the global reward.
[0011] Further, step 3.2 in the consensus phase is specifically as follows: Step 3.2.1: Obtain the number of opponent targets and obtain the task-target allocation scale through rule-driven. Step 3.2.2: The aircraft broadcasts the bid list, including the task-target pair and the aircraft number, to the neighbors within the communication range through the local communication network. Step 3.2.3: Receive the bid list of neighboring aircraft and update the local bid list. Step 3.2.4: Check the task allocation status and determine the set of unfilled task-target pairs. Step 3.2.5: Re-select the value, retain the group of aircraft with the highest bid, and resolve the allocation conflict. Step 3.2.6: Check whether the consistency converges. If it does not converge, repeat steps 3.2.1 to 3.2.4, using the maximum number of iterations as the termination condition or until the consistency converges.
[0012] Furthermore, the collaborative strategy optimization in step 4 is specifically as follows: based on the proximal policy optimization of the heterogeneous aircraft cluster, with the task-goal pair, aircraft state, and environmental state as inputs, reward functions and loss functions are designed for different tasks to optimize the aircraft cooperation strategy.
[0013] Furthermore, the multi-agent proximal policy optimization algorithm of the collaborative optimization module includes the following steps: Step 4.1: Taking the task-goal pair, aircraft state, and environmental state as inputs, output the expected position and attitude of the UAV. For the search and tracking task, the action space includes movement and scanning, and for the encirclement task, the action space includes movement and capture; Step 4.2: For the search task, design a reward function to maximize the coverage rate and a loss function to optimize the sensor data acquisition efficiency; Step 4.3: For the encirclement task, design a reward function to increase the success rate of encirclement and a loss function to reduce interference and conflict; Step 4.4: For the tracking task, design a reward function to improve the continuous tracking accuracy and a loss function to optimize the endurance efficiency.
[0014] A multi-scale cross-domain autonomous confrontation decision-making system for heterogeneous aircraft clusters. The system uses the above-mentioned multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters. The system includes: Adversarial task set establishment module: Establish an adversarial task set for the heterogeneous aircraft cluster; Based on the adversarial task set in step 1, determine the opponent's target; Distributed task allocation module: Based on the opponent's target, use the extended consensus auction algorithm to obtain the allocation of the task-goal pair; Collaborative strategy optimization module: Based on the allocation of the task-goal pair, through the multi-agent proximal policy optimization algorithm of the collaborative optimization module, realize the reinforcement learning of the heterogeneous aircraft cluster.
[0015] The beneficial effects of the present invention are: The present invention reduces the communication burden and improves the task allocation efficiency through the distributed extended consensus auction algorithm.
[0016] The present invention uses the multi-agent proximal policy optimization algorithm for reinforcement learning to optimize the multi-task collaboration performance.
[0017] The present invention supports the collaboration of heterogeneous aircraft and adapts to complex dynamic environments. Description of the Drawings
[0018] Figure 1 It is the method flow chart of the present invention.
[0019] Figure 2It is a block diagram of the extended consensus auction algorithm of the present invention.
[0020] Figure 3 It is a flowchart of the multi-agent proximal policy optimization algorithm of the present invention. Specific embodiments
[0021] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0022] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0023] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the present application specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0025] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application, but the present application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.
[0026] Embodiment 1 An embodiment of the present invention provides a multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, which realizes efficient task allocation and heterogeneous aircraft cluster cooperation through an extended consensus auction algorithm and a multi-agent proximal policy optimization algorithm. As Figure 1 shown, the method includes the following steps: Step 1: Establish an adversarial task set for the heterogeneous aircraft cluster; Step 2: Based on the adversarial task set in Step 1, determine the opponent's goals; Step 3: Based on the opponent's goals in Step 2, use the distributed task allocation module to obtain the allocation of task-goal pairs; Step 4: Based on the task-goal pairs in Step 3, through collaborative policy optimization, achieve reinforcement learning for the heterogeneous aircraft cluster.
[0027] Furthermore, the heterogeneous aircraft cluster in Step 1 specifically includes various types of aircraft, and each type of aircraft has specific performance parameters including reconnaissance ability, endurance ability, and interference ability; The adversarial task set in Step 1 includes search tasks , encirclement tasks and tracking tasks , and each task has a defined location, ability requirements, and priority.
[0028] Furthermore, the opponent's goals in Step 2 specifically mean that some positions are known, and the positions and threat levels are obtained through dynamic updates.
[0029] Furthermore, the distributed task allocation module in Step 3 specifically uses the extended consensus auction algorithm, comprehensively considering task adaptability, aircraft performance, target characteristics, and environmental factors, and calculating the value function in the auction stage and allocating task-goal pairs in the consensus stage.
[0030] As Figure 2 shown, furthermore, the extended consensus auction algorithm specifically includes the following steps: Step 3.1: In the auction stage, each aircraft traverses the task-goal pairs (i.e., calculates the value function corresponding to the auction stage), generates a bid list based on the value function, and the value function is comprehensive task adaptability, target matching, distance efficiency, and priority; Step 3.2: In the consensus stage, obtain the number of opponent's goals, obtain the task-goal allocation scale through rule-driven, the aircraft broadcasts the bid list through the local communication network, receives the bid information of neighbor aircraft, updates the consensus table, retains the aircraft group with the highest bid, and iterates multiple times to achieve task allocation consistency.
[0031] Furthermore, Step 3.1 in the auction stage specifically is: Step 3.1.1: Each aircraft defines a matching matrix according to task adaptability, and calculates task adaptability in combination with aircraft attributes; Step 3.1.2: According to target matching, determine the minimum sensor range and endurance time required for the task; Step 3.1.3: Calculate the distance efficiency between the aircraft and the target; Step 3.1.4: Calculate the priority according to the mobility and importance of the target; Step 3.1.5: Generate a bid list and select the task-goal pair that maximizes the global reward.
[0032] Furthermore, step 3.2 in the consistency phase is specifically as follows: Step 3.2.1: Obtain the number of opponent's goals and obtain the task-goal allocation scale through rule-driven; The aircraft broadcasts the bid list including the task-goal pair and the aircraft number to the neighbors within the communication range through the local communication network; Receive the bid lists of neighboring aircraft and update the local bid list; Check the task allocation status and determine the set of task-goal pairs that are not fully allocated; Re-perform value selection, retain the group of aircraft with the highest bid, and resolve the allocation conflict; Check whether the consistency converges. If it does not converge, repeat steps 3.2.1 to 3.2.4, with the maximum number of iterations as the termination condition or until the consistency converges, that is, stop the loop when reaching the maximum number of steps in the case of non-convergence.
[0033] Furthermore, the collaborative strategy optimization in step 4 is specifically based on proximal policy optimization of heterogeneous aircraft clusters (multi-agent), with the task-goal pair, aircraft state, and environmental state as inputs, designing reward functions and loss functions for different tasks, and optimizing the aircraft cooperation strategy.
[0034] As Figure 3 shown, furthermore, the multi-agent proximal policy optimization algorithm of the collaborative optimization module includes the following steps: Step 4.1: Take the task-goal pair, aircraft state, and environmental state as inputs, and output the expected position and attitude of the UAV. For the search and tracking task, the action space includes movement and scanning, and for the encirclement task, the action space includes movement and capture; For the search task, design a reward function to maximize the coverage rate and a loss function to optimize the sensor data acquisition efficiency; For the encirclement task, design a reward function to improve the encirclement success rate and a loss function to reduce interference conflicts; For the tracking task, design a reward function to improve the continuous tracking accuracy and a loss function to optimize the endurance efficiency.
[0035] The specific embodiment is as follows: System Composition: This system consists of 200 heterogeneous aircraft, labeled U1 to U200, which are divided into four types: robotic arm aircraft, flapping wing aircraft, fixed-wing aircraft, and quadrotor aircraft. The adversarial task set includes search, encirclement, and tracking, targeting 10 enemy drones.
[0036] Task allocation uses an extended consensus auction algorithm, which is divided into the following steps: Auction Phase: The aircraft Calculate bids for all tasks and targets according to the value function. The goal of the value function is to quantify the benefits of the aircraft performing tasks against enemy targets, in the form of a triple indicating that the aircraft performs tasks against enemy targets. Considering task suitability, aircraft performance, enemy target characteristics, and environmental factors, ensure that the allocation result maximizes the global reward;
[0037] Among them, is the bid of the aircraft for the task-target pair (t, g), is the task suitability coefficient, is the task suitability function, is the target matching coefficient, is the target matching function, is the flight efficiency coefficient, is the distance function, is the priority coefficient, is the importance function; For task suitability, define the matching matrix and the attributes of each aircraft ,
[0038] Among them, is the maximum aircraft attribute, is the matching matrix, is the attribute of each aircraft, is the weight, i is, n is is the weighted normalization term, and n is the total number of aircraft; For target matching, define the minimum sensor range required for the task and the minimum endurance time ,
[0039] Among them, is the importance weight, is the aircraft speed, is the sensor range, is the minimum sensor range, is the minimum endurance time, is the endurance time; Calculate the distance efficiency, the distance between the aircraft and the target ,
[0040] where, is the target reference position, is the maximum distance threshold; Calculate the priority, considering the mobility of the enemy target and the importance of the target
[0041]
[0042] where, is the target mobility, is the importance coefficient; Each aircraft traverses all task - target pairs , generates a bid list , using the valid task - target list :
[0043] where, is the bid, is the lowest bid in the list except , select the task - target pair that maximizes the reward :
[0044] where, is the task, is the target, is the bid of the aircraft for the task - target pair (t, g); In the consistency phase, obtain the number of opponent's targets and the number of our aircraft , calculate the task - target allocation scale through rule - driven
[0045]
[0046] The average ability of each aircraft is estimated through real - time status, and the proportion of reserve resources reflects the current resource occupancy situation.
[0047] Design constraint conditions for task requirements to limit the maximum aircraft requirements and the minimum aircraft requirements ,
[0048] Among them, is the task-goal allocation scale, j is the task-goal pair; The aircraft exclusive constraint ensures that each aircraft is assigned only one task-goal pair,
[0049] To ensure that the total number of allocated aircraft does not exceed the number of available aircraft ,
[0050] Among them, is the total number of allocated aircraft; Solve task allocation conflicts by exchanging bid information with neighbors through the local communication network; The aircraft broadcasts bids to neighbors within the communication range , the task-goal pair selection and the aircraft number .
[0051]
[0052] Receive the bid list of neighbors and update the local bid list :
[0053]
[0054]
[0055] Among them, is the set of neighbor aircraft bids, is the temporary bid list of aircraft i at time t+1, is the bid list of aircraft i at time t, is the set of bids of each aircraft i for task-goal pair j, is the bid of the neighbor aircraft, () is to merge multiple bid sources, is the neighbor aircraft; Check the allocation status, update the available task-goal pairs, and determine the set of unfilled task-goal pairs , re - conduct value selection,
[0056] For each task, retain the aircraft with the highest bid for assignment, re - assign aircraft after conflict resolution
[0057]
[0058] Among them, is the task assignment of aircraft i at time t + 1, is the aircraft with the highest bid, is the winner set of task - target pair j at time t + 1; Check whether the consistency converges
[0059] Among them, is the winner set of task - target pair j, is the winner set of task - target pair j in the previous round; In the collaborative strategy optimization part, the collaborative strategy is trained based on the multi - agent proximal policy optimization algorithm. The inputs include task - target pairs, aircraft states including position, speed, sensor data, and environmental states. The action space includes moving, scanning, capturing, etc.; The reward function and loss function are designed for different tasks: Search task : The reward function encourages maximizing the coverage rate, and the loss function optimizes the sensor data acquisition efficiency. Task Reward function:
[0060] Among them, is the reward function at time t, is the target position, is the minimum distance, is the intra - group distance, is the aircraft position, is the penalty variable; Loss function
[0061] Among them, is the policy loss function of task , is the empirical estimate expectation, is the advantage function, is the reward ratio, is the clipping range, is a policy parameter, is a clipping operation; is a task value function loss, is a value function, is a state, is to make a decision, is a discounted return, is a discount factor, is a single-step reward; Surrounding task : The reward function encourages the surrounding success rate, and the loss function reduces interference conflicts.
[0062] Task Reward function:
[0063] Among them, is the reward function at the t-th moment, is the state covariance, is the observation information, is a weight parameter, is a penalty variable; Loss function
[0064] Among them, is the policy loss function, is the advantage function, is the value function loss function, is the value function, is the state, is the action, is the discounted return, is the covariance regularization weight, is the variance term of the state covariance; Tracking task : The reward function encourages continuous tracking accuracy, and the loss function optimizes the endurance efficiency.
[0065] Task Reward function:
[0066] Among them, is the reward function at the t-th moment, is the energy loss; Loss function
[0067] Among them, is the policy loss function, is the advantage function, is the value function loss function, is the value function, is the state, is the action, is the discounted return, is the energy loss, is the regularization weight; The expected speed of the optimized aircraft is output during training After that, considering that the expected speed output by reinforcement learning may have noise or mutations, a moving average is used to smooth the speed sequence, and then the time step is dynamically adjusted, and dynamically adjusted according to the distance from the target to implement a collaborative strategy and improve the task completion efficiency.
[0068] Embodiment 2 An embodiment of the present invention provides a heterogeneous aircraft cluster multi-scale cross-domain autonomous confrontation decision-making system. The system uses the heterogeneous aircraft cluster multi-scale cross-domain autonomous confrontation decision-making method as described in Embodiment 1. The system includes: Adversarial task set establishment module: Establish an adversarial task set for the heterogeneous aircraft cluster; Based on the adversarial task set in step 1, determine the opponent's target; Distributed task allocation module: Based on the opponent's target, use the extended consensus auction algorithm to obtain the allocation of task-target pairs; Collaborative strategy optimization module: Based on the allocation of task-target pairs, through the multi-agent proximal policy optimization algorithm of the collaborative optimization module, realize the reinforcement learning of the heterogeneous aircraft cluster.
Claims
1. A multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, characterized in that The method includes the following steps: Step 1: Establish an adversarial task set for the heterogeneous aircraft cluster; Step 2: Determine the opponent's targets based on the adversarial task set in Step 1; Step 3: Based on the opponent's targets in Step 2, use the distributed task allocation module to obtain the allocation of task-target pairs; Specifically, the extended consensus auction algorithm is adopted. In the auction phase, each aircraft traverses the task-target pairs and generates a bid list based on a value function, where the value function is a combination of task suitability, target matching, distance efficiency, and priority; Each aircraft defines a matching matrix according to task suitability and calculates the task suitability in combination with the aircraft attributes; Determine the minimum sensor range and endurance required for the task according to target matching; Calculate the distance efficiency between the aircraft and the target; Calculate the priority according to the maneuverability and importance of the target; Generate a bid list and select the task-target pair that maximizes the global reward; Step 4: Based on the task-target pairs in Step 3, through collaborative policy optimization, achieve reinforcement learning of the heterogeneous aircraft cluster.
2. The method according to claim 1, wherein The heterogeneous aircraft cluster in Step 1 specifically includes multiple types of aircraft, and the performance parameters of each type of aircraft include reconnaissance ability, endurance ability, and interference ability; The adversarial task set in Step 1 includes search tasks, encirclement tasks, and tracking tasks, and each task has a defined location, ability requirements, and priority.
3. The method according to claim 1, wherein The opponent's targets in Step 2 specifically mean that some positions are known, and the positions and threat levels are obtained through dynamic updates.
4. The method according to claim 1, characterized in that The distributed task allocation module in Step 3 specifically means that based on the extended consensus auction algorithm, comprehensively considering task suitability, aircraft performance, target characteristics, and environmental factors, calculate the task-target pairs through the value function in the auction phase and allocate them in the consensus phase.
5. The method according to claim 4, characterized in that The extended consensus auction algorithm specifically is: In the consensus phase, obtain the number of opponent's targets, obtain the task-target allocation scale through rule-driven. The aircraft broadcasts the bid list through the local communication network, receives the bid information of neighboring aircraft, updates the consensus table, retains the group of aircraft with the highest bid, and iterates multiple times to achieve task allocation consistency.
6. The method according to claim 5, wherein The specific consensus phase is: Step 3.1: Obtain the number of opponent's targets and obtain the task-target allocation scale through rule-driven; Step 3.2: The aircraft broadcasts the bid list, including task-target pairs and aircraft numbers, to the neighbors within the communication range through the local communication network; Step 3.3: Receive the bid list of neighboring aircraft and update the local bid list; Step 3.4: Check the task allocation status and determine the set of task-target pairs that are not fully allocated; Step 3.5: Re-select the value, retain the group of aircraft with the highest bid, and resolve the allocation conflict; Step 3.6: Check whether the consistency converges. If it does not converge, repeat Steps 3.1 to 3.4, using the maximum number of iterations as the termination condition or until the consistency converges.
7. The method according to claim 4, wherein The collaborative policy optimization in Step 4 specifically means that based on proximal policy optimization of the heterogeneous aircraft cluster, using task-target pairs, aircraft states, and environmental states as inputs, design reward functions and loss functions for different tasks to optimize the aircraft cooperation strategy.
8. The method according to claim 7, wherein The multi-agent proximal policy optimization algorithm of the collaborative optimization module includes the following steps: Step 4.1: Taking task-goal pairs, aircraft states, and environmental states as inputs, output the desired positions and attitudes of the UAVs. For the search and tracking tasks, the action space includes moving and scanning, and for the encirclement task, the action space includes moving and capturing; Step 4.2: For the search task, design a reward function to maximize the coverage rate and a loss function to optimize the sensor data acquisition efficiency; Step 4.3: For the encirclement task, design a reward function to improve the encirclement success rate and a loss function to reduce interference and conflicts; Step 4.4: For the tracking task, design a reward function to improve the continuous tracking accuracy and a loss function to optimize the endurance efficiency.
Citation Information
Patent Citations
Multi-robot task allocation method, electronic equipment and medium
CN116757420A
Air-ground unmanned system task matching and distribution method based on improved auction mechanism and pigeon flock evolutionary algorithm
CN117592751A
Multi-unmanned aerial vehicle real-time distributed task allocation and reallocation method
CN117892933A
Distributed auction algorithm-based heterogeneous multi-target multi-unmanned ship task allocation method
CN117930852A
Port multi-agent task allocation method and system based on reinforcement learning
CN119273104A