Multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft swarms

Through the distributed expansion consensus auction algorithm and multi-agent near-end strategy optimization algorithm, the communication burden and inefficiency in heterogeneous aircraft collaborative mission allocation are solved, and efficient task allocation and independent decision-making of heterogeneous aircraft clusters are achieved.

CN120353241BActive Publication Date: 2025-09-02TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510795538.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-02
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The prior art has problems of heavy communication burden and low allocation efficiency in the allocation of heterogeneous aircraft coordinated missions, and the differences in aircraft performance are not fully considered.

Method used

The distributed expansion consensus auction algorithm and multi-agent near-end strategy optimization algorithm are adopted to realize reinforcement learning of heterogeneous aircraft clusters by establishing an adversarial task set, determining opponent goals, assigning task-target pairs, and performing collaborative strategy optimization.

Benefits of technology

It reduces the communication burden, improves the task allocation efficiency, and supports the coordinated operation of heterogeneous aircraft in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353241B_ABST
    Figure CN120353241B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of multi-agent systems and unmanned aerial vehicle technology, and specifically relates to a multi-scale, cross-domain autonomous confrontation decision-making method for a heterogeneous aircraft cluster. Step 1: Establish a confrontation task set for a heterogeneous aircraft cluster; Step 2: Determine the opponent's target based on the confrontation task set in step 1; Step 3: Based on the opponent's target in step 2, use a distributed task allocation module to obtain the allocation of task-target pairs; Step 4: Based on the task-target pairs in step 3, implement reinforcement learning for the heterogeneous aircraft cluster through collaborative strategy optimization. The present invention is used to solve the problem of low task allocation efficiency of heterogeneous aircraft clusters in the prior art, and the problem that the multi-task collaborative strategy of heterogeneous aircraft can be further optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-agent systems and unmanned aerial vehicles, and specifically relates to a multi-scale cross-domain autonomous confrontation decision-making method for a heterogeneous aircraft cluster. Background Art

[0002] With the rapid innovation of drone technology, the application boundaries of aircraft swarms in various scenarios continue to expand. However, existing technologies have the following shortcomings in the collaborative task allocation and control of heterogeneous aircraft:

[0003] 1) Task allocation algorithms are mostly based on centralized computing, which has a heavy communication burden and is difficult to adapt to dynamic environments;

[0004] 2) The performance differences of heterogeneous aircraft are not fully considered, resulting in inefficient allocation.

[0005] Therefore, a distributed and robust autonomous decision-making method is urgently needed. Summary of the Invention

[0006] The present invention provides a multi-scale cross-domain autonomous confrontation decision-making method for a heterogeneous aircraft cluster, which is used to solve the problem of low efficiency in task allocation of heterogeneous aircraft clusters in the existing technology, and the problem that the multi-task coordination strategy of heterogeneous aircraft can be further optimized.

[0007] The present invention is achieved through the following technical solutions:

[0008] A multi-scale cross-domain autonomous confrontation decision-making method for a heterogeneous aircraft cluster, the method comprising the following steps:

[0009] Step 1: Establish a set of adversarial tasks for a heterogeneous aircraft cluster;

[0010] Step 2: Based on the adversarial task set in step 1, determine the opponent's goal;

[0011] Step 3: Based on the opponent's goal in step 2, the distributed task allocation module is used to obtain the assignment of task-goal pairs;

[0012] Step 4: Based on the task-goal pairs in step 3, reinforcement learning of heterogeneous aircraft clusters is achieved through collaborative strategy optimization.

[0013] Furthermore, the heterogeneous aircraft cluster in step 1 specifically includes multiple types of aircraft, and the performance parameters of each type of aircraft include reconnaissance capability, endurance capability, and jamming capability;

[0014] The confrontation task set of step 1 includes a search task, a roundup task, and a tracking task, each task having a defined location, capability requirement, and priority.

[0015] Furthermore, the opponent target in step 2 is specifically that part of the position is known, and the position and threat level are obtained through dynamic updating.

[0016] Furthermore, the distributed task allocation module of step 3 specifically adopts an extended consensus auction algorithm, comprehensively considers task adaptability, aircraft performance, target characteristics and environmental factors, and allocates task-target pairs through value function calculation in the auction stage and consistency stage.

[0017] Furthermore, the extended consensus auction algorithm specifically includes the following steps:

[0018] Step 3.1: In the auction phase, each aircraft traverses the mission-target pairs and generates a bid list based on a value function that integrates mission suitability, target matching, distance efficiency, and priority.

[0019] Step 3.2: In the consistency phase, the number of adversary targets is obtained, and the task-target allocation scale is obtained through rule-driven methods. The aircraft broadcasts the bid list through the local communication network, receives bid information from neighboring aircraft, updates the consistency table, and retains the aircraft group with the highest bid. Multiple iterations are performed to achieve task allocation consistency.

[0020] Furthermore, step 3.1 of the auction stage is specifically as follows:

[0021] Step 3.1.1: Define a matching matrix for each aircraft based on its mission suitability and calculate mission suitability based on the aircraft attributes.

[0022] Step 3.1.2: Determine the minimum sensor range and endurance required for the mission based on target compatibility.

[0023] Step 3.1.3: Calculate the range efficiency between the aircraft and the target;

[0024] Step 3.1.4: Calculate the priority based on the mobility and importance of the target;

[0025] Step 3.1.5: Generate a list of bids and select the task-target pair that maximizes the global reward.

[0026] Furthermore, step 3.2 of the consistency phase is specifically as follows:

[0027] Step 3.2.1: Obtain the number of adversary targets and obtain the task-target allocation scale through rule-driven methods;

[0028] Step 3.2.2: The aircraft broadcasts a bid list, including the mission-target pair and aircraft number, to neighbors within the communication range via the local communication network.

[0029] Step 3.2.3: Receive the bid list of neighboring aircraft and update the local bid list;

[0030] Step 3.2.4: Check the task allocation status and determine the set of unfilled task-target pairs;

[0031] Step 3.2.5: Reselect the value, retain the aircraft group with the highest bid, and resolve the allocation conflict;

[0032] Step 3.2.6: Check whether the consistency has converged. If not, repeat steps 3.2.1 to 3.2.4 with the maximum number of iterations as the termination condition or until the consistency converges.

[0033] Furthermore, the collaborative strategy optimization in step 4 is specifically based on the proximal strategy optimization of the heterogeneous aircraft cluster, taking the task-target pair, aircraft state and environment state as input, designing reward functions and loss functions for different tasks, and optimizing the aircraft collaborative strategy.

[0034] Furthermore, the multi-agent proximal strategy optimization algorithm of the collaborative optimization module includes the following steps:

[0035] Step 4.1: Taking the mission-target pair, the aircraft state, and the environment state as input, output the desired position and attitude of the UAV. For the search and tracking mission, the action space includes movement and scanning; for the roundup mission, the action space includes movement and capture.

[0036] Step 4.2: For the search task, design a reward function to maximize coverage and a loss function to optimize sensor data acquisition efficiency;

[0037] Step 4.3: For the encirclement task, design a reward function to improve the encirclement success rate and a loss function to reduce interference conflicts;

[0038] Step 4.4: For the tracking task, design a reward function to improve continuous tracking accuracy and a loss function to optimize endurance efficiency.

[0039] A heterogeneous aircraft swarm multi-scale cross-domain autonomous confrontation decision-making system, the system using the above-mentioned heterogeneous aircraft swarm multi-scale cross-domain autonomous confrontation decision-making method, the system comprising:

[0040] Adversarial task set establishment module: establishes an adversarial task set for heterogeneous aircraft clusters;

[0041] Based on the adversarial task set in step 1, determine the opponent's target;

[0042] Distributed task allocation module: Based on the opponent's goals, the extended consensus auction algorithm is used to obtain the allocation of task-goal pairs;

[0043] Collaborative Strategy Optimization Module: Based on the allocation of task-goal pairs, the collaborative optimization module's multi-agent proximal strategy optimization algorithm is used to achieve reinforcement learning of heterogeneous aircraft clusters.

[0044] The beneficial effects of the present invention are:

[0045] The present invention reduces the communication burden and improves the efficiency of task allocation through a distributed extended consensus auction algorithm.

[0046] The present invention uses a multi-agent proximal strategy optimization algorithm to perform reinforcement learning to optimize multi-task collaborative performance.

[0047] The present invention supports the collaboration of heterogeneous aircraft and adapts to complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flow chart of the method of the present invention.

[0049] Figure 2 This is a block diagram of the extended consensus auction algorithm of the present invention.

[0050] Figure 3 It is a flow chart of the multi-agent proximal strategy optimization algorithm of the present invention. DETAILED DESCRIPTION

[0051] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present application with unnecessary details.

[0052] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0053] It should also be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0054] The following is a clear and complete description of the technical solutions in the embodiments of this application in conjunction with the drawings in the specification of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0056] Implementation Method 1

[0057] The embodiment of the present invention provides a multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, which realizes efficient task allocation and heterogeneous aircraft cluster collaboration by expanding the consensus auction algorithm and the multi-agent proximal strategy optimization algorithm. Figure 1 As shown, the method includes the following steps:

[0058] Step 1: Establish a set of adversarial tasks for a heterogeneous aircraft cluster;

[0059] Step 2: Based on the adversarial task set in step 1, determine the opponent's goal;

[0060] Step 3: Based on the opponent's goal in step 2, the distributed task allocation module is used to obtain the assignment of task-goal pairs;

[0061] Step 4: Based on the task-goal pairs in step 3, reinforcement learning of heterogeneous aircraft clusters is achieved through collaborative strategy optimization.

[0062] Furthermore, the heterogeneous aircraft cluster of step 1 specifically includes multiple types of aircraft, each type of aircraft having specific performance parameters including reconnaissance capability, endurance capability, and jamming capability;

[0063] The adversarial task set in step 1 includes the search task , roundup mission and tracking tasks ,Each task has a defined location, capacity requirements, and priority.

[0064] Furthermore, the opponent target in step 2 is specifically that part of the position is known, and the position and threat level are obtained through dynamic updating.

[0065] Furthermore, the distributed task allocation module of step 3 specifically adopts an extended consensus auction algorithm, comprehensively considers task adaptability, aircraft performance, target characteristics and environmental factors, and allocates task-target pairs through value function calculation in the auction stage and consistency stage.

[0066] like Figure 2 As shown, further, the extended consensus auction algorithm specifically includes the following steps:

[0067] Step 3.1: During the auction phase, each aircraft traverses the mission-target pairs (i.e., the calculated value function corresponding to the auction phase) and generates a bid list based on the value function, which is a comprehensive combination of mission suitability, target matching, distance efficiency, and priority.

[0068] Step 3.2: In the consistency phase, the number of adversary targets is obtained, and the task-target allocation scale is obtained through rule-driven methods. The aircraft broadcasts the bid list through the local communication network, receives bid information from neighboring aircraft, updates the consistency table, and retains the aircraft group with the highest bid. Multiple iterations are performed to achieve task allocation consistency.

[0069] Furthermore, step 3.1 of the auction stage is specifically as follows:

[0070] Step 3.1.1: Define a matching matrix for each aircraft based on its mission suitability and calculate mission suitability based on the aircraft attributes.

[0071] Step 3.1.2: Determine the minimum sensor range and endurance required for the mission based on target compatibility.

[0072] Step 3.1.3: Calculate the range efficiency between the aircraft and the target;

[0073] Step 3.1.4: Calculate the priority based on the mobility and importance of the target;

[0074] Step 3.1.5: Generate a list of bids and select the task-target pair that maximizes the global reward.

[0075] Furthermore, step 3.2 of the consistency phase is specifically as follows:

[0076] Step 3.2.1: Obtain the number of adversary targets and obtain the task-target allocation scale through rule-driven methods;

[0077] Step 3.2.2: The aircraft broadcasts a bid list, including the mission-target pair and aircraft number, to neighbors within the communication range via the local communication network.

[0078] Step 3.2.3: Receive the bid list of neighboring aircraft and update the local bid list;

[0079] Step 3.2.4: Check the task allocation status and determine the set of unfilled task-target pairs;

[0080] Step 3.2.5: Reselect the value, retain the aircraft group with the highest bid, and resolve the allocation conflict;

[0081] Step 3.2.6: Check whether the consistency has converged. If not, repeat steps 3.2.1 to 3.2.4, using the maximum number of iterations as the termination condition or until the consistency converges. If not, stop the loop when the maximum number of steps is reached.

[0082] Furthermore, the collaborative strategy optimization in step 4 is specifically based on the proximal strategy optimization of the heterogeneous aircraft cluster (multi-agent), taking the task-target pair, aircraft state and environment state as input, designing reward functions and loss functions for different tasks, and optimizing the aircraft collaborative strategy.

[0083] like Figure 3 As shown, further, the multi-agent proximal strategy optimization algorithm of the collaborative optimization module includes the following steps:

[0084] Step 4.1: Taking the mission-target pair, the aircraft state, and the environment state as input, output the desired position and attitude of the UAV. For the search and tracking mission, the action space includes movement and scanning; for the roundup mission, the action space includes movement and capture.

[0085] Step 4.2: For the search task, design a reward function to maximize coverage and a loss function to optimize sensor data acquisition efficiency;

[0086] Step 4.3: For the encirclement task, design a reward function to improve the encirclement success rate and a loss function to reduce interference conflicts;

[0087] Step 4.4: For the tracking task, design a reward function to improve continuous tracking accuracy and a loss function to optimize endurance efficiency.

[0088] The specific embodiments are:

[0089] The system consists of 200 heterogeneous aircraft, labeled U1 to U200, divided into four types: robotic aircraft, flapping-wing aircraft, fixed-wing aircraft, and quadrotor aircraft. The confrontation mission set includes search, capture, and tracking, targeting 10 enemy drones.

[0090] Task allocation uses the extended consensus auction algorithm, which is divided into the following steps:

[0091] Auction Phase: Aircraft According to the value function for all tasks and goals Calculate bid The goal of the value function is to quantify the benefits of the aircraft performing the mission against the enemy target, in the form of a triple , indicating that the aircraft is performing a mission against an enemy target. Taking into account mission suitability, aircraft performance, enemy target characteristics, and environmental factors, the allocation result is guaranteed to maximize the global reward;

[0092]

[0093] in, is the bid of the aircraft for the task-target pair (t,g), is the task adaptability coefficient, is the task adaptation function, is the target matching coefficient, is the target matching function, is the flight efficiency coefficient, is the distance function, is the priority coefficient, is the importance function;

[0094] For task adaptability, define the matching matrix and each aircraft's attributes ,

[0095]

[0096] in, is the matrix of maximum aircraft attributes, is the matrix of each aircraft attribute, is the weight, n for is the weighted normalization term, n is the total number of aircraft;

[0097] For the target matching function , defining the minimum sensor range required for the mission and minimum battery life ,

[0098]

[0099]

[0100] in, is the importance weight, is the aircraft speed, is the sensor range, is the minimum sensor range, For the minimum endurance time, For battery life;

[0101] Calculate range efficiency, the distance between the aircraft and the target ,

[0102]

[0103] in, is the target reference position, is the maximum distance threshold;

[0104] Calculate priority, taking into account enemy target mobility Importance of goals

[0105]

[0106] in, For target mobility, is the importance coefficient;

[0107] Each aircraft traverses all mission-target pairs , generate a bid list , using the list of valid tasks-targets :

[0108]

[0109] in, To bid, Remove The lowest bid among all the other bids is used to select the task-target pair that maximizes the reward. :

[0110]

[0111] in, For the task, For the goal, is the bid of the aircraft for the task-goal pair (t, g);

[0112] In the consistency phase, get the number of opponent targets and the number of our aircraft , calculate the task-goal allocation scale through rule-driven

[0113]

[0114] The average capability of each aircraft is estimated through real-time status, The reserve resource ratio reflects the current resource usage.

[0115] Design constraints on mission requirements to limit the maximum aircraft requirements and minimum aircraft requirements ,

[0116]

[0117] in, Assigning a size to the task-goal, j is a task-goal pair;

[0118] The exclusive aircraft constraint ensures that each aircraft is assigned to only one mission-target pair.

[0119]

[0120] To ensure that the total number of assigned aircraft does not exceed the number of available aircraft ,

[0121]

[0122] in, is the total number of aircraft assigned;

[0123] Exchange bidding information with neighbors through the local communication network to resolve task allocation conflicts;

[0124] The aircraft sends a signal to neighbors within the communication range. Broadcast Bids , task-goal pair selection and aircraft number .

[0125]

[0126] Receiving neighbors Bid list, update local bid list :

[0127]

[0128]

[0129]

[0130] in, Bid set for neighbor aircraft, is the temporary bidding list of aircraft i at time t+1, is the bid list of aircraft i at time t, is the set of bids for each aircraft i on the task-target pair j, Bids for neighbor aircraft, () is to merge multiple bidding sources, Flying vehicles for neighbors;

[0131] Check the allocation status, update the available task-target pairs, and determine the set of unfilled task-target pairs , re-select the value,

[0132]

[0133] For each task, keep the highest bid Assign aircraft to each unit and reallocate aircraft after resolving conflicts

[0134]

[0135]

[0136] in, is the task allocation of aircraft i at time t+1, For the aircraft with the highest bid, is the set of winners of task-target pair j at time t+1;

[0137] Check if the consistency has converged

[0138]

[0139] in, is the set of winners for task-target pair j, is the set of winners of task-target pair j in the previous round;

[0140] In the collaborative strategy optimization part, the collaborative strategy is trained based on a multi-agent proximal strategy optimization algorithm. The input includes the task-target pair, the aircraft state including position, speed, sensor data, and the environment state. The action space includes movement, scanning, capture, etc.

[0141] Reward functions and loss functions are designed for different tasks:

[0142] Search Tasks :The reward function encourages maximizing coverage, and the loss function optimizes the efficiency of sensor data collection; Task Reward function:

[0143]

[0144] in, is the reward function at time t, is the target position, is the minimum distance, is the intra-group distance, is the aircraft position, is the penalty variable;

[0145] Loss Function

[0146]

[0147] in, For the task The policy loss function is is the empirical estimate expectation, is the advantage function, is the reward ratio, is the cropping range, is the strategy parameter, For the cutting operation;

[0148] For the task The value function loss, is the value function, For status, To make up your mind, In return for the discount, is the discount factor, For single-step rewards;

[0149] roundup mission : The reward function encourages the encirclement success rate, and the loss function reduces the interference conflict.

[0150] Task Reward function:

[0151]

[0152] in, is the reward function at time t, is the state covariance, For observation information, are weight parameters, is the penalty variable;

[0153] Loss Function

[0154]

[0155] in, is the policy loss function, is the advantage function, is the value function loss function, is the value function, For status, For action, In return for the discount, is the covariance normalization weight, is the variance term of the state covariance;

[0156] Tracking Tasks : The reward function encourages continuous tracking accuracy, and the loss function optimizes endurance efficiency.

[0157] Task Reward function:

[0158]

[0159] in, is the reward function at time t, is the energy loss, The group for task j3;

[0160] Loss Function

[0161]

[0162] in, is the policy loss function, is the advantage function, is the value function loss function, is the value function, For status, For action, In return for the discount, is the energy loss, is the regularization weight;

[0163] Training outputs the expected vehicle speed Finally, considering that the expected speed output by reinforcement learning may have noise or mutation, the moving average is used to smooth the speed sequence, and then the time step is dynamically adjusted according to the distance from the target. , implement collaborative strategies and improve task completion efficiency.

[0164] Implementation Method 2

[0165] An embodiment of the present invention provides a heterogeneous aircraft swarm multi-scale cross-domain autonomous confrontation decision-making system, which uses the heterogeneous aircraft swarm multi-scale cross-domain autonomous confrontation decision-making method described in Embodiment 1. The system includes:

[0166] Adversarial task set establishment module: establishes an adversarial task set for heterogeneous aircraft clusters;

[0167] Determine the adversary's goals based on the adversarial task set;

[0168] Distributed task allocation module: Based on the opponent's goals, the extended consensus auction algorithm is used to obtain the allocation of task-goal pairs;

[0169] Collaborative Strategy Optimization Module: Based on the allocation of task-goal pairs, the collaborative optimization module's multi-agent proximal strategy optimization algorithm is used to achieve reinforcement learning of heterogeneous aircraft clusters.

Claims

1. A multi-scale cross-domain autonomous confrontation decision-making method for heterogeneous aircraft clusters, characterized by: The method comprises the following steps: Step 1: Establish a set of adversarial tasks for a heterogeneous aircraft cluster; Step 2: Based on the adversarial task set in step 1, determine the opponent's goal; Step 3: Based on the opponent's goal in step 2, the distributed task allocation module is used to obtain the assignment of task-goal pairs; Specifically, an extended consensus auction algorithm is used. During the auction phase, each aircraft traverses the mission-target pairs and generates a bid list based on a value function that integrates mission suitability, target matching, distance efficiency, and priority. Each aircraft defines a matching matrix based on mission suitability, and calculates mission suitability based on aircraft attributes; Determine the minimum sensor range and endurance required for the mission and calculate the target matching function; Calculate the distance efficiency between the aircraft and the target; Calculate priorities based on target mobility and importance; Generate a list of bids and select the task-target pair that maximizes the global reward; Step 4: Based on the task-goal pairs from step 3, implement reinforcement learning for heterogeneous aircraft swarms through collaborative strategy optimization. The extended consensus auction algorithm specifically includes the following steps: in the consistency phase, the number of adversary targets is obtained, the task-target allocation scale is determined through rule-driven methods, the aircraft broadcasts a bid list through the local communication network, receives bid information from neighboring aircraft, updates the consistency table, retains the aircraft group with the highest bid, and iterates multiple times to achieve task allocation consistency; The collaborative strategy optimization in step 4 is specifically based on the proximal strategy optimization of the heterogeneous aircraft cluster, taking the task-target pair, aircraft state and environment state as input, designing reward functions and loss functions for different tasks, and optimizing the aircraft collaborative strategy.

2. The method according to claim 1, characterized in that The heterogeneous aircraft cluster in step 1 specifically includes multiple types of aircraft, and the performance parameters of each type of aircraft include reconnaissance capability, endurance capability, and jamming capability; The confrontation task set of step 1 includes a search task, a roundup task, and a tracking task, each task having a defined location, capability requirement, and priority.

3. The method according to claim 1, characterized in that Specifically, the opponent's target in step 2 is partially known, and the position and threat level are obtained through dynamic updating.

4. The method according to claim 1, characterized in that The distributed task allocation module of step 3 is specifically based on the extended consensus auction algorithm, comprehensively considering task adaptability, aircraft performance, target characteristics and environmental factors, and allocating task-target pairs through auction stage value function calculation and consistency stage.

5. The method according to claim 1, characterized in that: The consistency phase is specifically as follows: Step 3.1: Obtain the number of adversary targets and obtain the task-target allocation scale through rule-driven methods; Step 3.2: The aircraft broadcasts a bid list, including the mission-target pair and aircraft number, to neighbors within the communication range through the local communication network. Step 3.3: Receive the bid list of neighboring aircraft and update the local bid list; Step 3.4: Check the task allocation status and determine the set of unfilled task-target pairs; Step 3.5: Reselect the value, retain the aircraft group with the highest bid, and resolve the allocation conflict; Step 3.6: Check whether the consistency has converged. If not, repeat steps 3.1 to 3.4 with the maximum number of iterations as the termination condition or until the consistency converges.

6. The method according to claim 1, characterized in that The multi-agent proximal policy optimization algorithm of the collaborative optimization module includes the following steps: Step 4.1: Taking the mission-target pair, the aircraft state, and the environment state as input, output the desired position and attitude of the UAV. For the search and tracking mission, the action space includes movement and scanning; for the roundup mission, the action space includes movement and capture. Step 4.2: For the search task, design a reward function to maximize coverage and a loss function to optimize sensor data acquisition efficiency; Step 4.3: For the encirclement task, design a reward function to improve the encirclement success rate and a loss function to reduce interference conflicts; Step 4.4: For the tracking task, design a reward function to improve continuous tracking accuracy and a loss function to optimize endurance efficiency.

Citation Information

Patent Citations

  • Distributed auction algorithm-based heterogeneous multi-target multi-unmanned ship task allocation method

    CN117930852A

  • Port multi-agent task allocation method and system based on reinforcement learning

    CN119273104A