Air-ground unmanned cluster conflict resolution method based on multi-agent reinforcement learning

By employing an air-to-ground hierarchical collaborative mechanism and a multi-head attention mechanism, the problems of insufficient real-time performance and generalization ability in air-to-ground unmanned swarm conflict resolution are solved, achieving efficient air-to-ground unmanned swarm conflict resolution and improving the real-time performance and policy convergence speed of the swarm.

CN121325577APending Publication Date: 2026-01-13BEIHANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511333462.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies for resolving conflicts in unmanned air-to-ground swarms suffer from poor real-time performance, slow response to environmental changes, and exponential growth in algorithm complexity with swarm size. Traditional multi-agent reinforcement learning frameworks are prone to the "curse of dimensionality" in high-dimensional state-action spaces, resulting in slow policy convergence and weak cross-scenario generalization ability.

Method used

A hierarchical air-to-ground collaborative mechanism is adopted, treating each subgroup as an autonomous intelligent agent. At the local layer, conflict resolution strategies for small-scale air-to-ground formations are learned, and at the global layer, multi-subgroup parallel expansion is achieved through decentralized negotiation. A multi-head attention mechanism is introduced to weight and focus on the state of neighboring subgroups and obstacle information. Combined with the design of safety constraint filters and reward functions, decision-making accuracy and generalization ability are improved.

Benefits of technology

It significantly improves the real-time performance and cluster scalability of the algorithm, increases the policy convergence speed and decision accuracy, and enhances the generalization ability to different scenarios and scale changes. Simulation results show that the success rate and average time are better than existing methods in narrow passage and high-density obstacle scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121325577A_ABST
    Figure CN121325577A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles and unmanned vehicles, in particular to an air-ground unmanned cluster conflict resolution method based on multi-agent reinforcement learning, which comprises the following steps of: establishing an air-ground unmanned cluster navigation environment, determining obstacles and a plurality of air-ground subgroups, and determining that routes from each air-ground subgroup to a target navigation point have conflict positions; determining a state space and an action space of an air space subgroup and a reward and punishment function of an air space subgroup action based on the air space unmanned cluster navigation environment; establishing strategy networks including a local Q network, a local strategy network, a hybrid network and an attention network; determining a network loss function; performing reinforcement learning training on the strategy network to obtain a trained strategy network; controlling the air-ground subgroup to navigate by a decision action output by the trained strategy network; according to the invention, the conflict of multiple groups of clusters can be resolved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicles (UAVs) and unmanned vehicles (UAVs), specifically to a method for resolving conflicts in air-to-ground unmanned swarms based on multi-agent reinforcement learning. Background Technology

[0002] With the increasing maturity of flight control, sensing, and computing technologies for unmanned autonomous systems such as unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs), the overall performance and intelligence level of air-to-ground heterogeneous platforms continue to improve. Their combat functions are showing a diversified and collaborative development trend, making them a key force in future information-based joint operations. Especially in complex and rapidly evolving battlefield environments, single platforms have inherent limitations in situational awareness, path planning, and collaborative decision-making. However, air-to-ground unmanned swarms composed of multiple UAVs and UGVs can achieve wide-area coverage in space and complementary advantages in function through network networking and information sharing, significantly enhancing overall mission completion efficiency and survivability.

[0003] However, when air-to-ground unmanned swarms perform tasks in dynamic, multi-threat environments, they often face highly coupled spatial conflicts—including formation collisions between homogeneous platforms, trajectory intersections between heterogeneous platforms, and temporary conflicts with external dynamic obstacles. Untimely conflict resolution directly weakens the swarm's collaborative effectiveness and can even lead to mission failure. Existing conflict resolution methods mainly include rule-based collision avoidance strategies, centralized optimization algorithms, and heuristic search. However, in large-scale heterogeneous swarm scenarios, these methods generally suffer from poor real-time performance, slow response to environmental changes, and exponential growth in algorithm complexity with swarm size.

[0004] Multi-Agent Reinforcement Learning (MARL) offers a data-driven, online adaptive solution to complex conflict resolution tasks. By enabling aerial drones and ground-based unmanned vehicles to interact and learn as cooperative agents in simulated or real-world environments, MARL can discover optimal or near-optimal collaborative action strategies through continuous trial and error. However, traditional MARL frameworks are prone to the "curse of dimensionality" in high-dimensional state-action spaces, resulting in slow policy convergence, susceptibility to local optima, and weak cross-scenario generalization ability. Summary of the Invention

[0005] In view of the above problems, the present invention provides a method for resolving conflicts in air-to-ground unmanned swarms based on multi-agent reinforcement learning, which solves the technical problem of how to resolve conflicts in multiple swarms in the prior art.

[0006] This invention provides a method for resolving conflicts in air-to-ground unmanned swarms based on multi-agent reinforcement learning, comprising the following steps:

[0007] Step S1: Establish an air-ground unmanned cluster navigation environment, including: identifying obstacles and multiple air-ground subgroups, each air-ground subgroup having a target navigation point, the routes from each air-ground subgroup to its target navigation point having conflicting locations, and each air-ground subgroup not being able to pass through the conflicting locations simultaneously;

[0008] Step S2: Determine the state space, action space, and reward / penalty function of the air-ground subgroup's actions based on the air-ground unmanned cluster navigation environment.

[0009] Step S3: Establish the policy network, including the local Q-network, local policy network, hybrid network, and attention network; determine the network loss function;

[0010] Step S4: Based on the state space, action space, reward and penalty function of the action of the empty-ground subgroup, and network loss function, perform reinforcement learning training on the policy network to obtain the trained policy network.

[0011] Step S5: The decision action output by the trained policy network controls the air-to-ground subgroup to navigate, thereby resolving conflicts in the air-to-ground unmanned cluster.

[0012] Preferably, in step S1, each of the airspace subgroups includes a drone and multiple unmanned vehicles; the drone is able to detect unmanned vehicles in the cone-shaped area directly below it, and the unmanned vehicles are able to detect other unmanned vehicles in the surrounding circular area;

[0013] Step S1 also includes: determining the maximum speed and maximum acceleration of the drone and the unmanned vehicle, determining the perception radius of the drone and the unmanned vehicle, and determining the minimum distance and expected distance of the drone and the unmanned vehicle.

[0014] Preferably, step S2 specifically includes:

[0015] Step S2-1: Based on the position, velocity, major axis length, tilt angle, and obstacle position of the virtual elliptical region where the air-ground subgroup is located, determine the content of the air-ground subgroup's own state information and external observation information.

[0016] Step S2-2: Determine the action range of the air-to-ground subgroup based on the UAV's acceleration control parameters, the major axis of the elliptical region, and the tilt angle range.

[0017] Step S2-3: Determine the target navigation reward, external obstacle collision penalty, internal virtual area collision penalty, virtual area planning tilt angle change penalty, and virtual area roundness deviation penalty, and synthesize them to obtain the reward and penalty function.

[0018] Preferably, step S2-2 specifically includes:

[0019] UAVi represents the i-th drone in an air-to-ground drone swarm, and UAVi's actions. Defined as:

[0020] a i =(u i ,L i ,θ i )∈{(u,L,θ)|-a max ≤u i ≤a max ,L mix ≤L≤L max ,-θ max ≤θ≤θ max}

[0021] Among them, a i Indicates the action of UAVi. u represents a real number field with dimensions of 1×4. i L represents the acceleration of UAVi. i ,θ i L represents the major axis and inclination angle of the virtual elliptical region corresponding to UAVi. min ,L max Let θ represent the minimum and maximum values ​​of the major axis of the virtual elliptical region, respectively. max This indicates the maximum tilt angle.

[0022] Preferably, in steps S2-3:

[0023] The target navigation reward is used to reward the air-ground subgroup for approaching the corresponding target navigation point; the external obstacle collision penalty is used to punish the air-ground unmanned subgroup for approaching obstacles; the internal virtual region collision penalty is used to punish the overlap between the virtual elliptical regions of the unmanned subgroup; the virtual region planning tilt angle change penalty is used to punish the change in the orientation of the virtual elliptical regions of the unmanned subgroup at different times; the virtual region roundness deviation penalty is used to punish the behavior of the virtual elliptical regions of the unmanned subgroup deviating excessively from the circle in geometric shape.

[0024] The reward and punishment function r i The expression for (k) is:

[0025] r i (k)=α1r T +βr o +λr i +μr θ +ωr r

[0026] Where, r T ,r o ,r i ,r θ ,r rα1, β, λ, μ, ω represent the target navigation reward, external obstacle collision penalty, internal virtual area collision penalty, virtual area planning tilt angle change penalty, and virtual area roundness deviation penalty, respectively, with adjustment weights.

[0027] Preferably, step S3 specifically includes:

[0028] Step S3-1: Establish local Q-network, local policy network, hybrid network, and attention network;

[0029] Step S3-2: Determine the loss function of the local Q network, the loss functions of the two local policy networks, the loss function of the attention network, and the loss function of the hybrid network.

[0030] Preferably, step S3-2 specifically includes:

[0031] The loss function of a local Q-network is defined as:

[0032]

[0033] in, Represents the loss function. R represents the expected value based on the sampled data. t Indicates the single-step reward, and γ represents the discount factor. This represents the estimation of the target's global Q-value. Indicates the evaluation of the global Q-value, s t ,τ t ,a t represents the state, trajectory, and action at time t, respectively, and j2 represents the number of the hybrid network;

[0034] The loss function of an attention network is defined as:

[0035]

[0036] in, Represents the loss function of the attention network;

[0037] The loss function for the two local policy networks is defined as follows:

[0038]

[0039] in, Let π(a) represent the loss function of a local policy network. t ∣τ t ) represents the global specific strategy at time t; q mix Represents the q-value of the hybrid network. Indicates π i The mathematical expectation of the payoff under the strategy, This represents the individual's specific strategy at time t. This represents the global Q-value under the target parameters;

[0040] The loss function of a hybrid network is defined as:

[0041]

[0042] Where L(α) represents the loss function of the hybrid network, Indicates based on a t ~π t Mathematical expectation under action distribution It represents the mathematical expectation of a uniform distribution.

[0043] Preferably, step S4 specifically includes:

[0044] Step S4-1: Initialize network parameters, including initializing the parameters of the local Q network, local policy network, hybrid network and attention network, initializing the replay buffer, and setting the maximum number of training epochs;

[0045] Step S4-2: Each agent obtains the global state and individual observations of the current and next time moments through observation, determines and executes actions by the local policy network, interacts with the environment, obtains global rewards, and adds the global state, individual observations, determined actions, global state and individual observations of the next time moment to the replay buffer.

[0046] Step S4-3: Randomly sample data from the replay buffer to obtain sampled data, and update the parameters of the policy network based on the sampled data;

[0047] Step S4-4: Return to step S4-2 to execute the next round until the maximum number of training rounds is reached, and obtain the trained policy network.

[0048] Preferably, step S4-3 specifically includes:

[0049] The local Q-network and attention network are updated using the following formula:

[0050]

[0051]

[0052] Where ← indicates assignment or update, α Q This indicates the update step size, and j2 represents the hybrid network number. This represents the gradient operator for evaluating hybrid networks. This represents the loss function used to evaluate the hybrid network;

[0053] The local policy network is updated using the following formula:

[0054]

[0055] Where, α π Indicates the update step size. L represents the gradient operator of a local policy network. π (θ) represents the loss function of the local policy network;

[0056] Update the hybrid network using the following formula

[0057]

[0058] Where τ represents the soft update weight, This represents the target hybrid network parameters.

[0059] Preferably, step S4 further includes:

[0060] Set the working status flags for each UAVi work (i), with an initial value of s work (i) = working; when agent i successfully reaches its corresponding target navigation point, set s. work (i) = success; when agent i touches an external obstacle, set s. work (i) = dead;

[0061] If and only if the working status flag of agent i is s work When (i) = working, perform steps S4-2 and S4-3 on it.

[0062] Compared with the prior art, the present invention has at least the following beneficial effects:

[0063] (1) The present invention adopts an air-ground hierarchical collaborative mechanism, treating each subgroup as an autonomous intelligent agent. It learns conflict resolution strategies for small-scale air-ground formations at the local layer, and then achieves parallel expansion of multiple subgroups through decentralized negotiation at the global layer. This avoids the "curse of dimensionality" caused by training directly in a huge state-action space, and significantly improves the real-time performance and cluster scalability of the algorithm.

[0064] (2) This invention introduces a multi-head attention mechanism into the policy network, which weights and focuses on the states of neighboring subgroups, obstacle information, and its own dynamics, enabling the agent to automatically identify and prioritize the most critical elements for conflict resolution. This mechanism effectively reduces redundant information interference and improves decision-making accuracy; on the other hand, it significantly enhances the policy's generalization ability to different scenarios and scales, accelerates training convergence, and improves the interpretability and robustness of the method.

[0065] (3) The present invention embeds collision penalty, passage efficiency reward and cooperative stability reward into the reward function, and performs real-time pruning of high-risk actions in conjunction with safety constraint filter, which not only ensures the safety and reliability of the conflict resolution process, but also accelerates the convergence of the strategy. Simulation results show that in typical narrow passage and high-density obstacle scenarios, the success rate and average time of the present invention are better than the existing rule-driven or centralized optimization methods. Attached Figure Description

[0066] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0067] Figure 1 This is a schematic diagram illustrating multiple air-to-ground unmanned cluster navigation and obstacle avoidance tasks provided by the present invention.

[0068] Figure 2 This is a schematic diagram of the unmanned vehicle and its communication structure model provided by the present invention.

[0069] Figure 3 This is a schematic diagram of the unmanned vehicle and its perception model provided by the present invention.

[0070] Figure 4 This is a schematic diagram illustrating the attention network processing of an indefinite number of external observations provided by the present invention.

[0071] Figure 5 A schematic diagram of the reward function component design provided by this invention.

[0072] Figure 6 This is a schematic diagram of a scenario for resolving conflicts in an unmanned air-to-ground cluster provided by the present invention.

[0073] Figure 7 A schematic diagram of the training reward curve provided by the present invention.

[0074] Figure 8 This is a schematic diagram of the training local Q-network loss curve provided by the present invention.

[0075] Figure 9 This is a schematic diagram of the training local policy network loss curve provided by the present invention.

[0076] Figure 10 This is a snapshot diagram of a simulation of conflict resolution for unmanned air-to-ground clusters provided by the present invention.

[0077] Figure 11 This is a schematic diagram of the distance change curve between a UAV and its nearest UAV, provided by the present invention.

[0078] Figure 12 This is a schematic diagram of the distance change curve between the UAV and the target provided by the present invention.

[0079] Figure 13The flowchart shows the air-to-ground unmanned swarm conflict resolution method based on multi-agent reinforcement learning provided by this invention. Detailed Implementation

[0080] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0081] This invention proposes a conflict resolution method for air-to-ground unmanned swarms based on multi-agent reinforcement learning. By designing key technologies such as hierarchical collaborative state representation, hybrid action decomposition mechanism and safety constraint filter, it effectively reduces the learning dimension, improves policy convergence efficiency, and significantly enhances real-time conflict resolution and robust collaborative capabilities in dynamic battlefield environments.

[0082] like Figure 1 As shown, this invention discloses a method for resolving conflicts in air-to-ground unmanned swarms based on multi-agent reinforcement learning. The specific implementation steps are as follows:

[0083] Step S1: Establish an air-ground unmanned cluster navigation environment, including obstacles and multiple air-ground subgroups. Each air-ground subgroup has a target navigation point. There are conflicting locations on the routes from each air-ground subgroup to its target navigation point. Each air-ground subgroup cannot pass through the conflicting locations at the same time.

[0084] In this step, the present invention first identifies the unmanned swarm navigation scenario, specifically the unmanned swarm conflict resolution scenario.

[0085] like Figure 1 As shown, the unmanned swarm conflict resolution scenario involves multiple air-to-ground sub-swarms, each consisting of one UAV in the air and multiple unmanned vehicles on the ground. Each air-to-ground sub-swarm has a target navigation point, and the routes from each sub-swarm to its target navigation point have conflict locations, i.e., narrow passages, through which the sub-swarms cannot pass simultaneously.

[0086] If at any given moment only one empty subgroup occupies the narrow passage and all subgroups pass through smoothly in the negotiated order without collision, then the conflict is considered successfully resolved. If two or more subgroups enter the narrow passage at the same time, causing a passage conflict, or if any drone or unmanned vehicle collides or gets stuck in the narrow passage, then the conflict is considered to have failed to resolve.

[0087] For multiple air-to-ground subgroups to complete their navigation and obstacle avoidance tasks, the following control objectives must be met: 1) The air-to-ground unmanned cluster navigates to the corresponding target endpoint; 2) The air-to-ground unmanned cluster avoids external static obstacles; 3) The air-to-ground unmanned cluster maintains cohesion and avoids internal collisions; 4) Conflicts between air-to-ground subgroups are resolved. The expressions controlling the air-to-ground subgroups are specifically described by formulas (1.1)-(1.4), which are detailed below.

[0088] Control objective 1 requires the existence of a finite time T. a ,satisfy

[0089]

[0090] Where, x a (T a ) is T a The location of the drone at any time, x d Let x be the coordinates of the target endpoint, ||·|| represent the modulus, ε1 and ε2 are preset positive constants, and x g,i′ (T a ) is T a Let N be the position of the i′-th unmanned vehicle at time i, and N be the total number of unmanned vehicles. The above expression indicates that the positions of all drones and unmanned vehicles in the air-to-ground unmanned swarm can converge to the vicinity of the target endpoint within a finite time.

[0091] Control objective 2 requires heterogeneous clusters to meet the following requirements:

[0092]

[0093] Where, x g,i′ (t) represents the position of the i′-th driverless car at time t, x k d represents the coordinates of the spatial obstacle point. safe,g This is the obstacle avoidance safety distance set by the autonomous vehicle, where t is the distance between time 0 and T. a At any point in time between moments, Let O represent any set of obstacles in space. The above expression represents the autonomous vehicle on the ground maintaining a safe distance from obstacles throughout its coordinated movement.

[0094] Control objective 3 requires heterogeneous clusters to meet the following requirements:

[0095]

[0096] Where, x g,i′ (t) and x g,j′ (t) represents the positions of the i′ and j′-th unmanned vehicles at time t, and d int Indicates the safe distance for collision avoidance between autonomous vehicles; Indicates existence, R cIt is the cluster cohesion constant. The above expression indicates that unmanned air-to-ground swarms maintain clustering and avoid collisions during coordinated movement.

[0097] Control objective 4 reflects the mission's requirement for safe and efficient coordination between air-to-ground subgroups. Let the coordinates of the formation center of air-to-ground subgroup 1 at time t be x. a,1 (t)=(x a1,1 (t),x a2,1 (t)), where the major and minor axes of the formation parameters of the empty subgroup 1 are a1(t) and b1(t) respectively, and the orientation angle is θ1(t); the center coordinates of the formation of the empty subgroup 2 are x a,2 (t)=(x a1,2 (t),x a2,2 (t)), where the major and minor axes are a2(t) and b2(t) respectively, and the orientation angle is θ2(t). Control objective 4 requires the following conditions to be met:

[0098]

[0099] Where, x a,1 ,x a,2 Let a1, b1, θ1, a2, b2, θ2 represent the center positions of the formations of empty subgroup 1 and empty subgroup 2, respectively. Let a1, b1, θ1, a2, b2, θ2 represent the major axis length, minor axis length, and orientation angle of empty subgroup 1 and empty subgroup 2, respectively. Let φ represent the angle of the line connecting the center positions of the two empty subgroups.

[0100] For the air-to-ground unmanned swarm navigation scenario, this invention constructs kinematic, communication, and perception models for the air-to-ground unmanned swarm, which are described in detail below:

[0101] (1) Kinematic model of unmanned swarm in air and ground

[0102] The kinematic model of a nonholonomic mobile autonomous vehicle can be represented as:

[0103]

[0104] in, It is the rate of change of the Cartesian coordinates at the center of the autonomous vehicle, v g,i′ It is linear velocity, θ g,i′ It's the azimuth. It is the rate of change of azimuth, ω g,i′ It is angular velocity, J g,i′ It is the moment of inertia of the driverless car, f g,i′ and τ g,i′ These are the forces and torques acting on the driverless vehicle.

[0105] To avoid the nonholonomic constraints introduced by the above kinematic model, the nonholonomic motion autonomous vehicle model is transformed into dual integrator dynamics using input-output feedback linearization technology. To simplify problem analysis and coordination algorithm design, the kinematic model of the autonomous vehicle is taken as follows:

[0106]

[0107] in Indicates the position of the driverless car i′, v g,i′ This represents the speed of the driverless car i′. Let u represent the acceleration of the autonomous vehicle i′. g,i′ Let N represent the control quantity of driverless car i', and let N represent the total number of driverless cars.

[0108] The kinematic model of the drone is as follows:

[0109]

[0110] Where, x i y i and z i V represents the position of the UAV in the inertial coordinate system. i , χ i and γ i These represent the drone's speed, yaw angle, and pitch angle, respectively. x i y i and z i rate of change, V i , χ i and γ i rate of change, L i D i and T i These represent the lift, drag, and thrust of the drone, respectively. It is the drone's roll angle, m i The mass of the drone is g, and g is the acceleration due to gravity.

[0111] Using a second-order integrator model for simplification, the kinematic model of the UAV can be represented as:

[0112]

[0113] in, Indicates the position of drone i, v i Indicates the speed of drone i. Indicates v i rate of change, u i This represents the control quantity of drone i.

[0114] (2) Air-to-ground unmanned cluster communication model

[0115] In the aforementioned air-to-ground unmanned swarm, the UAVs interact with the unmanned vehicles via broadcast. Let r be the communication radius between the UAVs and the unmanned vehicles. c,a (If the communication radius covers all unmanned vehicles), then the set N of unmanned vehicles that communicate and interact with the drone is... a for:

[0116] N a ={j'∈v:||x a -x g,j' ||<r c,a} (1.9)

[0117] Where v = {1, 2, ..., N} represents the ID of all autonomous vehicles, x a For the drone's location, x g,j' Let r be the location of the j'-th driverless car. c,a This refers to the communication radius between drones and unmanned vehicles.

[0118] The autonomous vehicle establishes communication links with neighboring autonomous vehicles within its range. For the i'-th autonomous vehicle, the set of autonomous vehicles it communicates with is... for:

[0119]

[0120] Where v'={1,2,…,i'-1,i'+1,…,N} represents the ID number after excluding the i′-th autonomous vehicle, x g,i Let i' be the location of the i-th driverless car.

[0121] This establishes the detection and communication constraints for the air-to-ground unmanned swarm. N a The unmanned vehicles in the text are called the neighbors of drones, and will The driverless car in the diagram is called the neighbor of the i′-th driverless car.

[0122] like Figure 2 The diagram illustrates the unmanned vehicle and its communication structure model according to the present invention.

[0123] (3) Unmanned Cluster Perception Model

[0124] The drones and unmanned vehicles in the air-to-ground unmanned swarm possess heterogeneous perception capabilities. The drone's onboard sensors can detect a cone-shaped area directly below it, which can be modeled as a cone (height is the drone's flight altitude h, and the base radius is the detection radius r). p If the detection area of ​​the UAV is:

[0125] D a ={x a,k |||x a,k -xa ||<r p (1.11)

[0126] Among them, D a The set of obstacles within the area detected by the drone; x a,k The x represents the two-dimensional coordinates of the obstacle point detected by the drone. a The two-dimensional position of the drone is represented by r. p This indicates the radius of the bottom surface of the detection area.

[0127] Autonomous vehicles perceive obstacles based on distance sensors and can detect a circular area around the vehicle. This detection area can be represented as:

[0128]

[0129] in, Let x be the set of obstacles detected by the i'-th autonomous vehicle; g,k The two-dimensional coordinates r of the obstacle point detected by the autonomous vehicle p,g Let r be the detection radius, and satisfy r d,g <r d,a .

[0130] In the air-to-ground unmanned swarm perception model, the perception models of drones and unmanned vehicles are as follows: Figure 3 As shown.

[0131] Through the above steps, this invention determines the kinematic models of each UAV and unmanned vehicle in the air-to-ground unmanned swarm, and determines the relevant range limitations for communication and perception of each UAV and unmanned vehicle. Thus, an air-to-ground unmanned swarm navigation environment is established for subsequent unmanned swarm conflict resolution planning.

[0132] Step S2: Determine the state space, action space, and reward / penalty function of the air-ground subgroup's actions based on the air-ground unmanned cluster navigation environment.

[0133] The following provides a detailed description of the state space, action space, and reward / penalty function for the actions of the empty land subgroup as determined by this invention.

[0134] (1) State space

[0135] For unmanned aerial vehicles (UAVs), this invention aims to utilize control methods to achieve control of UAVs within an air-to-ground UAV swarm and to enable UAVs to plan their routes within virtual areas. In multi-agent navigation tasks with obstacle avoidance, the goal of the air-to-ground UAV swarm is to reach the target point as quickly as possible while ensuring safety, and to avoid collisions with other swarms or obstacles. Therefore, the observations of upper-level UAVs should include information such as the relative position of the target point to the UAV.

[0136] Let UAVi denote the i-th UAV in the air-to-ground UAV swarm. The state space of the air-to-ground subswarm corresponding to this UAV represents the observation state o of UAVi. i Space, observation state o i Divided into its own state information and external observation information The following is a detailed description.

[0137] (1-1) UAVi's own status information

[0138] Specifically, UAVi's own state information is represented as in These represent the normalized relative position, relative velocity, major axis length of the ellipse, and ellipse inclination angle of UAVi, respectively, and their specific definitions are as follows:

[0139]

[0140] Where x i , Representing the position of UAVi and its target position respectively, r c This represents the communication radius of UAVi.

[0141]

[0142] Where v i ,v max These represent UAVi's speed and maximum speed, respectively.

[0143]

[0144] Where L i ,L max These represent the length of the major axis of the ellipse and the maximum velocity of the virtual region planned by UAVi, respectively.

[0145]

[0146] Where θ i ,θ max These represent the ellipse tilt angle and maximum tilt angle of the virtual region planned by UAVi, respectively.

[0147] (1-2) UAVi External Observation Information

[0148] Specifically, UAVi's own external observation information includes ((obs) j ) k ,(Ne m ) n), where k and n represent the number of obstacles observed by the UAVi and the number of neighboring UAVs, respectively. The external observations of the UAVi can refer to the air-to-ground unmanned swarm perception model determined in this invention.

[0149] obs j Ne m The expressions are as follows:

[0150]

[0151] Among them, obs j Mathematical modeling of obstacle j r p Let J represent the coordinates of obstacle j and the sensing radius of UAVi, respectively.

[0152]

[0153] Among them, Ne m Indicates the neighbor observation modeling of UAVm, x m ,v m ,L m ,θ m These represent the coordinates, velocity, major axis of the planned virtual region ellipse, and inclination angle of the planned virtual region ellipse, respectively.

[0154] In practical applications, the number of obstacles and other UAVs observed by UAVs is dynamically changing, i.e., external observation information (obs) j ) k ,(Ne m ) n In this paper, k and n are dynamically changing and uncertain. To address the issue of dynamically changing dimensions of the external observation information of UAVi, an attention network is set up in subsequent steps. The original external observation information with uncertain dimensions and the determined self-state information are input into the attention network. Attention weights are assigned according to the importance of each observation, and finally, a fixed-dimensional external observation information is output. Among them, obs i Represents the complete obstacle observation modeling of UAVi, Ne i This represents the modeling of all neighbor observations of UAVi. This represents a real number field with a dimension of 1×8.

[0155] Finally, the observation status of UAVi was obtained. For the global environmental information S, this invention sequentially stitches together the observation states of all UAVs according to their distances from their respective target points, from smallest to largest. i This yields the final global environment information S.

[0156] S=(o1,...,on (1.19)

[0157] (2) Action space

[0158] Let UAVi represent the i-th UAV in the air-to-ground UAV swarm. The action space of the corresponding air-to-ground sub-swarm of this UAV represents the action a of UAVi. i The space. Based on the motion model of the UAV in the above steps, continuous motion control of the UAV is achieved based on acceleration u, and virtual region planning is performed for the UAV based on the major axis L and tilt angle θ of the determined virtual elliptical region. The present invention will control the UAV's motion. Defined as acceleration control quantity and virtual region planning parameter e i =(L i ,θ i The combination of ) is:

[0159] a i =(u i ,e i )=(u i ,L i ,θ i (1.20)

[0160] Among them, a i Indicates the action of UAVi, u i U represents the acceleration of UAVi, u i Constrained within the upper and lower acceleration intervals u i =clip(u i ,-a max ,a max Within ) where a max L represents the maximum acceleration of the UAV. i ,θ i This represents the major axis and tilt angle of the virtual elliptical region corresponding to UAVi.

[0161] Virtual area planning parameter e i =(L i ,θ i Constrained within the feasible solution space of the virtual region planning, the expression is:

[0162] (L i ,θ i )∈{(L,θ)|L min ≤L≤L max ,-θ max ≤θ≤θ max}(1.21)

[0163] Among them, L min ,L max Let θ represent the minimum and maximum values ​​of the major axis of the virtual elliptical region, respectively.max This indicates the maximum tilt angle.

[0164] In summary, UAVi's actions Defined as:

[0165] a i =(u i ,L i ,θ i )∈{(u,L,θ)|-a max ≤u i ≤a max ,L mix ≤L≤L max ,-θ max ≤θ≤θ max}(1.22)

[0166] (3) Reward and punishment functions

[0167] The reward and penalty function is the core driving force of the multi-agent reinforcement learning framework, and its quality directly determines the policy search direction and convergence quality. In the air-to-ground unmanned swarm conflict resolution task, the agent not only needs to minimize the collision risk between itself and obstacles and neighbors, but also needs to simultaneously achieve multiple task objectives such as goal arrival coordination. In the air-to-ground heterogeneous swarm scenario, the aerial UAV needs to learn a refined motion control strategy for itself, and simultaneously plan the safety dynamic boundary of the entire unmanned vehicle swarm. The learning task presents a high degree of complexity with "individual-group" dual-scale coupling. This multi-objective, multi-temporal and spatiotemporal decision structure means that the native environment only returns reward signals in a very few times when the entire game is successful or unsuccessful, resulting in an extremely sparse reward space and difficulty in effectively propagating policy gradients. To alleviate early exploration stagnation and improve sample efficiency, this invention introduces "Reward Shaping" in addition to the original sparse reward: through components such as potential function progress reward, continuous risk penalty, and energy consumption regularization, it provides the agent with denser and directional feedback, guiding the learning process to quickly approach the optimal cooperative strategy while maintaining safety.

[0168] Specifically, such as Figure 5 As shown, the reward function under the Reward Shaping assistance based on the present invention includes the following components.

[0169] (3-1) Target navigation reward

[0170] The core idea of ​​target navigation reward is to encourage unmanned subgroups of air and ground to approach their corresponding target navigation points. For UAVi, this invention defines the target navigation power field as:

[0171] Φ T (x i ,x i,T )=||x i -xi,T || (1.23)

[0172] Where, Φ T (x i ,x i,T ) represents UAVi's target navigation field of influence, x i,T Representing the distance to the target point for UAVi, and combining the concept of reward shaping, the target navigation reward for UAVi at time k is defined as:

[0173]

[0174] Where, x i (k),x i (k-1) represents the position of UAV at time k and time k-1 respectively, γ is the training discount factor, and d0 represents the distance scale for determining successful arrival at the target point, which is a strictly positive constant.

[0175] (3-2) External obstacle collision penalty

[0176] The core idea of ​​external obstacle collision penalty is to penalize the behavior of unmanned subgroups in open areas that excessively approach external obstacles within a certain range. For UAVi, this invention defines the external obstacle collision force field as:

[0177]

[0178] Where, Φ o (x i ,x o ) represents the external obstacle collision force field of UAVi, x o d represents the coordinates of the nearest obstacle to UAVi. safe Representing the safety distance, which is a strictly positive constant, and combining the concept of reward-shaping, the external obstacle collision penalty of UAVi at time k is defined as:

[0179]

[0180] Where, x i (k),x i (k-1) represent the positions of UAVi at time k and time k-1, respectively.

[0181] (3-3) Collision penalty for internal virtual regions

[0182] The core idea of ​​internal virtual region collision penalty is to penalize the phenomenon of overlapping planned virtual regions when multiple unmanned subgroups of open areas occur. This penalty can effectively avoid collisions between unmanned vehicles in different subgroups, for UAVi and its perceived neighbors N. iSpecifically, this invention defines UAVi and its neighbor UAVj∈N i Internal virtual region collision penalty r i (j) is:

[0183]

[0184] Among them, E i E j S represents the virtual regions planned by UAVi and UAVj respectively, and S(·) represents the calculated area. max Indicates the maximum area.

[0185] The collision penalty for UAVi's internal virtual region can be defined as its collision with all its neighbors N. i The pairwise collision penalty is expressed as:

[0186]

[0187] Where, N i This represents the set of neighbors of UAVi.

[0188] (3-4) Penalty for changes in tilt angle in virtual area planning

[0189] The core idea of ​​the virtual region planning tilt angle variation penalty is to constrain the changes in the orientation of the virtual region planned by the upper-level UAV at different times to the greatest extent possible, preventing distortion caused by the change in the orientation of the virtual region planned by the UAV in continuous time intervals, which would lead to position convergence oscillations of the lower-level UAV cluster. This invention penalizes the behavior of frequent changes in the tilt angle of the virtual region planned by the air-to-ground UAV sub-swarm. For UAVI, this invention defines the virtual region planning tilt angle variation penalty as follows:

[0190] r θ (θ i (k),θ i (k-1))=-||θ i (k)-θ i (k-1)|| 2 (1.29)

[0191] Where, θ i (k),θ i (k-1) represent the tilt angles of UAV at time k and k-1, respectively.

[0192] (3-5) Virtual region roundness deviation penalty

[0193] The core idea of ​​virtual region roundness deviation penalty is to suppress excessive deviation of the virtual region's geometry from a circle during the planning phase. If upper-layer UAVs frequently elongate or flatten the ellipse to meet local obstacle avoidance or formation requirements, it will cause the passable corridors of the lower-layer UAV cluster to fluctuate in width, leading to a series of chain effects such as queue acceleration-deceleration oscillations and increased path replanning times. Therefore, this invention constrains the "roundness" of the virtual region at continuous time intervals: when the difference between the major semi-axis L(t) and the minor semi-axis b(t) determined by the area-preserving relationship is too large, the ellipse is judged to have significantly deviated from a circle, and a negative reward is given. Specifically, for UAVs, this invention defines the virtual region roundness deviation force field as:

[0194]

[0195] Among them, L i S represents the length of the major axis of the ellipse representing the virtual region planned by UAVi. area Let represent the known fixed area of ​​the virtual region ellipse, which is a strictly positive constant, and |·| denote the calculation of the absolute value. Combining the reward shaping concept, the virtual region roundness deviation penalty of UAVi at time k is defined as:

[0196] r r (L i (k),L i (k-1))=γΦ r (L i (k-1))-Φ r (L i (k))(1.31)

[0197] Among them, L i (k),L i (k-1) represent the major axes of the virtual region ellipse planned by UAVi at time k and time k-1, respectively.

[0198] In summary, the total reward / penalty function r of UAVi at time k is... i The expression for (k) is:

[0199] r i (k)=α1r T +βr o +λr i +μr θ +ωr r (1.32)

[0200] Where, r T ,r o ,r i ,r θ ,r rα1, β, λ, μ, ω represent the target navigation reward, external obstacle collision penalty, internal virtual area collision penalty, virtual area planning tilt angle change penalty, and virtual area roundness deviation penalty, respectively, with adjustment weights.

[0201] By defining the reward and penalty function, this invention will use the reward and penalty function to guide the learning process in subsequent reinforcement learning training, so as to quickly approach the optimal cooperative strategy while maintaining safety.

[0202] Step S3: Establish the policy network, including the local Q-network, local policy network, hybrid network, and attention network; determine the network loss function;

[0203] This invention proposes the MASAC-A (Attentional Multi-Agent SoftActor-Critic) algorithm, aiming to organically integrate the mature off-policy learning and maximum entropy regularization mechanism of the SAC algorithm, the self-attention mechanism that enhances information interaction and value aggregation capabilities among agents, and the linear decomposition architecture of "centralized training and distributed execution" in Value Decomposition Networks (VDN). Specifically, MASAC-A assumes that the overall joint Q-value can be written as a linear weighted sum of the local Q-values ​​of each agent. Based on this, it simultaneously introduces a double-Q objective network and entropy enhancement strategy optimization, and adds an attention mechanism to the individual observation module of each agent to achieve scalability of the observation state of each agent. In this way, the algorithm inherits the advantages of SAC in high sample efficiency and exploration stability in single-agent continuous control, effectively solves the non-stationarity and credit allocation problems in multi-agent environments using value decomposition methods, and ultimately achieves efficient and stable learning for large-scale collaborative tasks based on the self-attention mechanism.

[0204] This invention establishes a policy network for outputting decision-making actions, including a local Q-network, a local policy network, a hybrid network, and an attention network, described in detail below. The agent described in this step represents a drone. The functions and relationships of the four networks are as follows:

[0205] (1) Local Critic Network

[0206] Estimate the Q-value of a single agent under its local observations and actions to provide a reference for the action value of subsequent local policy network output actions.

[0207] (2) Local Actor

[0208] Based on the continuous action distribution of local information output, the actual observed output action of a single agent is generated.

[0209] (3) Mixing Network

[0210] The local Q-values ​​of each agent are combined with the global state to form the team value, and then the local networks are updated in reverse based on the team value.

[0211] (4) Self-Attention Encoder

[0212] It aggregates observations of entities such as neighbors, targets, and teammates with variable numbers and unpredictable order, producing fixed-dimensional context embeddings. It features permutation invariance and masking capabilities, mitigating the problems of variable input length and noise redundancy.

[0213] Overall input-output relationship

[0214] 1. The single observation is split into a fixed-dimensional part and a variable entity set. The variable part is first processed by self-attention to obtain fixed-dimensional features.

[0215] 2. The local policy network receives the representations from both and outputs the action distribution and samples the actions.

[0216] 3. Local Q-networks receive the same representations and actions and output the corresponding local values.

[0217] 4. Hybrid networks aggregate all local values ​​and global states into team value, which is used for target construction and error backpropagation during training.

[0218] 5. During execution, it relies only on local information and attention features; during training, it additionally uses global state to provide conditions for the hybrid network.

[0219] Detailed description is as follows:

[0220] (1) Local Q-network

[0221] Local Q-networks estimate the local value function q for each agent i. i (τ i ,a i This is reflected in a given self-observation-action history τ. i and local action a i The contribution of the agent to the global reward is then considered.

[0222] In the MASAC-A framework, all agents share the parameters of a local Q-network, and different agents reuse the same weights to ensure training efficiency and scalability.

[0223] (2) Local policy network

[0224] Local policy networks learn for each agent a history of observations and actions τ. i The continuous action distribution π i (a i |τ i During the execution phase, the agent selects actions to interact with the environment based on this distribution.

[0225] In the MASAC-A framework, all agents share the parameters of a local policy network, and different agents reuse the same weights. However, one-hot encoding is used at input to distinguish different agents, so as to ensure training efficiency and scalability while preserving the heterogeneity of different agents' action selection.

[0226] (3) Hybrid Network

[0227] The hybrid network combines the outputs of all local Q-networks {q} 1 ,...,q N} is associated with the global state S and mapped to a centralized joint Q-value Q. tot (s,τ,a). Each hybrid network consists of a Hypernet NetworkModule and a Linear Combination Module. The Hypernet NetworkModule takes the global state S as input and generates the linear weight vector {k} of the hybrid network. i (s)} and bias b(s). The linear combination module generates a linear weight vector {k} based on the hybrid network. i (s)} and bias b(s) and the outputs {q} of all local Q-networks 1 ,...,q N Perform linear combinations:

[0228]

[0229] Where s, τ, and a represent the state, historical trajectory, and action of a single agent, respectively.

[0230] (4) Attention Network

[0231] Attention networks are used to perform context-aware weighted aggregation of a variable number of features observed by each agent from neighboring agents and obstacles. Specifically, MASAC-A constructs a variable-length sequence {e1,...,e...} of the current agent's local Q-value, representing the value of the indicated local value function, along with the encoded vectors of all its neighbors and obstacles. K The input is fed into a multi-head scaled dot product self-attention layer, through a shared linear projection matrix W = (W Q W K W VConstruct the Query / Key / Value pairs, and then obtain the attention weights through a multi-head attention mechanism. Furthermore, the original sequence is weighted and summed to generate a context representation of length K+1, highlighting information about key entities.

[0232] All agents share the same set of self-attention module parameters. Different agents only need to input their own entity feature sequences. There is no need to manually design the network structure for a variable number of neighbors and obstacles. This ensures that the parameter scale is controllable and also enables flexible fusion of multi-source and variable inputs, providing more discriminative joint Q-value aggregation features for subsequent hybrid networks.

[0233] Furthermore, this invention also incorporates an attention network to process the raw external observation information of uncertain dimensions to obtain external observation information of fixed dimensions. For example... Figure 4 As shown, in practical applications, the number of obstacles and other UAVs observed by UAVs is dynamically changing, i.e., the external observation information (obs) j ) k ,(Ne m ) n In the given information, k and n are dynamically changing and uncertain. To address the issue of dynamically changing dimensions of the external observation information of UAVi, an attention network is set up. The original external observation information with uncertain dimensions and the fixed self-state information are input into the attention network. Attention weights are assigned according to the importance of each observation, and finally, a fixed-dimensional external observation information is output.

[0234] To ensure stability in multi-agent training environments, MASAC-A continues the approach of using an evaluation-target dual-network structure and a soft-threshold operator for updates. MASAC-A is based on two sets of local Q-networks to fuse the states of all agents and fit the global value Q. tot (s,τ,a), specifically, the loss functions of the two local Q-networks are defined as:

[0235]

[0236] in, Represents the loss function. R represents the expected value based on the sampled data. t Indicates the single-step reward, and γ represents the discount factor. This represents the estimation of the target's global Q-value. Indicates the evaluation of the global Q-value, s t ,τ t ,a t These represent the state, trajectory, and action at time t, respectively.

[0237] Based on the original Soft Actor-Critic (SAC) algorithm, the method of this invention also uses two hybrid networks. j2 represents the number of the hybrid network, and the smaller of their values ​​is used as the objective, specifically in the form of:

[0238]

[0239] in, Indicates π θ The mathematical expectation of the payoff under the strategy, s represents the global Q-value of the j2 network. t+1 ,τ t+1 ,a t+1 Let π(a) represent the state, trajectory, and action at time t+1, respectively, where α represents the weight, and π(a) represents the weight. t+1 ∣τ t+1 ) represents the global specific strategy at time t+1, q mix Represents the q-value of the hybrid network. Represents the local q value. This represents the individual's specific strategy at time t+1.

[0240] The output of the attention network directly affects the input of the hybrid network. As a feature preprocessing module of the local Q-network, it is attached to the loss backpropagation stage of the local Q-network when updating network parameters. Specifically, its loss function is defined as:

[0241]

[0242] in, Represents the loss function of the attention network;

[0243]

[0244] in, Let π(a) represent the loss function of a local policy network. t ∣τ t ) represents the global specific strategy at time t; q mix Represents the q-value of the hybrid network. Indicates π i The mathematical expectation of the payoff under the strategy, This represents the specific strategy employed by an individual at time t.

[0245] Here, α represents the weight of the policy entropy, generally referred to as the temperature parameter. The temperature parameter α needs to be set to different values ​​in different training stages or different tasks. This is because the required degree of exploration varies in different states: in some states, a better policy has been learned, and α should be reduced to a very small value to weaken exploration; while in other states, it is still uncertain which actions are good and which are bad, so the degree of exploration needs to be increased. The SAC algorithm proposes to reconstruct the original soft policy iteration process into a constrained optimization problem, that is, while optimizing the policy to maximize the cumulative discount reward, the average entropy of the policy is kept at a fixed value, while allowing the action entropy in each state to be variable. Specifically, the loss function of the hybrid network is defined as:

[0246]

[0247] Where L(α) represents the loss function of the hybrid network, Indicates based on a t ~π t Mathematical expectation under action distribution It represents the mathematical expectation of a uniform distribution.

[0248] Through the above steps, this invention establishes a policy network and determines the network loss function for subsequent reinforcement learning training.

[0249] Step S4: Based on the state space, action space, reward and penalty function of the action of the empty-ground subgroup, and network loss function, perform reinforcement learning training on the policy network to obtain the trained policy network.

[0250] In this step, the process of training the policy network through reinforcement learning includes the following steps:

[0251] Step S4-1: Initialize network parameters, including initializing the parameters of the local Q network, local policy network, hybrid network and attention network, initializing the replay buffer, and setting the maximum number of training epochs.

[0252] Specifically, the local policy network is initialized with parameter θ; the attention network is initialized with parameter w; and two independent hybrid networks are initialized, each containing an evaluation hybrid network and a target hybrid network; the parameters of the two evaluation hybrid networks are φ1 and φ2, respectively, and the parameters of the corresponding two target hybrid networks are as follows: and

[0253] In this step, the replay buffer D is initialized, the replay buffer size is set, and the replay buffer is used to train the policy and value network; the maximum number of training epochs M is set.

[0254] Step S4-2: Each agent obtains the global state and individual observations of the current and next time moments through observation, determines and executes actions by the local policy network, interacts with the environment, obtains global rewards, and adds the global state, individual observations, determined actions, global state and individual observations of the next time moment to the replay buffer.

[0255] Step S4-3: Randomly sample data from the replay buffer to obtain sampled data, and update the parameters of the policy network based on the sampled data.

[0256] The steps of updating the parameters of the policy network based on the sampled data specifically include:

[0257] The local Q-network and attention network are updated using the following formula:

[0258]

[0259] Where ← indicates assignment or update, α Q This indicates the update step size, and j2 represents the hybrid network number. This represents the gradient operator for evaluating hybrid networks. This represents the loss function used to evaluate the hybrid network;

[0260] The local policy network is updated using the following formula:

[0261]

[0262] Where, α π Indicates the update step size. L represents the gradient operator of a local policy network. π (θ) represents the loss function of the local policy network;

[0263] Update the hybrid network using the following formula

[0264]

[0265] Where τ represents the soft update weight, This represents the target hybrid network parameters.

[0266] Step S4-4: Return to step S4-2 to execute the next round until the maximum number of training rounds is reached, and obtain the trained policy network.

[0267] In this step, after multiple rounds of training, the final local policy network and the parameters Q of the policy network are obtained. φ ,π θ , as a trained policy network.

[0268] In particular, to further simulate real-world scenarios, this invention truncates useless segments in a given round and sets a working state flag s for each agent during training. work (i), whose initial value is s work (i) = working, if and only if the working status flag of agent i is s. work (i) When working, this invention performs normal state updates, action selection, and reward calculation on agent i; when agent i successfully reaches its corresponding target navigation point, that is, when the second term of formula (1.24) is satisfied, this invention considers agent i to have successfully completed the task in the current game round and sets its working state flag s. work (i) = success, and will not participate in the state update, action selection, or reward calculation of agent i thereafter; when agent i touches an external obstacle, that is, when the second term of formula (1.26) is satisfied, the present invention considers agent i to have died in the task of the current game round, and sets its working state flag s. work (i) = dead, and will not participate in the state update, action selection, or reward calculation of agent i thereafter. This design makes the agent's behavior in game rounds more consistent with real-world task scenarios and avoids redundant and invalid data after completing the task or dying prematurely, thus improving the overall training efficiency.

[0269] In this step, the present invention executes the MASAC-A algorithm, and the specific execution flow of the MASAC-A algorithm is shown in Table 1.

[0270] Table 1

[0271]

[0272] Step S5: The decision action output by the trained policy network controls the air-to-ground subgroup to navigate, thereby resolving conflicts in the air-to-ground unmanned cluster.

[0273] The policy network trained by this invention can be deployed on drones in a real environment. The decision actions output by the trained policy network control the air-to-ground sub-swarm for navigation, thereby resolving conflicts in the air-to-ground drone swarm.

[0274] To illustrate the effectiveness of the method proposed in this invention, the following detailed description of the above technical solution of this invention is provided through a specific embodiment.

[0275] Example 1

[0276] This embodiment discloses a method for resolving conflicts in air-to-ground unmanned swarms based on multi-agent reinforcement learning.

[0277] The specific implementation steps are as follows:

[0278] 1. Set the parameters required for resolving conflicts in an unmanned cluster scenario on an open ground.

[0279] The conflict resolution scenario for air-to-ground unmanned swarms mainly consists of four types of entities: obstacles, drones, unmanned vehicles, and target navigation points. Figure 6 As shown.

[0280] The specific parameters for the air-to-ground unmanned cluster conflict resolution scenario are shown in Table 2.

[0281] Table 2

[0282]

[0283]

[0284] 2. Training strategy for resolving conflicts in unmanned air-to-ground clusters

[0285] For drone group i, the observed state O i The input policy network outputs corresponding actions in the action space A for the UAV to execute and calculates the corresponding reward r. The MASAC-A multi-agent reinforcement learning framework is used to update the parameters of the policy network until the network parameters converge.

[0286] The hyperparameters of the MASAC-A multi-agent reinforcement learning framework are shown in Table 3.

[0287] Table 1

[0288]

[0289] 3. Output and analyze simulation results

[0290] Some training curves of MASAC-A are as follows: Figure 7 , Figure 8 , Figure 9 As shown in the figure, the MASAC-A algorithm has converged within the reward function framework of this invention, with its reward curve converging at 285 (2.2) and the Critic loss function converging at 1.1 (0.08).

[0291] Snapshots and partial data curves from the simulation of conflict resolution in unmanned air-to-ground clusters are shown below. Figure 10 , Figure 11 , Figure 12 As shown.

[0292] Depend on Figure 10 As shown, at t=0s, four groups of unmanned aerial vehicles (UAVs) started from the ground. Figure 4Each of the four groups of unmanned aerial vehicles (UAVs) randomly initializes 10 initial states for each UAV and begins navigation. At t=40s, each of the four groups encounters obstacles of different types. The upper-level UAV adjusts the elliptical virtual region based on perception information to achieve upper-level motion planning for the UAVs. At t=80s, the four groups of UAVs have basically completed their independent obstacle avoidance tasks and are about to perceive each other's UAVs and begin conflict resolution tasks. At t=100s, the four groups of UAVs perceive each other and maintain a certain degree of relative position and Lyapunov stability in the virtual region. However, there are strong collisions between adjacent UAVs in the navigation direction of their respective target points. At t=120s, the four groups of UAVs learn to rotate clockwise in coordination while maintaining relative position and Lyapunov stability in the virtual region to complete the conflict resolution of their respective target navigation tasks. At t=160s, the four groups of UAVs arrive at their respective target points and complete their respective navigation and obstacle avoidance tasks.

[0293] Figure 11 The dynamic distance changes between the four air-to-ground unmanned clusters and their nearest neighbor clusters are displayed throughout the mission. It can be seen that the distance between the unmanned clusters remains greater than 20m throughout the entire process, maintaining the internal security between the air-to-ground unmanned clusters during conflict resolution.

[0294] Figure 12 The dynamic distance changes between the four groups of unmanned aerial vehicles (UAVs) and their corresponding target points throughout the mission are shown. It can be seen that the distance between the UAVs and their targets generally decreases throughout the process, with only minor fluctuations at certain moments due to internal collision avoidance and external obstacle avoidance. Ultimately, they successfully reach their respective target points, achieving distance convergence.

[0295] While the specific embodiments of the present invention depict actions or steps in a particular order, this should be understood as requiring such actions or steps to be performed in the shown specific order or sequential order, or requiring all illustrated actions or steps to be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations. The above descriptions are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention.

[0296] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Step S1: Establish an air-ground unmanned cluster navigation environment, including: identifying obstacles and multiple air-ground subgroups, each air-ground subgroup having a target navigation point, the routes from each air-ground subgroup to its target navigation point having conflicting locations, and each air-ground subgroup not being able to pass through the conflicting locations simultaneously; Step S2: Determine the state space, action space, and reward / penalty function of the air-ground subgroup's actions based on the air-ground unmanned cluster navigation environment. Step S3: Establish the policy network, including the local Q-network, local policy network, hybrid network, and attention network; determine the network loss function; Step S4: Based on the state space, action space, reward and penalty function of the action of the empty-ground subgroup, and network loss function, perform reinforcement learning training on the policy network to obtain the trained policy network. Step S5: The decision action output by the trained policy network controls the air-to-ground subgroup to navigate, thereby resolving conflicts in the air-to-ground unmanned cluster.

2. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 1, characterized in that, In step S1, each of the airspace subgroups includes a drone and multiple unmanned vehicles; the drone is able to detect unmanned vehicles in the cone-shaped area directly below it, and the unmanned vehicles are able to detect other unmanned vehicles in the surrounding circular area; Step S1 also includes: determining the maximum speed and maximum acceleration of the drone and the unmanned vehicle, determining the perception radius of the drone and the unmanned vehicle, and determining the minimum distance and expected distance of the drone and the unmanned vehicle.

3. The air-to-ground unmanned swarm conflict resolution method based on multi-agent reinforcement learning according to claim 2, characterized in that, Step S2 specifically includes: Step S2-1: Based on the position, velocity, major axis length, tilt angle, and obstacle position of the virtual elliptical region where the air-ground subgroup is located, determine the content of the air-ground subgroup's own state information and external observation information. Step S2-2: Determine the action range of the air-to-ground subgroup based on the UAV's acceleration control parameters, the major axis of the elliptical region, and the tilt angle range. Step S2-3: Determine the target navigation reward, external obstacle collision penalty, internal virtual area collision penalty, virtual area planning tilt angle change penalty, and virtual area roundness deviation penalty, and synthesize them to obtain the reward and penalty function.

4. The air-to-ground unmanned swarm conflict resolution method based on multi-agent reinforcement learning according to claim 3, characterized in that, Step S2-2 specifically includes: UAVi represents the i-th drone in an air-to-ground drone swarm, and UAVi's actions. Defined as: a i =(u i ,L i ,i i )∈{(u,L,θ)|-a max ≤u i ≤a max ,L mix ≤L≤L max ,-θ max ≤θ≤θ max } Among them, a i Indicates the action of UAVi. u represents a real number field with dimensions of 1×4. i L represents the acceleration of UAVi. i ,θ i L represents the major axis and inclination angle of the virtual elliptical region corresponding to UAVi. min ,L max Let θ represent the minimum and maximum values ​​of the major axis of the virtual elliptical region, respectively. max This indicates the maximum tilt angle.

5. The air-to-ground unmanned swarm conflict resolution method based on multi-agent reinforcement learning according to claim 4, characterized in that, In step S2-3: The target navigation reward is used to reward the air-ground subgroup for approaching the corresponding target navigation point; the external obstacle collision penalty is used to punish the air-ground unmanned subgroup for approaching obstacles; the internal virtual region collision penalty is used to punish the overlap between the virtual elliptical regions of the unmanned subgroup; the virtual region planning tilt angle change penalty is used to punish the change in the orientation of the virtual elliptical regions of the unmanned subgroup at different times; the virtual region roundness deviation penalty is used to punish the behavior of the virtual elliptical regions of the unmanned subgroup deviating excessively from the circle in geometric shape. The reward and punishment function r i The expression for (k) is: r i (k)=α1r T +βr o +λr i +μr θ +ωr r Where, r T ,r o ,r i ,r θ ,r r α1, β, λ, μ, ω represent the target navigation reward, external obstacle collision penalty, internal virtual area collision penalty, virtual area planning tilt angle change penalty, and virtual area roundness deviation penalty, respectively, with adjustment weights.

6. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 5, characterized in that, Step S3 specifically includes: Step S3-1: Establish a local Q-network, a local policy network, a hybrid network, and an attention network. The agent obtains the global state and individual observations through observation. Individual observations are split into fixed-dimensional features and a variable entity set. The variable entity set is input into the attention network to obtain local features. The local policy network includes two sets of local policy networks. The local features are input into the local Q-network and two sets of local policy networks respectively to obtain local Q-values ​​and actions; the hybrid network receives the local Q-values ​​and the global state to obtain the total value. Step S3-2: Determine the loss function of the local Q network, the loss functions of the two local policy networks, the loss function of the attention network, and the loss function of the hybrid network.

7. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 6, characterized in that, Step S3-2 specifically includes: The loss function of a local Q-network is defined as: in, Represents the loss function. R represents the expected value based on the sampled data. t Indicates the single-step reward, and γ represents the discount factor. This represents the estimation of the target's global Q-value. Indicates the evaluation of the global Q-value, s t ,τ t ,a t represents the state, trajectory, and action at time t, respectively, and j2 represents the number of the hybrid network; The loss function of an attention network is defined as: in, Represents the loss function of the attention network; The loss function for the two local policy networks is defined as follows: in, Let π(a) represent the loss function of a local policy network. t ∣τ t ) represents the global specific strategy at time t; q mix Represents the q-value of the hybrid network. Indicates π i The mathematical expectation of the payoff under the strategy, This represents the individual's specific strategy at time t. This represents the global Q-value under the target parameters; The loss function of a hybrid network is defined as: Where L(α) represents the loss function of the hybrid network, Indicates based on a t ~π t Mathematical expectation under action distribution It represents the mathematical expectation of a uniform distribution.

8. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 7, characterized in that, Step S4 specifically includes: Step S4-1: Initialize network parameters, including initializing the parameters of the local Q network, local policy network, hybrid network and attention network, initializing the replay buffer, and setting the maximum number of training epochs; Step S4-2: Each agent obtains the global state and individual observations of the current and next time moments through observation, determines and executes actions by the local policy network, interacts with the environment, obtains global rewards, and adds the global state, individual observations, determined actions, global state and individual observations of the next time moment to the replay buffer. Step S4-3: Randomly sample data from the replay buffer to obtain sampled data, and update the parameters of the policy network based on the sampled data; Step S4-4: Return to step S4-2 to execute the next round until the maximum number of training rounds is reached, and obtain the trained policy network.

9. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 8, characterized in that, Step S4-3 specifically includes: The local Q-network and attention network are updated using the following formula: Where ← indicates assignment or update, α Q This indicates the update step size, and j2 represents the hybrid network number. This represents the gradient operator for evaluating hybrid networks. This represents the loss function used to evaluate the hybrid network; The local policy network is updated using the following formula: Where, α π Indicates the update step size. L represents the gradient operator of a local policy network. π (θ) represents the loss function of the local policy network; Update the hybrid network using the following formula Where τ represents the soft update weight, This represents the target hybrid network parameters.

10. The method for resolving conflicts in unmanned air-to-ground swarms based on multi-agent reinforcement learning according to claim 9, characterized in that, Step S4 further includes: Set the working status flags for each UAVi work (i), with an initial value of s work (i) = working; when agent i successfully reaches its corresponding target navigation point, set s. work (i) = success; when agent i touches an external obstacle, set s. work (i) = dead; If and only if the working status flag of agent i is s work When (i) = working, perform steps S4-2 and S4-3 on it.

Citation Information

Cited By

  • A space-time grid closed-loop multi-machine obstacle avoidance anchor point allocation system and method

    CN122347887A

  • A space-time grid closed-loop multi-machine obstacle avoidance anchor point allocation system and method

    CN122347887B