Multi-agent formation control method and device based on graph attention network and residual dynamic capability perception sharing super network
By employing a multi-agent formation control method using graph attention networks and residual dynamic capability-aware shared supernetworks, the problems of slow training speed and difficulty in representing agent behavior independence in large-scale formation scenarios are solved, achieving efficient collaboration and reliable task execution.
Patent Information
- Application Number
- CN202511196750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-21
AI Technical Summary
Existing multi-agent formation control methods suffer from slow training speed, difficulty in convergence, and difficulty in representing the independence of agent behavior in large-scale formation scenarios, especially in the cooperative control of UAV/ship swarms, where global observation is difficult to achieve.
A multi-agent formation control method based on graph attention network and residual dynamic capability perception sharing supernetwork is adopted. By constructing a perception compression model and a decision model, and combining it with a proximal policy optimization algorithm, the efficient coordination and reliable task execution of the multi-agent formation are achieved.
It achieves efficient collaboration and reliable task execution in formation tasks of different sizes, and improves training efficiency and the ability to represent agent behavior independently.
Smart Images

Figure CN120993962A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-agent control, and particularly to a multi-agent formation control method and device based on a graph attention network and a residual dynamic capability perception shared super network. BACKGROUND
[0002] Existing multi-agent formation control research mainly focuses on the formation control of homogeneous unmanned platforms. In particular, in the field of formation control using multi-agent reinforcement learning (MARL), researchers generally design reasonable reward functions to encourage the agent group to achieve efficient collaboration during task execution while exhibiting the ability to adapt to environmental dynamics. The advantage of MARL is that it can optimize the interaction strategy between agents in a data-driven manner and fully utilize the collaborative characteristics of individuals and groups in complex tasks. According to the difference in information dependency between the training and execution stages, multi-agent formation control methods based on deep reinforcement learning can be roughly divided into two categories: (1) centralized execution, all agents can obtain global information, including the position and state information of all other agents, so as to directly refer to global data when selecting strategies; (2) centralized training and decentralized execution (CTDE), that is, global information is used to optimize strategies during centralized training, while in the execution stage, each agent can only make decisions based on its local information.
[0003] Although centralized control has shown good results in large-scale multi-agent formation control, it relies on global observation for training and execution, resulting in high input vector dimension, slow training speed, and difficulty in convergence. In addition, global observation is difficult to achieve in practical applications, especially in complex scenarios such as unmanned aerial / ship swarm collaborative control, where agents often cannot obtain global information. In contrast, the CTDE method uses local observation in the execution stage, making it more in line with actual needs. However, traditional CTDE methods are generally applied to scenarios with a small number of agents and are difficult to extend to large-scale formation scenarios.
[0004] Parameter sharing in multi-agent deep reinforcement learning plays a crucial role in allowing algorithms to be extended to a large number of agents. Parameter sharing between agents significantly reduces the number of trainable parameters, shortens training time to a manageable level, and enables more efficient learning. However, having all agents share the same parameters can also have a negative impact on learning: on the one hand, this approach lacks strong theoretical support; on the other hand, this approach is not conducive to the independent representation of different agent behaviors and generally only works when different agents perform very similar tasks. SUMMARY
[0005] Embodiments of the present application mainly aim to provide a multi-agent formation control method and device based on a graph attention network and a residual dynamic capability perception shared super network, so as to solve at least one of the problems in the prior art, and to achieve efficient cooperation and reliable task execution of multi-agent in different scale formation tasks.
[0006] To achieve the above-mentioned purpose, one aspect of the embodiments of the present application provides a multi-agent formation control method based on a graph attention network and a residual dynamic capability perception shared super network, which comprises the following steps: constructing a multi-agent formation of unmanned aerial vehicles and unmanned surface vehicles; constructing a target reward function of the multi-agent formation; constructing a perception compression model through a graph attention network; constructing a decision model through a residual dynamic capability perception shared super network; obtaining an initial control model based on a proximal policy optimization algorithm according to the perception compression model and the decision model; and realizing control of the multi-agent formation according to the target reward function and the target control model.
[0007] To achieve the above-mentioned purpose, another aspect of the embodiments of the present application provides a multi-agent formation control device based on a graph attention network and a residual dynamic capability perception shared super network, which comprises: a first module for constructing a multi-agent formation of unmanned aerial vehicles and unmanned surface vehicles; a second module for constructing a target reward function of the multi-agent formation; a third module for constructing a perception compression model through a graph attention network; a fourth module for constructing a decision model through a residual dynamic capability perception shared super network; a fifth module for obtaining an initial control model based on a proximal policy optimization algorithm according to the perception compression model and the decision model; a sixth module for training the initial control model to obtain a target control model; and a seventh module for realizing control of the multi-agent formation according to the target reward function and the target control model.
[0008] The embodiments of the present application have at least the following beneficial effects: the present application provides a multi-agent formation control method and device based on a graph attention network and a residual dynamic capability perception shared super network, which constructs a multi-agent formation of unmanned aerial vehicles and unmanned surface vehicles; constructs a target reward function of the multi-agent formation; constructs a perception compression model through a graph attention network; constructs a decision model through a residual dynamic capability perception shared super network; obtains an initial control model based on a proximal policy optimization algorithm according to the perception compression model and the decision model; trains the initial control model to obtain a target control model; realizes control of the multi-agent formation according to the target reward function and the target control model, and achieves efficient cooperation and reliable task execution of a multi-agent system in different scale formation tasks. BRIEF DESCRIPTION OF DRAWINGS
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0010] Figure 1 This is a flowchart of a multi-agent formation control method based on graph attention network and residual dynamic capability perception shared supernetwork provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the drone control speed provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of unmanned surface vessel speed control provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the expected distribution of UAVs and unmanned surface vessels within a single formation in a small-scale formation, provided by an embodiment of the present invention. Figure 5 This is a schematic diagram of the environmental layout in a small-scale formation provided by an embodiment of the present invention; Figure 6 This is a schematic diagram of the environmental layout in a large-scale formation provided by an embodiment of the present invention; Figure 7 This is a schematic diagram of obstacle detection for unmanned surface vessels provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the RCASH structure provided in an embodiment of the present invention; Figure 9 This is an overall structural diagram of the GAT+IPPO model provided in this embodiment of the invention; Figure 10 This is an overall structural diagram of the GAT+MAPPO model provided in this embodiment of the invention; Figure 11 This is a schematic diagram illustrating parameter sharing in the sub-grouping implementation provided by an embodiment of the present invention; Figure 12 This is a comparison chart of the reward curves of the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models under small-scale formation provided in the embodiments of the present invention. Figure 13 This is a comparison chart of the reward curves of the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models under large-scale formation provided in the embodiments of the present invention. Figure 14 This is a diagram showing the decision results after training the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models in a small-scale formation, as provided in the embodiments of the present invention. Figure 15This is a diagram showing the decision results after training the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models under large-scale formation provided in the embodiments of the present invention. Figure 16 This is a flowchart of the training process for UAV GAT+IPPO and unmanned surface vessel GAT+MAPPO+RCASH provided in an embodiment of the present invention. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the terms "first / S100" and "second / S200" in the specification, claims, and the aforementioned drawings can be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this invention, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.
[0012] like Figure 1 As shown, this embodiment of the invention provides a multi-agent formation control method based on graph attention networks and residual dynamic capability perception shared supernetworks, which may include, but is not limited to, steps S100 to S700: Step S100: Construct a multi-agent formation of drones and unmanned surface vessels; Step S200: Construct the target reward function for the multi-agent formation; Step S300: Construct a perceptual compression model using a graph attention network; Step S400: Construct a decision model by using the residual dynamic capability-aware shared supernetwork; Step S500: Based on the near-end policy optimization algorithm, the initial control model is obtained according to the perception compression model and the decision model; Step S600: Train the initial control model to obtain the target control model; Step S700: Control of the multi-agent formation is achieved based on the target reward function and the target control model.
[0013] In step S100 of some embodiments, a multi-UAV and unmanned surface vessel (USV) formation environment is constructed to simulate formation tasks of different scales, and obstacles are set in the environment to improve scene fidelity. For example, a multi-UAV and USV formation environment is constructed by setting the number of formations, starting position, target position, number of USVs within the formation, control speed and formation distribution of UAVs and USVs, number of obstacles and random number seed generation, and detection rays of USVs, etc., in which a multi-UAV and USV formation can be set up.
[0014] Optionally, random obstacle polygons can be generated using a mass model and the Shapely library in Python. In the mass model, external forces are ignored; after making a decision, the UAV or unmanned surface vessel (USV) is assumed to move at a constant linear velocity, meaning its position at the next decision execution time is equal to its expected position. The Shapely library supports the creation, spatial relationship judgment, and operations (such as intersection, union, and buffer analysis) of basic geometric types like points, lines, and polygons. The core module, shapely.geometry, provides an intuitive API, allowing for the rapid construction of geometric objects using classes like Point, LineString, and Polygon, and supports functions such as area calculation, distance measurement, and topological relationship verification. This makes it suitable for generating obstacle polygons and performing obstacle detection calculations within a mass model environment.
[0015] In some embodiments, step S100 may include, but is not limited to, steps S101 to S105: Step S101: Preset the first number, starting position, and target position of the multi-agent formation; Step S102: Preset the second number of unmanned surface vessels in each multi-agent formation; Step S103: Preset the control speed and formation distribution of the multi-agent formation; Step S104: Preset the third number of obstacles and generate the obstacle distribution according to the random number seed; Step S105: Set up detection rays with equal angle distribution on the unmanned surface vessel.
[0016] In steps S101 to S102 of some embodiments, a first number of multi-agent formations, a starting position, a target position, and a second number of unmanned surface vessels (USVs) within the multi-agent formations are preset. Optionally, the multi-agent formations are divided into small-scale formations and large-scale formations, and the agents include drones and USVs.
[0017] For example, a small-scale formation is set up with 3 formations, each with 1 drone acting as the leader (subscript and superscript in full text). Used to indicate the navigator), 5 unmanned surface vessels act as followers (superscript for full text). (Used to represent followers), there are a total of 18 agents. The three formations are initially arranged in a triangle. Looking along the x-axis from the negative to the positive direction, formation 1 is in the middle, and formations 2 and 3 are equidistantly distributed on either side behind formation 1. Their target positions are assigned to formations 1, 2, and 3 to perform the tasks of moving forward, moving left-forward, and moving right-forward, respectively. What is the starting position of formation 1? and target location The starting position of formation 2 and target location The starting position of formation 3 and target location They are respectively: Formation 1: ; Formation 2: ; Formation 3: ; A large-scale formation is configured with 10 formations, each with one drone as the navigator and 10 unmanned surface vessels as followers, totaling 110 intelligent agents. The 10 formations initially form a circular distribution, evenly distributed on a circle with a radius of 500m. The target positions are evenly distributed on a circle with a radius of 1000m. Considering the 10 formations as a whole, they perform a distributed task. Therefore, the starting position of the large-scale formation n is... and target location It can be represented as: Formation n: ; In step S103 of some embodiments, the control speed and formation distribution of the drone and unmanned surface vessel are set. Exemplarily, both the drone and the unmanned surface vessel have discrete motion spaces, and the drone has five control speeds. These represent five decision actions: forward, left, right, backward, and stationary, with a speed range set to 1 m / s (e.g., ...). Figure 2 (As shown); the unmanned surface vessel is also designed with five control speeds. These represent five decision actions: moving forward, moving to the left front, moving to the right front, moving backward, and staying still, with a speed range set to 0.75 m / s (e.g., ...). Figure 3 As shown), then the following expression exists: ; ; Set decision execution time The time is 10s. In the point mass model, neglecting the influence of external forces, after making a decision, the UAV and unmanned surface vessel are assumed to move at a constant linear velocity according to the control speed. That is, their position at the next decision execution time is equal to the expected position. Therefore, the expected position is... =Current position Decision speed × Decision Execution Time In both small and large formations, the minimum displacement distance required for each formation is 500m. Considering the speed range of the UAV is 0.75m / s and the speed range of the unmanned surface vessel is 1m / s, and the decision-making execution time is 10s, based on the slower speed of the UAV as the lower limit of the number of decision steps, each decision can only result in a maximum displacement of 7.5m (excluding the case of staying in place). Therefore, the number of decision steps T needs to be at least 500m / 7.5m≈67. To make the entire decision-making process more diverse and allow the UAV to have the opportunity to choose a strategy of staying in place for adjustment, this embodiment of the invention selects a larger number of decision steps T=100.
[0018] In some embodiments, the three-dimensional position of a drone in a sub-formation is set as ,in, respectively drones in Coordinates on the axis. Since the control of the drone actually only uses a point mass model, we will only discuss the two-dimensional drone, with the two-dimensional projection point of the drone navigator. The speed is used to characterize the drone, thus simplifying the discussion of the drone to a two-dimensional space similar to that of an unmanned surface vessel. The location of the drone's two-dimensional projection point. Defined as in The plane, that is, the projection of the unmanned surface vessel onto the sea surface, is: drones in the formation Corresponding target point It can be represented as: For unmanned surface vessel (USV) swarms within sub-formations, unmanned surface vessels The corresponding sub-target point location (i.e., the desired location). It can be represented as: To maintain a proper formation between the navigator and followers, the followers should be evenly distributed around the navigator. On a horizontal plane, the followers should be distributed at equal angular intervals within a circle centered on the navigator with a radius of [missing information]. On the circumference. Set the initial angle offset. Then the expected angle of each follower is: ; in, The total number of unmanned surface vessels in the sub-formation; Let be the desired radial distance, representing the expected distance between each follower and the leader, i.e., the formation radius. On a two-dimensional plane, the radial vector... This represents the vector pointing towards the formation circle in the navigator's coordinate system, and can also represent the expected position of a follower relative to the navigator, defined as: ; In order to obtain the desired position of each follower Rotation matrices can be used To calculate the displacement of the follower relative to the leader. For example, the... The relative displacement of the followers It can be obtained by multiplying the rotation matrix and the radial vector: ; Wherein, rotation matrix Indicates an angle in a two-dimensional plane The rotation matrix is defined as: ; Therefore, followers Expected position It can be accessed through the navigator's projection point With rotational displacement The superposition of these elements gives the following expression: ; Expanding further, one can gain followers. The desired position coordinates are in the following form: ; In subsequent model training, to improve computational efficiency, vectorized calculation of all follower sub-target points can be used. Define the angle vector. : ; in, The matrix consisting of all the sub-target points of the followers is: ; in, It is a vector of all 1s. and For vectors that are operated on element by element.
[0019] In some embodiments, in small-scale formations, a formation radius is set. The danger radius of collisions between drones and unmanned surface vessels is... ,like Figure 4 The diagram shows the expected distribution of UAVs and unmanned surface vessels (USVs) within a single small-scale formation. The black dots represent the projections of the navigator UAV onto the plane where the USV is located. The green circle represents the formation area, where the unmanned surface vessels (USVs) are expected to be distributed at equal angles. The red circle represents the danger zone for the USVs. In large-scale formations, the formation radius is set. The danger radius of collisions between drones and unmanned surface vessels is... .
[0020] In step S104 of some embodiments, a third number of obstacles is set and a random number seed is generated to generate an obstacle distribution. Optionally, 15 obstacles are set in a small-scale formation and 50 obstacles are set in a large-scale formation. The deployment range is determined by the starting position and target position of the UAV. The obstacle deployment range in a small-scale formation is as follows: ; The obstacle deployment range in large-scale formations has been slightly expanded: ; To prevent decision-making tasks from falling into dead zones, a certain area centered on the starting and final target positions of each unmanned surface vessel (USV) is excluded from the deployment range; these areas are collectively referred to as restricted zones. For example, given a given USV... Acquire unmanned surface vessels The starting position of the follower at the initial moment. and the final target location And set the initial restricted area at the initial moment. and the ultimate target restricted area : ; ; in, The danger radius can be optionally set to 20m, making the danger radii the same for both drones and unmanned surface vessels. This then defines the no-go zone formed by the unmanned surface vessels. for: ; In some embodiments, the starting position of each drone will be used as the center. The area within the radius is excluded from the deployment range, and the target location of each UAV is used as the center. This excludes areas within the radius of the target formation from the deployment range, thus preventing obstacles from interfering with the start and completion phases of the mission by being placed within the initial and target positions of the formation. This represents the distribution radius of the formation at the initial moment. Assuming the decision-making task can ensure that the UAVs reach the vicinity of the target point during the later stages of training, and the unmanned surface vessels also reach the vicinity of their respective sub-target points according to the predetermined formation, given a certain UAV... According to the starting position and target location The expression, drone As the navigator, the initial starting position is The target location is The following initial restricted areas and final target restricted areas are set respectively: ; ; The formation exclusion zone, that is, the exclusion zone composed of drones. for: ; Based on the restricted areas formed by unmanned surface vessels and unmanned aerial vehicles, the total restricted areas can be calculated as follows: ; In some embodiments, the obstacle generation steps are as follows: Step 1.1: Within the deployment area A center point is randomly generated within the region. Randomly select an integer within the given vertex range [3,8] as the number of vertices of the obstacle polygon. Calculate the average angle score for each vertex. .
[0021] Step 1.2: For the vertices of the obstacle polygon , Angle center value Within the angular offset range Randomly select a floating-point number As an angular offset, the vertex is calculated. actual angle And select an integer within the given radius range [5,9]. And randomly select a floating-point number within the radius offset range [0.7, 1.3]. As the radius offset rate, the radius offset rate is compared with the radius. Multiply them, and the result is used as the radius of the vertex. The coordinates of the vertex are obtained using the polar coordinate calculation method: ; ; The coordinates of the vertices can be written in matrix form as follows: ; Iterate through each vertex to obtain the coordinates of all vertices, and then apply the buffer coefficient. For the polygon defined by its vertices, an outwardly expanding buffer is generated, resulting in a new polygonal region that surrounds the original polygon. Rounded corners are added to increase the smoothness of the new region, thus generating the final obstacle range. .
[0022] Step 1.3: Determine and Is there any overlap? Then the range of the obstacle shall be used. If the obstacle is not selected, the process returns to step 1.1 to continue generating obstacles until the number of attempts exceeds the maximum number of attempts (set to 10 times the number of obstacles). (or the number of obstacles that meet the conditions reaches the required number of obstacles) Until the requirements are met.
[0023] Ultimately, the initial distributions of small-scale and large-scale formations are as follows: Figure 5 and Figure 6 As shown. The pentagrams represent the target points of the corresponding formations, and the gray polygons represent obstacles.
[0024] In step S105 of some embodiments, detection rays are set at equally angularly distributed for each unmanned surface vessel to detect obstacle information. For example... Figure 7 As shown, unmanned surface vessel There are a total Strips are distributed at equal angles, with a maximum length of Detection rays. Figure 7 There are 3 detection rays (such as) Figure 7 The length of the three red lines in the image is less than [the length of the three red lines in the image]. This indicates that there is an obstacle in that direction, and the distance between the unmanned surface vessel and the obstacle is... , Indicates the first A detection ray.
[0025] In step S200 of some embodiments, a target reward function is designed for the agent formation based on the task objective of multi-agent cooperation. This target reward function considers not only the cooperation and task completion degree between agents, but also factors such as formation stability and collision avoidance, ensuring that each agent employs the optimal strategy when performing the task. In the target reward function designed for agent formation, a leader-follower model is selected. Formation design and control are performed based on tasks such as path planning, formation maintenance, and obstacle avoidance. Optionally, the target reward function is obtained by designing proximity rewards, position rewards, direction rewards, collision avoidance rewards, arrival rewards, and obstacle rewards.
[0026] In some embodiments, step S200 may include, but is not limited to, steps S201 to S210: Step S201: Obtain proximity reward based on the agent's first position at the current moment, the agent's second position at the previous moment, and the agent's target position at the current moment; Step S202: Obtain the location reward based on the first location and the target location; Step S203: Obtain the first direction vector from the second position to the target position; Step S204: Obtain the second direction vector from the second position to the first position; Step S205: Obtain the direction reward based on the angle between the first direction vector and the second direction vector, the first position, and the second position; Step S206: Preset the danger radius of the intelligent agent collision and obtain the first distance between the intelligent agent and the target position at the current moment; Step S207: Obtain the arrival reward based on the first distance and the danger radius; Step S208: Obtain collision avoidance reward based on the danger radius and the second distance between agents; Step S209: Obtain obstacle reward based on the third distance between the unmanned surface vessel and the obstacle; Step S210: Obtain the target reward function based on proximity reward, position reward, direction reward, collision avoidance reward, arrival reward, and obstacle reward.
[0027] In steps S201 to S210 of some embodiments, for the Navigator UAV, the altitude information of the target position is ignored in this embodiment of the invention; the UAV only needs to reach the target position in the x and y axis planes. Therefore, for the Navigator UAV: (1) Design a reward close to the drone exist Compared to time By getting closer to its target location, a drone can be obtained. Positive rewards are awarded based on the distance traveled; conversely, if the drone... exist Compared to time The further away from the target location a drone is at any given moment, the greater the penalty will be. Proximity reward Defined as: ; in, yes Time Drone The first position; yes Time Drone The second position; yes Time Drone The target location.
[0028] (2) Design a positional reward system to penalize the drone for remaining stationary before reaching the target location. The farther the drone's current location is from the target location, the greater the penalty. (Drone) Location Rewards Defined as: ; (3) Design a directional reward system to penalize drones whose speed and direction deviate from the desired direction; the greater the deviation angle, the greater the penalty. (Drone) Directional rewards Defined as: ; In the formula, It is the angle between the first direction vector and the second direction vector, where, for the Navigator UAV, the first direction vector is the angle between the UAV and the second direction vector. exist Location of drones at any time exist The direction vector of the target position at a given time, and the second direction vector is the UAV's... exist Location of drones at any time exist The direction vector of the position at that moment. Then... The definition of is: ; (4) Design collision avoidance rewards to penalize drones. Other drones This poses a collision risk. If other drones... In drones Within the danger radius, drones Distance from other drones The closer, the more drones The greater the punishment, the better. (Drones) Collision avoidance reward Defined as: ; in, This indicates that all other drones Perform summation; Danger radius; for Time Drone With drones The second distance. Collision avoidance rewards ensure [safety] within the drone's range. Enter other drones Penalties are imposed when the drones are within a safe threshold, thereby encouraging them to maintain a safe distance and preventing collisions between the drones and unmanned surface vessels in their sub-formation.
[0029] (5) Design an arrival reward: if the drone reaches a certain range of the target location, it will receive a positive reward to encourage the drone to maintain its position after reaching the target location. Drone Arrival reward Defined as: ; In the formula, ; in, for Time Drone The initial distance to the target location is ignored at this point; the drone only needs to be at the target location's altitude. Reaching the target position within the axis plane is sufficient. Reaching the item to the left of the reward... Used for judgment Is it less than the threshold? When the value is less than the threshold, the numerator and denominator cancel each other out to get 1, and the final result is the term on the right. The item on the right uses The function structure allows for larger rewards the closer to the target point. However, an upper limit coefficient is set to prevent the rewards from becoming too large. Ensure that the summation of any single term in the formula does not exceed a certain value. And when Greater than the threshold When this happens, the result of the item on the left is 0, and no reward is generated. The threshold is set to half of the danger radius, i.e. This prevents the reward from having too broad an impact and encourages drones to stay more precisely near the target point.
[0030] For follower unmanned surface vessels: (1) Design a reward close to the target, if unmanned surface vessel exist Compared to time By getting closer to its target location, a relationship with the unmanned surface vessel can be obtained. Positive rewards are directly correlated with the distance traveled; conversely, if the unmanned surface vessel... exist Compared to time The further away from the target location you are from at any given time, the greater the penalty will be. Unmanned surface vessel (USV) Proximity reward Defined as: ; in, yes Unmanned Surface Vessel The first position; yes Unmanned Surface Vessel The second position; yes Unmanned Surface Vessel The target location.
[0031] (2) Design a positional reward system to penalize the unmanned surface vessel (USV) for remaining stationary before reaching the target point. The farther the USV's current position is from the target position, the greater the penalty. Location Rewards Defined as: ; (3) Design a directional reward system to penalize unmanned surface vessels (USVs) for deviating from the desired direction in speed and direction; the greater the deviation angle, the greater the penalty. Unmanned Surface Vessel Directional rewards Defined as: ; in, It is the angle between the first direction vector and the second direction vector. For the follower unmanned surface vessel (USV), the first direction vector is the angle between the USV and the second direction vector. exist The direction vector from the current position to its target position, and the second direction vector is the unmanned surface vessel. exist The position of the moment to its location The direction vector of the position at that moment. Then... Defined as: ; (4) Design collision avoidance rewards to penalize unmanned surface vessels. Other unmanned surface vessels This poses a collision risk. If other unmanned surface vessels... On unmanned surface vessel Within the danger radius, unmanned surface vessels Distance from other unmanned surface vessels The closer you are, the greater the punishment. Unmanned surface vessel. Collision avoidance reward Defined as: ; in, This indicates that all other unmanned surface vessels... Perform summation; Danger radius; for Unmanned Surface Vessel Unmanned Surface Vessel The second distance. Collision avoidance rewards ensure safety on the unmanned surface vessel. Enter other unmanned surface vessels Penalties are imposed when the distance is within a safe threshold, thereby encouraging unmanned surface vessels (USVs) to maintain a safe distance and avoid collisions between them.
[0032] (5) Design an arrival reward: if the unmanned surface vessel (USV) reaches a certain range from the sub-target point, it will receive a positive reward to encourage the USV to reach the sub-target point more accurately and maintain its position. Unmanned Surface Vessel Arrival reward Defined as: ; In the formula, ; in, for Unmanned Surface Vessel The first distance from the target location.
[0033] (6) An obstacle reward system is designed to prevent the unmanned surface vessel (USV) from colliding with obstacles. If the USV moves away from an obstacle, it receives a positive reward; otherwise, it incurs a negative penalty. Additionally, closer proximity to an obstacle incurs an extra penalty. The obstacle reward is defined as: ; in, for Unmanned Surface Vessel The The length of the detection ray. In the formula for obstacle reward, the first two terms... Reflects unmanned surface vessels The change in the total length of the detection ray, as the total length changes with time... When the length increases, it indicates that the unmanned surface vessel (USV) is moving away from the obstacle as a whole, and the first two terms calculate a positive reward; conversely, a negative penalty is given. Therefore, the first two terms are mainly used to incentivize the USV to move away from the obstacle. Since the point mass model does not model the collision between the USV and the obstacle, using only the first two terms as obstacle rewards may lead some USVs to choose a locally optimal solution that directly passes through the obstacle. Therefore, a correction term is constructed. This avoids the local optimal solution where the unmanned surface vessel chooses a path that directly passes through obstacles, thus facilitating the subsequent migration process from a point mass model to a simulation environment with collision modeling. The coefficients in the correction term... In the length of the probe ray When the radius is less than the danger radius, it means that the unmanned surface vessel is "activated" when it is about to collide with an obstacle in a certain direction, thus achieving the right-hand item. Additional penalty, length of the detection ray The smaller the value, the greater the penalty. This is achieved by introducing a penalty cap coefficient. To prevent excessive penalties from causing system instability, the maximum penalty is limited to no more than [a certain value]. .
[0034] In step S300 of some embodiments, a Graph Attention Network (GAT) is used as a perceptual compression model, which is used for feature aggregation of neighboring nodes. This is achieved by inputting the original node vector matrix. and adjacency matrix By using the perceptual compression model, the updated node vector of the output can be obtained. This refers to the perceptual compressed feature vector. The adjacency matrix is a matrix used to represent the connection relationships between nodes in a graph (or network). For example, the main steps are as follows: Step 3.1: Calculate the attention coefficient. GAT uses different nodes to represent different agents and their probe rays. Each node calculates the attention coefficient for its neighboring nodes. , representing a node (i.e., intelligent agent) ) for nodes (i.e., neighboring intelligent agents) The level of attention given to ). Assuming Feature dimension is , Feature dimension is Then concatenate the feature vectors The feature dimension is Subsequently, it is compared with trainable vectors of the same feature dimensions. Perform attention weighting (i.e., ... After transpose and Multiply. Then the nodes and nodes Attention coefficient Defined as: ; in, It is a trainable weight matrix used to map the features of nodes to a new space. and These are nodes and nodes Features mapped to a new space. It is a trainable vector used to calculate the attention coefficients between nodes. This is achieved by concatenating the feature vectors of node pairs. After performing a linear transformation, the attention coefficients are calculated using a dot product operation. It is for each node For all neighboring nodes Perform a softmax operation to ensure the attention coefficient. The sum of these is 1. Therefore, the expression for the attention coefficient can be written as: ; in, Indicates that, except for the numbered The number of intelligent agents other than the intelligent agent itself.
[0035] Step 3.2: Aggregate the features of neighboring nodes. For a node... GAT will aggregate the features of neighboring nodes by weighting, where the features of each neighbor are weighted according to their corresponding attention coefficient. Weighted. Then the node In the The feature update rule for the layer is: ; in, For nodes In the The feature matrix of the layer; For nodes In the The feature matrix of the layer; For the first The trainable weight matrix of the layer; The activation function is used to introduce more nonlinear components into the model. Optionally, the activation function can be the ELU function, whose calculation rules are as follows: ; By reducing the impact of bias offset through the ELU function, the normal gradient is made closer to the unit natural gradient, thereby accelerating learning towards zero mean, enhancing learned features, and mitigating the gradient vanishing problem. It can be set to the default value of 1.0.
[0036] Step 3.3: Node Information Activation and Suppression. Calculate the next layer node feature matrix. Previously, adjacency matrices were used. Node feature matrix Activation and inhibition operations are performed to leverage the advantages of the graph network's adjacency topology. Optionally, the adjacency relationship between agents is determined based on set conditions; if two agents are not associated, the adjacency matrix is updated accordingly. The value at the corresponding position in the middle is set to If two agents are related, the value at the corresponding position is set to 1. In this way, an adjacency matrix can be constructed. Thus, the node feature matrix Activation and inhibition operations are performed. The following expression holds: ; in, For nodes In the The result matrix of the layer, through Activated or inhibited after dot product; For nodes In the The feature matrix of the layer; This is a matrix dot product operation, where the two matrices being multiplied have the same dimensions, and the result is obtained by multiplying element-wise. Similarly, we have... In the formula, For nodes In the The result matrix of the layer, through Activated or inhibited after dot product; For nodes In the The feature matrix of the layer.
[0037] In some embodiments, assuming there are three agents, and the relationships between them are as follows: agent 1 and agent 2 are directly related, agent 2 and agent 3 are directly related, and agent 1 and agent 3 are not directly related, then there is an adjacency matrix. In the adjacency matrix, 0 indicates that the agent is not associated with itself.
[0038] Step 3.4: Node Information Compression and Mapping. The GAT network can employ a multi-head attention mechanism to enhance the model's expressive power and stability by representing the high-dimensional input feature vector. After weighted aggregation through the multi-head attention mechanism, the input feature vector is mapped to a lower-dimensional dynamic obstacle avoidance feature space. This design ensures that the output vector contains sufficient information while significantly reducing data storage and transmission overhead. In the multi-head attention mechanism, multiple attention heads independently calculate their respective attention coefficients and feature aggregations, and then the results are combined (usually concatenated or averaged). For example, nodes using the multi-head attention mechanism... In the The feature update rule for a layer can be expressed as: ; in, Indicates the number of attention heads; Indicates the first Attention coefficients calculated from individual attention heads; This represents a concatenation operation along the feature dimension.
[0039] Optionally, the above nodes In the The feature update rules for a layer are written in matrix form: ; In the formula, It is a dynamically generated attention weight matrix, where Represents a node and nodes Attention weights. It is the first The node feature matrix of the layer is initialized to the initial layer node feature matrix. , That is, the input node feature matrix contains the association information between drones and drones, and between unmanned surface vessels and unmanned surface vessels.
[0040] Then The outputs of each attention head are concatenated (or averaged) to obtain the matrix representation of node updates as follows: ; In the formula, Indicates the first The agent feature matrix of the layer; Indicates the activation function; Indicates the number of attention heads. ; Indicates the first Attention coefficient matrix of each attention head; Indicates the first The first attention head Layer feature matrix; Indicates the first The first attention head Layer linear transformation weight matrix; This indicates a splicing operation.
[0041] In step S400 of some embodiments, such as Figure 8 As shown, a Residual Capability-Aware Shared Hypernetwork (RCASH) is used as the decision model. The perceptual compressed feature vector output from the perceptual compressed model and the agent's observations are input into the decision model to generate neural network weights. These weights are then decoded along with the encoded vector generated from the agent's observations to obtain the decision and value estimation results. The constructed decision model includes: Step 4.1: Encode using an encoder. Capability-Aware Shared Hypernetwork (CASH) employs a network based on shared parameters. This is used to process local observations of the proxy and encode the observations. The mapping implemented by the encoder in CASH can be used... It means that among them The joint space of observations by all intelligent agents, i.e., global observations; This is the encoding vector space. Consider the time step... The Each intelligent agent, CASH will generate the encoded vector from the encoder. As input to the adaptive decoder, the expression for the encoded vector is: .
[0042] Step 4.2: Generate weights using a hyper adapter. To achieve sufficient flexibility in subsequent use of the adaptive decoder, RCASH combines the dynamic obstacle avoidance capability representation vector extracted from the GAT model, i.e., the perceptual compressed feature vector, and generates weights using a hyper adapter. At each time step Generate the first Adaptive decoder for each agent Neural network weights Then we have the following expression: In the formula, ;in, In time step intelligent agent Observations; In time step intelligent agent In this embodiment of the invention, the ability of GAT to perform time steps... Extracted perceptual compressed feature vector The capabilities of an intelligent agent are used to dynamically reflect the agent's capabilities. Its obstacle avoidance capabilities and performance.
[0043] Step 4.3: Decode using an adaptive decoder. Using the neural network weights provided by the super adapter The encoded vector generated by the encoder Decoding is achieved by performing matrix multiplication and addition. The decoded vector is then added to the encoded vector to complete the residual connection operation, and finally mapped to the vector space of value or policy. For example, the formulas used in an adaptive decoder include: ; ; In the formula, Indicates the first Each agent at each time step The output characteristics; Indicates an adaptive decoder; This represents the network weight of the super adapter; This represents the encoded vector generated by the encoder; This represents the output of the capability-aware shared hypernetwork; Indicates a fully connected layer; This represents the weight matrix in the neural network weights of the super adapter; This represents the bias matrix in the neural network weights of the super adapter. The operations in steps 4.1 through 4.3 allow CASH to flexibly encode policy formulations that can vary between different agents within a single network (i.e., the encoder, super adapter, and adaptive decoder in steps 4.1 through 4.3). Figure 8 As shown, a single network is replicated across different agents, meaning a single network simultaneously acts as an agent for several agents. These policies, based on different dynamic obstacle avoidance capabilities and observations, are processed through a super adapter to obtain different weights across different agents. These weights are then decoded using the encoded vectors obtained from the encoder, resulting in different policy outcomes (e.g., ...). Figure 8 The different colored adaptive decoder parts in different agents can obtain the policy or value estimate corresponding to their respective agents, even if the single network used is the same.
[0044] By adding a method similar to residual joins to CASH, we can obtain, for example... Figure 8 The RCASH structure shown adds an encoding vector to the output features. Furthermore, by removing the activation function to avoid disrupting the effect of the residual connections by using the activation function after the residual connections, the final output expression of the capability-aware shared supernetwork can be obtained as follows: .
[0045] In step S500 of some embodiments, based on the framework of the Proximal Policy Optimization (PPO) algorithm, the perceptual compression model is combined with the decision model to optimize the cooperative behavior of multiple agents. For example, each agent in PPO... Based on time Local observation and an independent strategy Generate an action To maximize discount accumulation rewards : ; in, Discount factor for future rewards; Indicates based on , I'm hoping for a draw; Indicates that in the state of Take action under the circumstances The benefits or rewards derived from actions. Each agent is based on its local state. Let's learn independent value functions .
[0046] There are two types of networks in PPO: Actor and Critic. The task of the Critic network is to learn the value function. That is, a state space Mapping to real numbers: The Actor network is responsible for learning the policy function. Specifically, the policy function Learning a mapping from observation The action mean and variance are mapped to a range (discrete action space) or to a Gaussian function (continuous action space) for subsequent action sampling. During policy updates, the Clipped Surrogate Objective (CPO) method from the PPO algorithm is used to limit the step size of policy updates, thereby enhancing model stability.
[0047] In some embodiments, each UAV agent independently optimizes its policy, while unmanned surface vessel (USV) agents share parameters within a formation. Optionally, the UAVs employ Independent Proximal Policy Optimization (IPPO), while the USVs in the UAV formation employ Multi-Agent Proximal Policy Optimization (MAPPO).
[0048] In some embodiments, step S500 may include, but is not limited to, steps S501 to S502: Step S501: Based on the independent near-end policy optimization algorithm and combined with the perception compression model, the initial control model of the UAV is obtained. Step S502: Based on the multi-agent near-end policy optimization algorithm, the perception compression model and the decision model are combined to obtain the initial control model of the unmanned surface vessel.
[0049] In step S501 of some embodiments, the UAV employs the Independent Proximal Policy Optimization (IPPO) algorithm combined with the Gaussian Compressed Atmosphere (GAT) model to obtain the initial control model GAT+IPPO for the UAV. In IPPO, each agent maintains a local policy function and a value function. During training, each agent independently samples and obtains trajectories, updating its own policy and value function based on these trajectories. The overall structure of the GAT+IPPO model is as follows: Figure 9 As shown.
[0050] In step S502 of some embodiments, the unmanned surface vessel (USV) in the UAV formation employs the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, combining the Gaussian Compressed Ability (GAT) model with the Decision Reduction (RCASH) model to obtain the initial control model of the USV: GAT+MAPPO+RCASH. In MAPPO, each agent has an independent Actor network but shares a centralized Critic network. The Critic evaluates the value of actions based on the global state, thus introducing global information to guide policy optimization during the training phase. This utilizes global information to guide decentralized policy optimization, rather than simply extending the single-agent algorithm. The overall structure of the GAT+MAPPO model is as follows: Figure 10 As shown. Since the tasks of different formations are different, the unmanned surface vessel (USV) intelligent agents share parameters within each formation, such as... Figure 11 As shown.
[0051] In some embodiments, for drones Both the Actor and Critic networks employ multi-layered neural network structures to generate optimal action policies and value functions for evaluating the current policy for each agent. The network consists of multiple fully connected layers, with activation functions (such as ReLU or Tanh) introduced between layers to enhance non-linear expressiveness. The final layer of the Actor network uses a Softmax activation function to map the output to the probability distribution of various actions, providing guidance for the agent's action selection. The Critic network takes the agent's state information as input and outputs a state-value function. This is used to measure the overall return of the current strategy in a given state. For example, at time t... drones Local state It is based on local observation and perceptual compression vector The resulting vector is obtained through merging, while the perceptual compression vector... It is composed of node feature matrix and adjacency matrix It is obtained through the GAT network. Therefore, the input layer of the UAV's multi-layer neural network is derived from... Mapped to the hidden layer, which is set to one layer with 64 neurons, in the Actor and Critic networks, the probability distribution and value estimate of various actions are finally obtained through the output layer. The ReLU function is used in the middle of the layer to increase the non-linear expressive power.
[0052] ; in, and These represent the multi-layer neural networks in the Actor and Critic networks, respectively. This represents the action output by the Actor network; This represents the value estimate of the Critic network output.
[0053] Combining the leader-follower model, for leader drones Assuming local observation The node feature matrix consists of the speed and position of the drone itself. Other drones The distance to its own drone is used to determine whether this distance is less than the drone's sensing radius. To construct the adjacency matrix At this point, the GAT model can be further understood as acting like radar, using the perception radius as the radar radius to detect other intelligent agents or obstacles within that radius. The perception radius of a drone. It should be greater than or equal to the formation radius. The embodiments of the present invention are set .
[0054] ; in, For other drones The location. When another drone is less than the sensing radius of this drone. When the adjacency matrix has an element of 1, it indicates a connection, meaning there is an association between the drone and another drone; when the distance between the drone and another drone is greater than or equal to the sensing radius... When the adjacency matrix is zero, the corresponding element is 0.
[0055] In some embodiments, for unmanned surface vessels Both the Actor and Critic networks employ an RCASH network structure to generate optimal action policies and value functions for evaluating the current policies for each agent. The network consists of an encoder, a hyper adapter, an adaptive encoder, and a final linear network layer, enabling flexible policy encoding and allowing policies to vary flexibly among different unmanned surface vessels (USVs) within a formation represented by a single network. The final layer of the Actor network uses a softmax activation function to map the output to the probability distribution of various actions, providing guidance for the agent's action selection. The Critic network takes the state information of all USVs in the formation as input and outputs the state value function of each USV. This is used to measure the benefits of the current strategy for each unmanned surface vessel under a given state.
[0056] ; In the formula, ; In this formula, each vector is formed by concatenating the vectors of several unmanned surface vessels within the formation, thus creating a global vector for the formation. and These represent residual dynamic capability-aware shared supernetworks in the Actor and Critic networks, respectively.
[0057] For drones Follower unmanned boat Assuming local observation In addition to the unmanned surface vessel's own position, it also includes the navigator drone. Velocity and position. Node feature matrix. Other unmanned surface vessels With its own unmanned surface vessel The distance is composed of [variable values], and it is also determined whether this distance is less than that of the unmanned surface vessel. Perception radius To construct the adjacency matrix At this point, the GAT model is equivalent to an unmanned surface vessel. Its own radar, therefore, other unmanned surface vessels are not limited to other unmanned surface vessels in the same formation, but also include unmanned surface vessels in other formations.
[0058] ; in, For other unmanned surface vessels The location. In addition, for unmanned surface vessels. Each probe ray is assigned an equivalent object as a probe ray node, with its position being... , Indicates the first A probe ray is used to calculate a new node feature matrix. and adjacency matrix The expression is as follows: ; Since the maximum length of the probe ray itself is In the absence of obstacles, the end of the probe ray is equivalent to a probe ray node, and the length of the probe ray is used as the data or information of the probe ray node. Therefore, ,in unmanned surface vessel The detection ray length is determined by multiplying the sensing radius by a coefficient close to 1, 0.99, to prevent issues with the adjacency matrix in the absence of obstacles. The detection ray node is also connected.
[0059] Then, the vectors are corrected to obtain the unmanned surface vessel. exist Actual local observations, actual node feature matrix, and actual adjacency matrix at time t: ; in, For unmanned surface vessels exist The mean length of all probe rays at any given time; the actual adjacency matrix. The calculation uniformly sets the sensing radius to Through the representation of the above vectors, the unmanned surface vessel (USV) can effectively utilize its own motion, obstacle avoidance status, and navigator status. At the same time, it can perceive nearby USVs and obstacles through the GAT model, providing sufficient and effective information for the decision-making model to achieve the USV's formation maintenance and collision avoidance tasks.
[0060] In some embodiments, an Actor network loss function is also designed. The first loss function for the independent Actor network of a drone is based on policy gradient theory, typically using an advantage function to guide policy updates. First Loss Function The format is:
[0061] in, This represents the advantage function, which measures the advantage of the current action relative to the average policy. It is a strategy ratio, used to measure new strategies. and old strategies Select Action The probability difference at that time. It is the entropy of the policy, which represents the randomness of the policy distribution. Adding it to the first loss function can encourage policy exploration. This is a hyperparameter used to balance the weights between the policy optimization loss and the entropy loss. By minimizing the first loss function, the drone's independent Actor network can learn the optimal action policy.
[0062] In practical applications, consider the time. , the expected function Expanding the calculations, for the drone intelligent agent... The first loss function of the independent Actor network can be expanded as follows: ; in, This represents the probability ratio between the old and new strategies. Used to restrict The update range should be narrowed to avoid excessive policy deviation. It is the entropy of the policy, representing the randomness of the policy distribution. It is an Actor network based on state The probability of the generated action. It is the entropy loss coefficient, which needs to be set with an initial value, and then decreases linearly as the number of training iterations increases. It is the maximum number of steps per agent within an exploration round. Represent the advantage function, and use the generalized advantage function (GAE). To indicate: ;in, These are GAE hyperparameters used to control the smoothness of the dominance function; It is a discount factor used to control the degree of influence of future expected rewards. yes The timing difference error at time t is defined as: ;in, yes Instant rewards for each moment; yes Value estimation at any given moment; yes Value estimation at any given moment.
[0063] For unmanned surface vessels (USVs) using RCASH networks, the calculation of the second loss function of the Actor network is first performed by splitting the process into individual losses for each USV within the formation. The average of these losses is then taken as the final loss, followed by one backpropagation. For example, the decision-making and calculation of the second loss function of the Actor network are performed by splitting the process into individual USV agents. The parts are handled separately: ; Calculate each unmanned surface vessel intelligent agent The losses are averaged to obtain the final loss, and then backpropagated to achieve implicit parameter sharing (explicit parameter sharing, where both input and output are features and variables of a single unmanned surface vessel agent; therefore, the saved shared parameters need to be read before calculating the output, and the new parameters are retained as shared parameters after network optimization). The second loss function... The expression is: ;in, This refers to the number of unmanned surface vessels (USVs) in the formation. Although the inputs and outputs of this implicit parameter sharing in the Actor network are global variables in form, the loss calculation and decision execution are still split and executed separately by each agent.
[0064] In some embodiments, a Critic network loss function is also designed. The Critic network loss function is as follows: [Further details on the Critic network loss function would be needed for the independent Critic network of an UAV and for an unmanned surface vessel using an RCASH network.] Both are expressed in the form of Mean Squared Error (MSE), used to minimize the difference between the predicted and actual values: ;in, Represents the true future return value; It is a value function, which can be understood as the value of a single drone or unmanned surface vessel based on its current state. The estimated future reward value. Minimizing the third loss function can improve the accuracy of the Critic network's evaluation of policy value, thereby enhancing the stability and efficiency of the entire algorithm.
[0065] Alternatively, the Critic network loss function can be expanded as follows: .
[0066] In practical applications, due to the generalized dominance function From timing difference error The instant reward calculated in the middle The correction, by Replace with This minimizes the mean squared error, resulting in a more accurate estimation of the value function. This is the value estimate calculated by the old Critic network that has not been updated. In this embodiment of the invention, mean squared error is used to measure the deviation between the predicted value and the target value. Therefore, the third loss function for the independent Critic network of the UAV can be further expressed as: ; For unmanned surface vessel agents using RCASH networks, the input to the Critic network is the global state. Local observation by each unmanned surface vessel intelligent agent Global observations from the merger And the perception compression vectors extracted from each unmanned surface vessel agent using GAT. Merged global-aware compressed vector Then, the value estimates of each unmanned surface vessel agent are obtained through the Critic network. Therefore, for unmanned surface vessel intelligent agents, Timing difference error at time point for: ; For unmanned surface vessel agents, the fourth loss function of the Critic network for: ; In step S600 of some embodiments, the model is trained and validated, and the effects of perception data compression, decision time, and task execution stability can be evaluated through formation tasks of different sizes. For example, training and validating the initial control model may include the following steps: Step 6.1: Network Initialization. Optionally, the parameters of the Actor and Critic networks are initialized using He initialization or Xavier initialization methods to ensure network stability in the early stages of training.
[0067] Step 6.2: Set relevant hyperparameters. Set the hyperparameters required for training the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models, including learning rate, batch size, and discount factor. (discount factor), entropy loss coefficient And gradient clipping threshold, etc. Optionally, for the GAT+IPPO model of UAVs, the initial learning rate is set to The learning rate is gradually reduced to a minimum using a cosine annealing decay method. The initial entropy loss coefficient is 0.5, and the maximum number of exploration rounds is... Set to 200, according to The method achieves a minimum value of 0.001 in the 100th round to improve stability in later training stages. For the GAT+MAPPO+RCASH model of the unmanned surface vessel, the initial learning rate is set to... The learning rate is gradually reduced to a minimum using a cosine annealing decay method. The initial entropy loss coefficient is 0.05, according to The method reaches a minimum value of 0.001 in the 100th round. Choosing a smaller value prevents excessive fluctuations in the shared network, which simultaneously performs decision-making and value estimation for multiple agents, during the early stages of training, thus enhancing training stability. The maximum number of steps an agent can take in one episode. Set to 100. During training, mini-batches are not used; instead, parameters are updated with the complete dataset. The Adam optimizer is chosen for its strong adaptability, which improves training efficiency.
[0068] Step 6.3: Train the UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models. First, run the current policy in the configured environment, collecting data on the states, actions, rewards, and next states of multiple agents. Then, based on the sampled data, calculate the advantage function. This guides policy updates. Finally, the parameters of the Actor network are updated using the policy gradient method, and the value function of the Critic network is optimized using the temporal difference method. Through iterative sampling and optimization of the above training process, the performance of the policy model is gradually improved until the reward value converges or reaches the preset number of training rounds.
[0069] Step 6.4: Verify Model Performance. Test the trained UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models in the same environment as the training environment. The tests mainly include: verifying the model's memory compression capabilities under different task scenarios, ensuring effective reduction of memory usage while preserving key neighborhood information; testing the time from perceived input to action generation in the Actor network to evaluate the model's adaptability in real-time tasks; statistically comparing the reward convergence curves and reward values of the UAV IPPO and UAV MAPPO models under different test environments; and evaluating the optimization effects of the model in terms of algorithm performance and generalization performance. For example, the testing process includes: (1) Memory usage compression capability verification: The GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models were run in small-scale and large-scale formations. The number of parameters of all decision models was recorded to evaluate the memory usage of the models. The test results are as follows: Figure 12 , Figure 13 As shown in Table 1.
[0070] Table 1
[0071] (2) Decision generation time test: The complete time from input to action generation of the Actor network of the GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models were tested respectively to evaluate the performance of the models in real-time response tasks. The test results are as follows: Figure 14 , Figure 15 And as shown in Table 2. Among them, Figure 14 Part (a) is a schematic diagram of the decision results after training the IPPO and MAPPO models of UAVs and unmanned surface vessels in a small formation, and Part (b) is a schematic diagram of the decision results after training the GAT+IPPO and GAT+MAPPO+RCASH models of UAVs and unmanned surface vessels in a small formation. Figure 15 Part (a) is a schematic diagram of the decision results after training the IPPO and MAPPO models of UAVs and unmanned surface vessels under large-scale formation, and Part (b) is a schematic diagram of the decision results after training the GAT+IPPO and GAT+MAPPO+RCASH models of UAVs and unmanned surface vessels under large-scale formation.
[0072] Table 2
[0073] (3) Comparison of training effects: The reward convergence curves of the GAT+IPPO and UAV GAT+MAPPO+RCASH models and the UAV IPPO and UAV MAPPO models were statistically analyzed to evaluate the training effect of the models.
[0074] (4) Generalization performance test: The random number seed for generating obstacle distribution in the training environment is 52. In this invention, 10 random number seeds from 2015 to 2024 are used to generate obstacles to obtain a new test environment. The rewards of GAT+IPPO and UAV GAT+MAPPO+RCASH models and UAV IPPO and UAV MAPPO models are compared to evaluate the generalization performance of the models. The test results are shown in Table 3.
[0075] Table 3
[0076] In step S700 of some embodiments, during the simulation, the agent continuously updates its actions according to the instructions of the control model. After each update, the current reward value is calculated, and the subsequent control strategy is adjusted based on the reward feedback, gradually moving towards maximizing the reward value. This drives the agent to adjust its position and speed to achieve formation and stability. By designing a target reward function to guide the agent's behavior, and using a target control model to enable the agent to autonomously learn and optimize the control strategy, combined with real-time environmental feedback and adjustments, precise and efficient control of the multi-agent formation is successfully achieved, enabling it to operate stably according to the expected formation.
[0077] like Figure 16 As shown, the training process for UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models includes: (1) Constructing the formation environment: Set the relevant parameters of the formation and obstacles, and set the detection rays with equal angle distribution for the unmanned surface vessel.
[0078] (2) Design reward function: Select the leader-follower model and design reward function based on the tasks of path planning, formation maintenance and obstacle avoidance.
[0079] (3) Construct the GAT model: aggregate neighbor node information to characterize the dynamic obstacle avoidance capability of the unmanned surface vessel.
[0080] (4) Constructing the RCASH model: Constructing the encoder, super adapter and adaptive decoder.
[0081] (5) Training decision algorithm: The IPPO algorithm is used for UAVs and the MAPPO algorithm is used for unmanned surface vessels in UAV formation to optimize the strategy and value estimation of each agent.
[0082] (6) Testing and validating the model: Test the trained UAV GAT+IPPO and UAV GAT+MAPPO+RCASH models, and compare the algorithm performance, algorithm effect and generalization performance.
[0083] This invention also provides a multi-agent formation control device based on graph attention networks and residual dynamic capability perception shared supernetworks, which can implement the aforementioned multi-agent formation control method based on graph attention networks and residual dynamic capability perception shared supernetworks. The device includes: a first module for constructing a multi-agent formation of UAVs and unmanned surface vessels; a second module for constructing a target reward function for the multi-agent formation; a third module for constructing a perception compression model using a graph attention network; a fourth module for constructing a decision model using a residual dynamic capability perception shared supernetwork; a fifth module for obtaining an initial control model based on a proximal policy optimization algorithm, the perception compression model, and the decision model; a sixth module for training the initial control model to obtain a target control model; and a seventh module for controlling the multi-agent formation based on the target reward function and the target control model.
[0084] This invention also provides an electronic device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the aforementioned multi-agent formation control method based on graph attention networks and residual dynamic capability perception sharing supernetworks. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0085] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned multi-agent formation control method based on graph attention networks and residual dynamic capability perception sharing supernetworks.
[0086] This invention also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned multi-agent formation control method based on graph attention networks and residual dynamic capability perception shared supernetworks.
[0087] In summary, the multi-agent formation control method and apparatus based on graph attention networks and residual dynamic capability perception sharing supernetworks of this invention have the following advantages: 1. This invention utilizes a Graph Attention Network (GAT) perception compression model to effectively aggregate and compress perception data. Simultaneously, it combines a shared supernetwork and reinforcement learning algorithms to optimize the multi-agent collaborative decision-making process, achieving efficient control of multi-UAV and unmanned surface vessel (USV) formations. This invention not only effectively addresses the complex changes in agent perception characteristics in dynamic environments but also improves system convergence performance and task stability by optimizing the perception and decision-making process, enabling efficient collaboration and reliable task execution of multi-agent systems in formation tasks of varying scales.
[0088] 2. The GAT model in this embodiment of the invention compresses and represents the dynamic obstacle avoidance capability of the unmanned surface vessel by introducing an attention machine. It effectively utilizes dynamic information while reducing the storage requirements and memory ratio of redundant data, enabling the system to maintain efficient operation in resource-constrained environments.
[0089] 3. In this embodiment of the invention, a residual dynamic capability sharing supernetwork (RCASH) is further introduced into the control of unmanned surface vessel formations. The dynamic obstacle avoidance capability representation of the GAT model is used as the input of dynamic capability, thereby reducing memory usage and improving sample utilization. At the same time, residual connection operations are used to increase gradients and enhance training effect.
[0090] 4. This embodiment of the invention combines the task of multi-UAV and unmanned surface vessel (USV) cooperative control with the leader-follower model, employing UAV GAT+IPPO and USV GAT+MAPPO+RCASH models. It maintains the high efficiency of IPPO for independent policy optimization and, through high-quality sensor input, outperforms traditional models in both policy stability and convergence speed. Compared to traditional UAV IPPO and USV MAPPO models, the UAV GAT+IPPO and USV GAT+MAPPO+RCASH models demonstrate significant advantages in training reward convergence values and task execution stability, exhibiting greater flexibility and robustness in complex dynamic environments.
[0091] 5. Through perception data compression, feature optimization, and parameter sharing technologies, the decision generation time of the model is significantly reduced, providing strong support for real-time response requirements. Furthermore, both the UAV GAT+IPPO and the UAV GAT+MAPPO+RCASH models maintain good performance in both small-scale and large-scale formation scenarios, demonstrating their excellent scalability and adaptability.
[0092] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.
Claims
1. A multi-agent formation control method based on graph attention networks and residual dynamic capability perception-shared supernetworks, characterized in that, Includes the following steps: Constructing multi-agent formations of drones and unmanned surface vessels; Construct the target reward function for the multi-agent formation; A perceptual compression model is constructed using a graph attention network; A decision-making model is constructed by perceiving and sharing a hypernetwork with residual dynamic capabilities. Based on the near-end policy optimization algorithm, an initial control model is obtained according to the perception compression model and the decision model; The initial control model is trained to obtain the target control model; The multi-agent formation is controlled based on the target reward function and the target control model.
2. The method according to claim 1, characterized in that, The construction of a multi-agent formation of drones and unmanned surface vessels includes the following steps: The first number, starting position, and target position of the multi-agent formation are preset; A second number of unmanned surface vessels is preset within each of the multi-agent formations; The control speed and formation distribution of the multi-agent formation are preset; A third number of obstacles is preset, and an obstacle distribution is generated based on a random number seed; The unmanned surface vessel is equipped with equally angularly distributed detection rays.
3. The method according to claim 1, characterized in that, The objective reward function for constructing the multi-agent formation includes the following steps: Based on the agent's first position at the current moment, the agent's second position at the previous moment, and the agent's target position at the current moment, obtain the proximity reward; Based on the first location and the target location, obtain a location reward; Obtain the first direction vector from the second position to the target position; Obtain the second direction vector from the second position to the first position; A directional reward is obtained based on the angle between the first direction vector and the second direction vector, the first position, and the second position. The danger radius of the collision of the intelligent agent is preset, and the first distance between the intelligent agent and the target position at the current moment is obtained; A reward is awarded based on the first distance and the danger radius. Based on the danger radius and the second distance between the agents, a collision avoidance reward is obtained; Obtain obstacle reward based on the third distance between the unmanned surface vessel and the obstacle; The target reward function is obtained based on the proximity reward, the position reward, the direction reward, the collision avoidance reward, the arrival reward, and the obstacle reward.
4. The method according to claim 1, characterized in that, The formula used to construct the perceptual compression model through the graph attention network includes: ; In the formula, Indicates the first The agent feature matrix of the layer; Indicates the activation function; Indicates the number of attention heads. ; Indicates the first Attention coefficient matrix of each attention head; Indicates the first The first attention head Layer feature matrix; Indicates the first The first attention head Layer linear transformation weight matrix; This indicates a splicing operation.
5. The method according to claim 1, characterized in that, The formulas used to construct the decision model through the residual dynamic capability-aware shared hypernetwork include: ; ; In the formula, Indicates the first Each agent at each time step The output characteristics; Indicates an adaptive decoder; This represents the network weight of the super adapter; This represents the encoded vector generated by the encoder; This represents the weight matrix in the network weights of the super adapter; This represents the bias matrix in the network weights of the super adapter; This represents the output of the capability-aware shared hypernetwork; This indicates a fully connected layer.
6. The method according to claim 1, characterized in that, The near-end policy optimization algorithm, based on the perception compression model and the decision model, obtains the initial control model, including the following steps: Based on the independent near-end policy optimization algorithm and combined with the perception compression model, the initial control model of the UAV is obtained; Based on the multi-agent near-end policy optimization algorithm, the perception compression model is combined with the decision model to obtain the initial control model of the unmanned surface vessel.
7. A multi-agent formation control device based on graph attention networks and residual dynamic capability perception sharing supernetworks, characterized in that, include: The first module is used to build a multi-agent formation of drones and unmanned surface vessels; The second module is used to construct the target reward function of the multi-agent formation; The third module is used to construct a perceptual compression model through a graph attention network; The fourth module is used to construct decision models by sensing the shared supernetwork through residual dynamic capabilities; The fifth module is used to obtain the initial control model based on the near-end policy optimization algorithm, the perception compression model, and the decision model. The sixth module is used to train the initial control model to obtain the target control model; The seventh module is used to control the multi-agent formation based on the target reward function and the target control model.
8. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.