Distributed formation control method driven by recursive balanced network under communication attack
By constructing an RBN controller and a heterogeneous communication network, and combining shrinkage mapping theory and multi-objective loss function, the robustness problem of multi-agent systems under DoS attacks is solved, and efficient and stable distributed formation control is achieved.
Patent Information
- Application Number
- CN202511543860.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-28
AI Technical Summary
Existing multi-agent systems struggle to maintain robust performance under heterogeneous communication links, denial-of-service attacks, and strict distributed information constraints. Traditional methods lack theoretical stability guarantees and quantitative characterization of complex attacks.
A distributed multi-agent system model is constructed, an RBN controller and a heterogeneous communication network are configured, stability is analyzed using contraction mapping theory, a multi-objective loss function is designed for training, and robustness is improved through a difficulty-first adversarial training strategy.
It achieves exponential convergence and robustness of the system under DoS attacks, improves computational efficiency by 50%, maintains 100% task completion and zero collisions under high-intensity attacks, and improves the real-time performance and engineering practicality of the system.
Smart Images

Figure CN121008593B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cooperative control technology for multi-agent systems, specifically to a distributed formation control method driven by a recursive balanced network under communication attacks. Background Technology
[0002] Cooperative control of multi-agent systems has broad application prospects in fields such as UAV formation, intelligent transportation, and distributed sensing. Although significant progress has been made in distributed formation control research, existing methods are generally based on the homogeneous communication assumption, making it difficult to cope with challenges such as heterogeneous communication links, denial-of-service (DoS) attacks, and strict distributed information constraints in real-world systems. Designing distributed control strategies that can adapt to asymmetric communication quality, resist malicious attacks, and maintain robust performance under conditions of limited local information has become a core bottleneck restricting its practical application.
[0003] In trajectory control, traditional methods typically require complete attitude information and rely on prior model knowledge. While recent research has improved control flexibility using neural network methods, it lacks theoretical stability guarantees. Particularly under strict distributed constraints, achieving stable trajectory tracking using only local information remains an open problem. Although contraction theory has shown potential in single-system control, extending it to multi-agent distributed scenarios faces numerous challenges.
[0004] Communication security is another key challenge facing multi-agent systems. DoS attacks severely threaten system performance by blocking communication channels. Existing research mainly focuses on single attack patterns, with insufficient theoretical analysis of combined attack scenarios. Although event triggering and adaptive strategies improve robustness to some extent, they struggle to cope with complex attacks such as selective interference and protocol interruption. More importantly, existing methods lack a quantitative characterization of the relationship between the degree of communication degradation and system stability.
[0005] Furthermore, the heterogeneity of agent communication capabilities is prevalent in practical deployments, yet research on this topic is insufficient. Most existing work assumes homogeneous communication networks, neglecting the natural differences in parameters such as bandwidth and latency. While preliminary explorations have been made in the cooperative control of heterogeneous multi-agent systems, a systematic approach to improving system fault tolerance through heterogeneous channel configuration remains lacking. The concept of effective capacity in information theory provides a theoretical foundation for heterogeneous network optimization, but its application in multi-agent formation control is still unexplored. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention proposes a distributed formation control method driven by a recursive balanced network under communication attacks, comprising the following steps:
[0007] S1. Construct a distributed multi-agent system model, configure an RBN controller for each agent, and construct a heterogeneous communication network architecture for the RBN controller and a DoS attack model.
[0008] S2. Optimize the RBN controller and implement a distributed RBN controller architecture. Analyze the stability of the RBN controller based on the shrinking map theory and prove the exponential convergence of the system state difference by constructing a Lyapunov function.
[0009] S3. The optimized RBN controller is trained using a difficulty-first adversarial training strategy. A multi-objective loss function is designed, and the training weights of each difficulty level are dynamically adjusted.
[0010] S4. Simulate the distributed multi-agent system optimized by the above steps, configure the simulation environment of the distributed multi-agent system under DoS attack, execute the formation trajectory tracking task, and test the robustness of the distributed multi-agent system under increasing DoS attack intensity.
[0011] Further, step S1 includes:
[0012] S1.1 Construct a distributed multi-agent system model consisting of n agents. The state vector of each agent i includes a position vector and a velocity vector. The communication relationship between agents is defined by an adjacency matrix. The neighbor set of each agent satisfies the preset maximum number of neighbors constraint.
[0013] S1.2 Design an RBN controller for each agent; the internal state evolution of the RBN controller is described by implicit difference equations, the nonlinear activation output is calculated by cascaded ReLU units, and the control output is generated by linear mapping;
[0014] S1.3 Construct a heterogeneous communication network, wherein the heterogeneous communication network includes at least two types of channels with different bandwidth, delay and packet loss rate parameters, and each type of channel is configured according to a preset ratio so that the average effective capacity of the network is equivalent to that of a homogeneous network.
[0015] S1.4 Establish a DoS attack model, which describes the impact of bandwidth flooding attacks, selective interference attacks, and protocol interruption attacks on communication channel performance through a time-varying attack strength function.
[0016] Furthermore, in step S2, a distributed RBN controller architecture is constructed, in which each agent is equipped with an independent RBN controller, receives local observation vectors and processes neighbor state information, and handles missing or delayed data caused by communication attacks through zero-padding.
[0017] The RBN controller includes a discrete-time state equation and a cascaded ReLU activation layer. The internal state is initialized using a zero vector, and cascaded ReLU activation, internal state update, and control output generation are performed at each time step.
[0018] The stability of the RBN controller is analyzed based on the shrinking mapping theory, and the exponential convergence of the system state difference is proved by constructing the Lyapunov function.
[0019] Set the global control gain parameter β as a shared parameter for all agents to adjust the control strength.
[0020] Furthermore, the stability of the RBN controller is analyzed based on the shrinking map theory. Let the state evolution matrix and the transition matrix be the state evolution matrix and the transition matrix, respectively, if the following conditions are satisfied:
[0021] Spectral radius condition: For constants Established;
[0022] Parameter constraints: ;
[0023] The input matrix is a nonlinear matrix. Connect the internal state to the nonlinear layer;
[0024] Then the RBN system with respect to metric It is exponentially contracting, and the rate of contraction does not exceed a constant. .
[0025] Further, step S3 includes:
[0026] S3.1 Divide the training scenario into multiple difficulty levels, including normal flight, light attack, medium attack, heavy attack and extreme attack; S3.2 Design a multi-objective loss function, which includes state tracking loss, collision avoidance loss, formation keeping loss, control regularization loss, boundary constraint loss and velocity constraint loss.
[0027] Furthermore, state tracking loss The distance from the agent to the target is measured using a weighted quadratic form of position and velocity, with position weights... Speed weight When the agent approaches the target at a distance of less than 2 meters, a proximity factor is introduced. ;
[0028] Collision avoidance of loss Based on the neighbor locations obtained through communication, for each pair of potentially colliding agents, when the distance is less than a safety threshold, the following calculations are performed: hour, To minimize the distance, apply a repulsive force. If the distance is less than Additional emergency avoidance items d represents the Euclidean distance between the two agents.
[0029] Formation maintains losses Based entirely on local communication information, each agent estimates the formation centroid based on the received neighbor positions, then calculates its own deviation from the ideal formation position. When communication is incomplete, the confidence factor is adjusted according to the number of available neighbors. ,in Let n be the set of neighbors that have successfully communicated, and n be the total number of agents.
[0030] Controlling regularization loss Limit control effort, weight Allowing for necessary large control inputs, boundary constraint loss Preventing the agent from leaving the designated airspace, velocity constraint loss The maximum speed is limited to 2 m / s.
[0031] Furthermore, in step S4:
[0032] S4.1 Configure a simulation environment for a distributed multi-agent system under a DoS attack;
[0033] S4.2, Perform formation trajectory tracking task;
[0034] S4.3 Test the robustness of the distributed multi-agent system under DoS attack;
[0035] S4.4. Conduct a comparative analysis of heterogeneous and homogeneous communication network configurations and a comprehensive performance comparison between the RBN controller and the traditional MPC method.
[0036] Compared with the prior art, the present invention has the following beneficial technical effects:
[0037] This invention proposes an innovative self-supervised trajectory control architecture, achieving multiple breakthroughs in theory, efficiency, and robustness in the field of distributed multi-agent formation control. Theoretically, it introduces contraction mapping theory into formation control under strict distributed constraints for the first time. Through parameterized design of recursive neural structures, it automatically satisfies the Lyapunov conditions, providing a mathematical guarantee for the system's exponential stability and filling the gap in existing methods lacking stability proofs when relying only on local states and limited neighbor information. Computationally, it employs explicit forward propagation instead of traditional implicit solving, significantly reducing computational complexity. The single-step inference time is only 23 milliseconds, improving efficiency by more than 50% compared to standard model predictive control. It also supports GPU parallel acceleration and can be scaled to large-scale real-time applications with more than 40 drones. In terms of robustness, it constructs an adversarial stability framework under combined communication attacks. Through a hard-first training strategy, it simultaneously resists five concurrent attacks, maintaining input-state stability even under extreme conditions where communication quality drops to 40%. In high-intensity attack tests, it achieves outstanding performance with 100% task completion and zero collisions. Furthermore, the heterogeneous network optimization mechanism based on effective capacity equivalence increases the survival probability of critical links by 40%, and the fully distributed architecture eliminates the risk of single point of failure. Through automated parameter tuning and simplified failure handling, the system achieves a trajectory tracking accuracy of 0.4 meters in embedded deployment, providing a complete solution with theoretical rigor, real-time performance and engineering practicality for resource-constrained combat environments. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a diagram of the distributed formation control system based on RBN of the present invention;
[0040] Figure 2 This is a diagram illustrating the distributed control trajectory of multiple UAVs swarms traversing a ring-shaped obstacle in 3D and XY / XZ planes under a communication attack, according to the present invention.
[0041] Figure 3 This is a graph showing the change of the first formation trajectory of the present invention with the degradation of communication quality;
[0042] Figure 4 This is a graph showing the change of the second formation trajectory of the present invention with the degradation of communication quality;
[0043] Figure 5 This is a graph showing the change of the third formation trajectory of the present invention with the degradation of communication quality;
[0044] Figure 6 This is a radar image comparison diagram of heterogeneous and homogeneous configurations for this invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0046] In the accompanying drawings of specific embodiments of the present invention, in order to better and more clearly describe the working principle of each component in the system and show the connection relationship of each part in the device, only the relative positional relationship between each component is clearly distinguished. It does not constitute a limitation on the signal transmission direction, connection sequence, or size, dimension, and shape of each part within the component or structure.
[0047] like Figure 1 The diagram shown illustrates the architecture of the RBN-based distributed formation control system of this invention. This invention proposes a distributed formation control method driven by a recursive balanced network under communication attacks. The control method includes the following steps:
[0048] S1. Construct a distributed multi-agent system model, configure an RBN controller for each agent, and build a heterogeneous communication network architecture and DoS attack model for the RBN controller.
[0049] S1.1 Construct a distributed multi-agent system model.
[0050] Consider by A multi-agent system consisting of quadcopter drones employs a fully distributed architecture. Each agent... state vector Includes position vector and velocity vector .
[0051] Communication relationships between agents are determined by adjacency matrices. Definition, where Represents intelligent agents In intelligent agents Within the communication range. Intelligent agent Neighbor set Satisfying cardinality constraint ,in This is the preset maximum number of neighbors.
[0052] S1.2, RBN controller architecture design.
[0053] A recursive equilibrium network controller is designed for each agent, whose internal state evolution is described by implicit difference equations; the nonlinear activation output of the controller is calculated by cascaded ReLU units, and the control output is generated by linear mapping.
[0054] S1.2.1 RBN Dynamics Design.
[0055] Each agent is equipped with an independent RBN controller, whose internal state The evolution is described by implicit difference equations:
[0056] ;
[0057] here, These are the state evolution matrix and the transition matrix, respectively. For nonlinear input matrices, For the observation input matrix, This is the internal state bias vector. Let be the dimension of the internal state of the RBN. Local observation vector. Includes its own state and the neighbor states received through communication. .
[0058] Nonlinear activation output Calculated using cascaded ReLU units Let be the number of units in the nonlinear activation layer, where the th is the number of units in the nonlinear activation layer. The pre-activation value of each unit is:
[0059] ;
[0060] m represents the m-th unit, which is activated by a scaling factor. Perform normalization. Matrix Connect the internal state to the nonlinear layer. To ensure the causality of forward computation using a strictly lower triangular matrix, Processing external input, This is a ReLU cell bias scalar.
[0061] Control output Generated via linear mapping:
[0062] ;
[0063] in , , To output the mapping matrix, To control the output bias vector.
[0064] S1.2.2 Parameterization and stability guarantee of RBN controller.
[0065] RBN parameters are optimized by the free matrix. and auxiliary matrix Obtain. Construct a symmetric positive definite matrix. ,in Let be the regularization parameter, and I be the parameter for network depth. Let... For the first Layered networks, By dimension After partitioning, the system matrix is extracted using the following relationship: , Lyapunov matrix ,as well as .
[0066] S1.3 Construct a heterogeneous communication network.
[0067] Channel Communication performance is determined by bandwidth. (bytes / second), transmission delay (seconds), packet loss rate and channel quality factor Characterization. Based on Wu-Negi theory, effective capacity is defined as... ,in For QoS parameters.
[0068] Heterogeneous networks include three channel types: Type A (high-speed channel) has bandwidth ,Delay and packet loss rate Type B (standard channel) maintains reference parameters , and Type C (remote channel) configuration is as follows: , and The configuration ratios are 25%, 50%, and 25% to ensure that the average effective capacity is equivalent to that of a homogeneous network.
[0069] S1.4 Establish a DoS attack model.
[0070] Attack strength function Describe the time-varying characteristics of the attack, during the preparation phase ( Linear growth to peak Active phase ( ) remain constant, recovery phase ( Linear decay.
[0071] Bandwidth flooding attacks degrade channel quality to The packet loss rate increased to Selective interference at high intensity ( The connection is completely disconnected when the intensity is low to medium, and the delay increases to [value missing]. Protocol interruption attacks amplify latency to The packet loss rate rose to .
[0072] S2. Optimize the RBN controller design and realize a distributed RBN controller architecture. Analyze the stability of the RBN controller based on the shrinking mapping theory and prove the exponential convergence of the system state difference by constructing a Lyapunov function.
[0073] S2.1 Construct a distributed RBN controller architecture.
[0074] The distributed controller designed in this step strictly adheres to the local information constraints defined in step S1. Each agent... Equipped with an independent RBN controller, the overall control architecture is as follows: Figure 1 As shown.
[0075] The red box shows the RBN controller for a single agent, which includes a discrete-time state equation and cascaded ReLU activation layers. The system achieves distributed control through local communication and dynamically selects training scenarios using a difficulty-first mechanism.
[0076] The RBN controller receives local observation vectors. , including its own state and neighbor states obtained through communication When communication is attacked, some neighbor states may be missing or delayed. The RBN controller handles missing data with zero-padding to ensure that the input dimension remains constant. .
[0077] RBN controller internal state The initialization uses a zero vector, that is... At each time step, the controller performs the following sequence of calculations: first, it calculates the nonlinear activation through cascaded ReLU layers. Then update the internal state. Finally, control output is generated. .
[0078] Global control gain parameters As a shared parameter for all agents, it is used to adjust the control strength: This parameter is adaptively adjusted during training, with an initial value set to 1.0.
[0079] S2.2 Analysis of the stability of the RBN controller based on the shrinking mapping theory.
[0080] The stability of RBN is guaranteed by the shrinking mapping theory, and distributed control is achieved through three layers of explicit computation, as shown in Algorithm 1. RBN uses explicit forward propagation, reducing computational complexity from... Down to ,in .
[0081] Theorem 1 (Shrinkability of Recursive Balanced Networks) Consider an RBN system, let... reversible, If the following conditions are met:
[0082] Spectral radius condition: For a certain Established;
[0083] Parameter constraints: ;
[0084] Then the RBN system with respect to metric It is exponentially contracting, with a contraction rate not exceeding [a certain percentage]. .
[0085] Proof. To simplify the notation, agent subscripts are omitted in this proof. Analyze the properties of a single RBN controller. , , Consider the implicit dynamic equations of the RBN system:
[0086] (1);
[0087] in , , This is the element-wise ReLU activation function.
[0088] Suppose that the RBN system has two solution trajectories. and , corresponding to the same external input sequence Define the state increment. .
[0089] Expand equation (A.1) for each of the two trajectories:
[0090] (2a);
[0091] (2b);
[0092] Subtracting the two equations yields the incremental dynamics:
[0093] (3);
[0094] in The incremental dynamics equations describe the evolution of deviations between the trajectories of different UAVs. In formation flying, this means that even if a UAV deviates from its trajectory due to communication interference, the deviation will not be amplified and propagated within the formation, ensuring the overall stability of the formation.
[0095] For the ReLU activation function, it satisfies the sector bounded condition. Specifically, for any ,have:
[0096] (4);
[0097] Applying this property to vector form, the sector-bounded characteristic of the ReLU activation function corresponds in engineering to the saturation protection mechanism of a controller. When the UAV suffers strong interference, the control output will not oscillate excessively or diverge, but will remain within a reasonable range. This mathematical constraint translates into safety assurance in actual flight: even under extreme attack scenarios, the UAV will not engage in dangerous maneuvers such as uncontrolled dives or rapid climbs, ensuring the safety of personnel and equipment. Definition The dimension of the nonlinear activation layer is a diagonal matrix. Its diagonal elements are:
[0098] (5);
[0099] From (A.4) we know .and then:
[0100] (6);
[0101] Substitute (A.6) into (A.3):
[0102] (7);
[0103] Left multiplication :
[0104] (8);
[0105] The shrinking property of the state transition matrix ensures that the state difference between any two drones in the formation decays exponentially over time. This ensures that even if some drones lose contact, the formation can still maintain stable flight, avoiding disintegration or collision. Define the Lyapunov function. :
[0106] (9);
[0107] in .matrix The positive definiteness is due to The invertibility is directly guaranteed by the Lyapunov function. This can be understood as an "energy" measure of the state deviation within the formation. The monotonically decreasing property of this function ensures that the drone formation can autonomously converge to the desired trajectory after being subjected to a communication attack, without human intervention.
[0108] To analyze the evolution of the Lyapunov function, substituting (A.8) yields:
[0109] (10a);
[0110] Will Substitute:
[0111] (10b);
[0112] Note ,therefore:
[0113] (10c);
[0114] Define the time-varying transition matrix Due to the Jacobian matrix satisfy We have:
[0115] (11);
[0116] To establish contractility, it is necessary to prove its existence. Make .
[0117] First, using matrix perturbation theory, for The spectral radii are:
[0118] (12);
[0119] because (because Given a diagonal matrix with diagonal elements in [0,1], we get:
[0120] (13);
[0121] Based on parameter constraints Selecting contraction factors satisfy:
[0122] (14);
[0123] Since the spectral radius of a symmetric matrix is equal to its spectral norm, and It is a symmetric matrix, and we have:
[0124] (15);
[0125] This means , equivalent to:
[0126] (16);
[0127] Multiply both sides by left Right multiplication :
[0128] (17);
[0129] Matrix inequalities The engineering implication is the "memory decay" characteristic of formation deviation. This ensures that the impact of historical disturbances on the current formation state decays exponentially, giving the system the ability to "forget" the effects of past attacks. In continuous combat environments, this characteristic ensures that the formation will not be affected by early attacks, thus improving the system's operational sustainability and mission reliability. Substituting (A.17) into (A.10c):
[0130] (18);
[0131] Recursively apply the above inequality:
[0132] (19);
[0133] Finally, using the property of Rayleigh quotient, for any :
[0134] (20);
[0135] Combining (A.19) and (A.20):
[0136] (twenty one);
[0137] Define condition number ,have to:
[0138] (twenty two);
[0139] because The system exhibits exponential contraction, with a contraction rate not exceeding [a certain percentage]. Shrinkage < The mathematical guarantee of <1 translates into concrete engineering performance: when communication quality degrades, formation deviation converges exponentially. For example, At a value of 0.8, the formation resynchronizes within 3-5 control cycles (0.15-0.25 seconds), meeting the response requirements of real-time formation control. The shrinking property is particularly important in distributed settings because it ensures that each local controller remains stable even when partial communication fails. Specifically, when neighbor information is missing, the controller degenerates into feedback control based on its own state; the shrinking property ensures the stability of this degenerate mode.
[0140] S3. The optimized RBN controller is trained using a difficulty-first adversarial training strategy. A multi-objective loss function is designed, and the training weights for each difficulty level are dynamically adjusted.
[0141] S3.1, Difficulty-first adversarial training strategy.
[0142] To improve the robustness of the controller under combined attacks, this invention proposes a Difficulty-First adversarial training strategy. This strategy divides the training scenarios into five difficulty levels: normal flight (no attack), light attack (1 attack source), medium attack (2 attack sources), heavy attack (3 attack sources), and extreme attack (4-5 attack sources).
[0143] The training process is divided into two phases. The basic training phase accounts for [percentage] of the total training rounds. Basic control capabilities are established using only normal flight scenarios. The adversarial training phase employs a periodic strategy, with each cycle consisting of 10 training rounds to ensure coverage of all difficulty levels.
[0144] The scene selection probability is dynamically adjusted based on the training progress. Let the training progress be... ,in This is the current round number. Based on the number of training rounds, This represents the total number of rounds. Early stage ( The focus is on normal and mild scenarios, with a focus on the mid-term ( Increase the proportion of heavy scenes, in the later stages ( The focus is on training in extreme scenarios.
[0145] The performance tracing mechanism records the most recent loss value for each difficulty level. ,in Indicates the difficulty level. Updated using an exponential moving average: The smoothing coefficient of the exponential moving average , is the most recent exponential moving average loss value for the d-th difficulty level. The worst-performing scenario type receives a higher training weight in the next cycle.
[0146] S3.2 Design of multi-objective loss function.
[0147] The loss function takes into account multiple control objectives, and the weights of each item are dynamically adjusted according to task priority and training phase.
[0148] State tracking loss The distance from the agent to the target is measured using a weighted quadratic form based on position and velocity. Position weights. Speed weight When the agent approaches the target (within 2 meters), a proximity factor is introduced. Enhanced accuracy requirements, among which Distance to the target.
[0149] Collision avoidance of loss Calculation is based on neighbor locations obtained through communication. For each pair of potentially colliding agents, when the distance is less than a safety threshold... hour, To minimize the distance, apply a repulsive force. If the distance is less than Additional emergency avoidance items .
[0150] Formation maintains losses It is entirely based on local communication information. Each agent estimates the formation centroid based on the received neighbor positions and then calculates its own deviation from the ideal formation position. When communication is incomplete, the confidence factor is adjusted according to the number of available neighbors. ,in The set of neighbors that have successfully communicated.
[0151] Controlling regularization loss Limit control effort, weight Allow for necessary large control inputs. Boundary constraint loss. Preventing the agent from leaving the designated airspace, velocity constraint loss The maximum speed is limited to 2 m / s.
[0152] The time-varying weight strategy enhances arrival accuracy in the later stages of the trajectory. Let the time weight be:
[0153] ,in This represents the trajectory length. The formation weight is reduced to 1% of its original value in the last 10 steps, allowing the agent to reach the target position first.
[0154] In practice, the RBN parameters are trained using the Adam optimizer, with the initial learning rate set to... The value is halved during the middle of training. To prevent numerical instability, matrix inversion is performed. Cholesky decomposition was used for calculation, and the regularization parameter was calculated. Ensure the condition numbers are good.
[0155] Gradient calculation uses automatic differentiation, avoiding the implicit function theorem required by REN. This allows RBN to be seamlessly integrated into modern deep learning frameworks, supporting GPU acceleration and batch training. Internal state Reset at the start of each training episode to ensure training independence.
[0156] Communication failure handling employs a zero-fill strategy, meaning that when a neighbor... When the status is not received, set This simple strategy performs well in practice because RBNs learn to recognize and adapt to zero-input patterns through training.
[0157] S4. Simulate the distributed multi-agent system optimized by the above steps, configure the simulation environment of the distributed multi-agent system under DoS attack, execute the formation trajectory tracking task, test the robustness of the distributed multi-agent system under increasing DoS attack intensity, conduct comparative analysis of heterogeneous and homogeneous communication network configurations, and compare the comprehensive performance of RBN controller and traditional MPC method.
[0158] S4.1 Configure a simulation environment for a distributed multi-agent system under a DoS attack.
[0159] This step describes the simulation environment configuration for the distributed multi-agent system under a DoS attack. The simulation aims to verify the ability of the proposed distributed RBN controller to maintain formation flight under conditions of communication disruption.
[0160] 1) Multi-agent system configuration
[0161] The simulation environment includes multiple quadcopter UAVs performing a formation flying mission. The system adopts a discrete-time dynamics model, and the main parameters are shown in Table 1.
[0162] Table 1 System Parameter Configuration
[0163]
[0164] The system supports three standard formation configurations. The geometric parameters and task settings for each formation are shown in Table 2.
[0165] Table 2 Formation Configuration and Mission Parameters
[0166]
[0167] Note: All formations will undertake the crossing mission, from... m flies towards m, passing through A circular obstacle in a plane (radius of the safety passage) m). Formation in - The plane retains geometric invariance, that is... ,in For the formation center of mass trajectory, Let be the fixed offset vector of agent i relative to the centroid. Let i be the desired position of agent i.
[0168] 2) Communication network architecture
[0169] The communication network system employs a predefined static topology, eliminating dependence on global location information. Under the standard homogeneous communication network configuration, all channels have the same parameters. To improve realism, a heterogeneous communication network configuration based on the Wu & Negi equivalent capacity theory was also implemented, ensuring that the overall network performance is equivalent to the homogeneous benchmark. Table 3 shows the communication channel parameters.
[0170] Table 3 Communication Channel Parameters
[0171]
[0172] equivalent capacity ,in For queue weight factors, For channel bandwidth, For packet loss rate, This is a delay.
[0173] The heterogeneous communication network configuration simulates the heterogeneity of real-world networks by adjusting the bandwidth, latency, and packet loss rate of various channel types, while maintaining an average equivalent capacity deviation of less than 5%. Type A channels represent high-speed, short-distance links, Type B channels represent standard links, and Type C channels simulate long-distance or low-quality links.
[0174] 3) DoS attack modeling
[0175] It implemented four types of DoS attacks, covering common attack patterns encountered in real-world networks. Attack strength. The specific parameters and impact mechanisms of various types of attacks determine the extent of their destructive power, as shown in Table 4.
[0176] Table 4. Parameters of the DoS Attack Model
[0177]
[0178] The impact of DoS attacks on network performance is quantified by the following formula:
[0179] Channel quality degradation: ;
[0180] Increased latency: ,in ;
[0181] The packet loss rate is increasing: ,in ;
[0182] q is a channel quality parameter. The rate of change of channel quality, The intensity of the DoS attack is τ, and the current communication latency is τ. Let k be the latency growth rate, k be the latency growth coefficient, and p be the packet loss rate. This represents the rate of change in packet loss rate.
[0183] The training phase employs a combination of eight preset DoS attack scenarios, gradually improving the system's robustness through a difficulty-first strategy. DoS attack opportunities are distributed between steps 10 and 130 to ensure coverage of all stages of the task.
[0184] S4.2 Execute formation trajectory tracking task.
[0185] This step verifies the universality of the distributed RBN controller architecture. Figure 2 The study showcased the obstacle-crossing trajectories of formations of 3 (triangular), 4 (square), and 5 (pentagonal) drones. After training with various attack scenarios, all formations achieved 100% mission success rate and zero collisions during the testing phase. Despite varying numbers of communication links (6-10), each formation completed obstacle crossing relying solely on local neighbor information, demonstrating adaptive deformation and rapid recovery capabilities. In particular, the stable control of the 5-drone formation in a 30-dimensional state space validated the effectiveness of the RBN shrinkage mapping properties, providing support for large-scale applications.
[0186] S4.3 Test the robustness of the distributed multi-agent system under DoS attacks.
[0187] This step evaluates the adversarial robustness and generalization ability of the RBN controller. Figures 3-5 The study demonstrates the trajectory performance of three formations under progressive DoS attacks: from no attack to extreme attacks (communication quality 37-39%). Crucially, the test attack patterns are completely different from the training phase—using new attack combinations, temporal sequences, and intensity distributions—validating the model's generalization ability.
[0188] The experiment revealed three key findings: (1) Even with a communication loss of more than 60%, all formations still completed their tasks (error < 0.4m), verifying the stability guarantee of ISS under 40% communication quality; (2) Trajectory perturbation and attack intensity have a nonlinear relationship, with mild attacks having almost no impact, while extreme attacks result in oscillations but still convergence; (3) The robustness provided by hard-first training enables the distributed multi-agent system to cope with unseen attack patterns, demonstrating strong zero-shot generalization ability. This provides a confidence guarantee for practical deployment.
[0189] S4.4. Conduct a comparative analysis of heterogeneous and homogeneous communication network configurations and a comprehensive performance comparison between the RBN controller and the traditional MPC method.
[0190] Figure 6 The radar charts compared the loss values of the three formations under increasing attack intensity (radial coordinates, inner circle is better). The heterogeneous communication network configuration (dashed line) outperformed the homogeneous (solid line) configuration in all scenarios, with an average loss reduction of 4-8%, among which the three-aircraft formation showed the most significant improvement (7.9%). Geometrically, the homogeneous communication network configuration exhibited an irregular contour and significant performance fluctuations; the heterogeneous communication network configuration maintained a smooth pentagon, with more balanced degradation. This advantage stems from the differentiated channel parameter design—the combination of Type A high-speed, Type B standard, and Type C long-range links effectively avoids cascading failures. Under the constraint of maintaining the equivalent average effective capacity, the heterogeneous communication network mechanism provides a practical robustness enhancement scheme for multi-agent communication in adversarial environments. Table 5 shows the comprehensive performance comparison of the RBN and MPC methods.
[0191] Table 5. Comparison of overall performance between RBN and MPC methods
[0192]
[0193] Table 5 compares the performance of RBN with four MPC variants. RBN has a significant advantage in computational efficiency, requiring only 23ms for a single inference step, which is 50% faster than standard MPC and 66% faster than robust MPC. This is due to its O(lq) time complexity. 2 The computational complexity of ) is far superior to that of MPC (O(N)). 3 p 2 In terms of scalability, RBN supports formations of more than 40 drones, while standard MPC only supports 8. RBN provides stability guarantees through shrinking maps, employs a fully distributed architecture, and supports GPU acceleration. Although parameter tuning is moderately difficult, its low hardware requirements make it more suitable for practical deployment. Overall, RBN significantly outperforms traditional MPC methods in terms of computational efficiency and system scalability while maintaining high task completion rates.
[0194] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0195] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A distributed formation control method driven by a recursive balanced network under communication attacks, characterized in that, Includes the following steps: S1. Construct a distributed multi-agent system model, configure an RBN controller for each agent, and construct a heterogeneous communication network architecture for the RBN controller and a DoS attack model. S2. Optimize the RBN controller and implement a distributed RBN controller architecture. Analyze the stability of the RBN controller based on the shrinking map theory and prove the exponential convergence of the system state difference by constructing a Lyapunov function. S3. The optimized RBN controller is trained using a difficulty-first adversarial training strategy. A multi-objective loss function is designed, and the training weights of each difficulty level are dynamically adjusted. S4. Simulate the distributed multi-agent system optimized by the above steps, configure the simulation environment of the distributed multi-agent system under DoS attack, execute the formation trajectory tracking task, and test the robustness of the distributed multi-agent system under increasing DoS attack intensity.
2. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 1, characterized in that, Step S1 includes: S1.1 Construct a distributed multi-agent system model consisting of n agents. The state vector of each agent i includes a position vector and a velocity vector. The communication relationship between agents is defined by an adjacency matrix. The neighbor set of each agent satisfies the preset maximum number of neighbors constraint. S1.2 Design an RBN controller for each agent; the internal state evolution of the RBN controller is described by implicit difference equations, the nonlinear activation output is calculated by cascaded ReLU units, and the control output is generated by linear mapping; S1.3 Construct a heterogeneous communication network, wherein the heterogeneous communication network includes at least two types of channels with different bandwidth, delay and packet loss rate parameters, and each type of channel is configured according to a preset ratio so that the average effective capacity of the network is equivalent to that of a homogeneous network. S1.4 Establish a DoS attack model, which describes the impact of bandwidth flooding attacks, selective interference attacks, and protocol interruption attacks on communication channel performance through a time-varying attack strength function.
3. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 1, characterized in that, In step S2, a distributed RBN controller architecture is constructed, in which each agent is equipped with an independent RBN controller, receives local observation vectors and processes neighbor state information, and handles missing or delayed data caused by communication attacks through zero-padding. The RBN controller includes a discrete-time state equation and a cascaded ReLU activation layer. The internal state is initialized using a zero vector, and cascaded ReLU activation, internal state update, and control output generation are performed at each time step. The stability of the RBN controller is analyzed based on the shrinking mapping theory, and the exponential convergence of the system state difference is proved by constructing the Lyapunov function. Set a global control gain parameter as a shared parameter for all agents to adjust the control strength.
4. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 3, characterized in that, The stability of the RBN controller is analyzed based on the shrinking mapping theory, provided that the following conditions are met: Spectral radius condition: For constants Established; Parameter constraints: ; The input matrix is a nonlinear matrix. Connect the internal state to the nonlinear layer. These are the state evolution matrix and the transition matrix, respectively; Then the RBN system with respect to metric It is exponentially contracting, and the rate of contraction does not exceed a constant. .
5. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 4, characterized in that, Step S3 includes: S3.1 Divide the training scenario into multiple difficulty levels, including normal flight, light attack, medium attack, heavy attack and extreme attack; S3.2 Design a multi-objective loss function, which includes state tracking loss, collision avoidance loss, formation keeping loss, control regularization loss, boundary constraint loss and velocity constraint loss.
6. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 5, characterized in that, State tracking loss The distance from the agent to the target is measured using a weighted quadratic form of position and velocity, with position weights... Speed weight When the agent approaches the target at a distance of less than 2 meters, a proximity factor is introduced. ; Collision avoidance of loss Based on the neighbor locations obtained through communication, for each pair of potentially colliding agents, when the distance is less than a safety threshold, the following calculations are performed: hour, To minimize the distance, apply a repulsive force. If the distance is less than Additional emergency avoidance items d represents the Euclidean distance between the two agents. Formation maintains losses Based entirely on local communication information, each agent estimates the formation centroid based on the received neighbor positions, then calculates its own deviation from the ideal formation position. When communication is incomplete, the confidence factor is adjusted according to the number of available neighbors. ,in Let n be the set of neighbors that have successfully communicated, and n be the total number of agents. Controlling regularization loss Limit control effort, weight Allowing for necessary large control inputs, boundary constraint loss Preventing the agent from leaving the designated airspace, velocity constraint loss The maximum speed is limited to 2 m / s.
7. The distributed formation control method driven by a recursive balanced network under communication attacks as described in claim 1, characterized in that, In step S4: S4.1 Configure a simulation environment for a distributed multi-agent system under a DoS attack; S4.2, Perform formation trajectory tracking task; S4.3 Test the robustness of the distributed multi-agent system under DoS attack; S4.
4. Conduct a comparative analysis of heterogeneous and homogeneous communication network configurations and a comprehensive performance comparison between the RBN controller and the traditional MPC method.
Citation Information
Patent Citations
Modulation strategy design method and system of DC-DC converter based on reinforcement learning
CN117318480A
Robot system, robot system control method, program, recording medium, and article manufacturing method
JP2016221659A