Multi-agent collaborative pursuit method based on dynamic compromise coefficient
By introducing an escapee velocity factor to optimize the trade-off coefficient and dynamically adjusting the pursuit strategy, the problem of low capture efficiency and insufficient coordination efficiency of multi-agent pursuit algorithms in complex scenarios is solved, achieving a more efficient and safer pursuit effect.
Patent Information
- Application Number
- CN202511647121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-24
AI Technical Summary
Existing multi-agent pursuit algorithms lack dynamic balancing mechanisms when facing high-speed moving targets, making them unable to adapt to dynamic environmental changes. This results in low capture efficiency, insufficient coordination efficiency, and the initial formation has a significant impact on algorithm performance, leading to path redundancy and collision risks.
By introducing a speed-sensitive factor for escapees and optimizing the trade-off coefficient design, a dynamically adjusted pursuit strategy is generated. By calculating the encirclement speed and pursuit speed of each pursuer, a coordinated pursuit is achieved.
It significantly improves the adaptability of the pursuit system to dynamic environments, shortens capture time, reduces total movement distance and number of collisions, and enhances the cooperation efficiency among pursuers.
Smart Images

Figure CN121559856A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of fully automated products and relates to a multi-agent cooperative pursuit method based on dynamic trade-off coefficients. Background Technology
[0002] In the study of multi-agent cooperative pursuit problems, although traditional methods have laid an important foundation for this field, their inherent limitations have gradually become apparent in real-world complex scenarios, making it difficult to meet the requirements for efficient and robust pursuit.
[0003] First, differential game theory (Isaacs), as the analytical framework for the classic pursuit-escape problem, is based on the core idea of describing the dynamic game process between the pursuer and the escapee by constructing the Hmilton-Jacobi-Isaacs (HJI) equations and solving for the optimal strategy. However, this theory was initially designed for a scenario where a single pursuer is fighting against a single escapee. In this simplified scenario, the pursuer can directly obtain the optimal control law by solving the HJI equations, and the computational complexity is still acceptable. But when the problem is extended to multiple pursuers cooperating to pursue a single escapee or multiple escapees, the state space dimension of the system grows exponentially with the number of pursuers (i.e., the so-called "curse of exponential growth").
[0004] Secondly, the Apollo Circle model (Sun et al.) is another classic approach to pursuit strategy design. Its basic principle is to model the pursuit process as the pursuer adjusting its position to keep the escapee always on the boundary of a dynamically changing "Apollo Circle"—the radius of which is related to the pursuer's speed, and the center of the circle is adjusted in real time according to the escapee's current position. Theoretically, when the pursuer can keep the escapee on the circumference, they can achieve encirclement or capture through continuous approach. However, the effectiveness of this model is highly dependent on fixed geometric constraints: for example, the radius of the circle is usually preset to a fixed value proportional to the pursuer's speed, and the trajectory of the circle's center also follows predefined geometric rules (such as always pointing towards the tangent direction between the escapee and the circle). This fixed nature results in a lack of flexibility in the model when faced with dynamic changes in the escapee's speed (e.g., sudden acceleration, deceleration, or change of direction): if the escapee's speed suddenly increases, the pre-set circle radius may be insufficient to cover its escape range, preventing the pursuers from adjusting their encirclement strategy in time; conversely, if the speed is too low, an excessively large circle radius will cause the pursuers to disperse excessively, reducing encirclement efficiency. Furthermore, the model does not consider the cooperative geometric relationships between multiple pursuers (e.g., the superposition and conflict of multiple Apollo circles), making it difficult to achieve efficient cooperation among multiple agents.
[0005] Finally, the original tradeoff coefficient design serves as a decision-making basis in multi-hunter cooperative strategies. Its core is to allocate priority or control among hunters by quantifying the relative positional relationship between hunters and escapees (such as coverage angle differences and distances). However, this design suffers from significant information inadequacy: the algorithm only focuses on the hunter's own state (position, orientation) and the static positional relationship with the escapee, neglecting the escapee's dynamic characteristics—especially the escapee's speed, direction, and magnitude. Ignoring this crucial information causes the tradeoff coefficient design to deviate from the actual optimal strategy, thereby reducing pursuit efficiency.
[0006] Due to the limitations of the traditional methods mentioned above, existing multi-hunter cooperative pursuit algorithms still face multiple defects in practical applications, which seriously restricts their practicality in complex scenarios (such as high-speed escape and dynamic environments).
[0007] First, encirclements are easily broken when escapees are moving at high speeds. When escapees are moving at high speeds, traditional encirclement strategies often fail to respond in time.
[0008] Secondly, the initial formation layout significantly impacts algorithm performance, but lacks systematic optimization. The initial positional distribution of multiple pursuers directly affects the efficiency of subsequent coordinated pursuit: a reasonable initial formation can shorten the pursuit path and reduce the probability of collaborative conflicts; while an unreasonable initial formation may require pursuers to move additionally to adjust their positions, or even prevent the formation of an effective encirclement due to an unbalanced initial distribution. However, most existing methods assume that the initial formation affects the final capture time and pursuit path length, and do not design adaptive initial adjustment strategies. This "black box" initial setting makes the algorithm performance overly dependent on initial conditions, making it difficult to guarantee stable pursuit results in practical applications.
[0009] Third, the high number of collisions and low coordination efficiency are problems. Multiple pursuers need to frequently adjust their positions to avoid collisions during coordinated pursuit, while maintaining a reasonable cooperative distance to maximize encirclement pressure. However, traditional methods typically focus only on the interaction between a single pursuer and the escapee when designing pursuit strategies, neglecting the dynamic coordination among multiple pursuers. This leads to unnecessary close-range contact or even collisions when pursuers adjust their positions due to information asymmetry or strategy conflicts. This not only increases ineffective path movement but may also further prolong capture time due to readjustment after collisions. Summary of the Invention
[0010] To address the shortcomings of existing multi-agent pursuit algorithms, such as the lack of a dynamic balancing mechanism in pursuit and encirclement strategies, resulting in low capture efficiency for high-speed moving targets, the inability of traditional trade-off coefficients to consider escapee speed factors and adapt to dynamic environmental changes, insufficient coordination efficiency among pursuers leading to path redundancy and collision risks, and poor algorithm adaptability under different initial formations, making it difficult to handle complex escape strategies, this invention adopts the following technical solution: a multi-agent cooperative pursuit method based on dynamic trade-off coefficients, comprising the following steps: S1: Obtain the initial position and maximum speed of the pursuer and the escapee, and the capture radius of the pursuer; S2: Determine if the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit and execute S3 when t≥t max The pursuit ended; S3: Obtain the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; S4: Based on the real-time location information of the escapee and the pursuer, a speed-sensitive factor of the escapee is introduced to obtain an improved trade-off coefficient and generate a dynamically adjusted pursuit strategy. S4: Based on the improved compromise coefficient, calculate the encirclement speed and pursuit speed of each pursuer, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, thereby realizing the coordinated pursuit of the escapee by the pursuers.
[0011] Furthermore: the improved compromise coefficient The expression is as follows:
[0012] in: For the escapee's speed-sensitive factor, As the bounding factor, For pursuit factors;
[0013] Where: k is the adjustment coefficient. Let V be the angle between the direction of the escapee's velocity and the direction of the pursuer's velocity, and Ve be the escapee's velocity. i Then it represents the speed of the pursuer i.
[0014] Furthermore, the dynamically adjusted pursuit strategy is as follows: >0: Increase the compromise coefficient after improvement Strengthen the surrounding speed component The pursuit strategy has been adjusted as follows: Increase the density of the encirclement and reduce the distance between adjacent pursuers; Adjust the heading angle to fill the encirclement gap; Reduce the radial velocity of the pursuer to prevent excessive compression from causing the escapee to break through; <0: Reduce the improved compromise factor Prioritize shortening the distance and adjust the pursuit strategy as follows: Increase the radial velocity component of the pursuer ; Maintain minimum enclosure density to prevent escape by detour.
[0015] Furthermore, this method is applicable when N pursuers surround and capture one escapee, where N ≥ 2.
[0016] Furthermore, the dynamically adjusted pursuit strategy also includes arranging the pursuit methods of the pursuers, including: Gate-shaped formation: Two pursuers are arranged in a "gate" shape, simulating an interception and encirclement strategy; the "gate" shape is symmetrically distributed on the left and right sides, leaving an escape passage in the middle; Triangular formation: Three pursuers form an acute or obtuse triangle to simulate a "focused" encirclement; the vertices of the acute or obtuse triangle are located on the escapee's likely path. Equilateral triangle formation: Three pursuers form a tight equilateral triangle, simulating a "high-pressure close-in" strategy; the center of the equilateral triangle is close to the escapee's initial position. Two pursuers pressure the escapee's movement, while one pursuer is in a predetermined trajectory: Two pursuers pressure the escapee's movement on one side, while the other pursuer is in the trajectory in the direction the escapee is being pressured, simulating a "path prediction" strategy. Trapezoidal formation: The four pursuers are arranged in a trapezoidal shape, simulating the "asymptotic compression" strategy.
[0017] A multi-agent cooperative pursuit device based on dynamic trade-off coefficients, comprising: Acquisition Module I: Used to acquire the initial position and maximum speed of both the pursuer and the escapee, as well as the capture radius of the pursuer; Judgment module: Used to determine whether the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit when t≥t max The pursuit ended; Acquisition Module II: Acquire the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; The generation module is used to generate a dynamic adjustment strategy based on the real-time location information of the escapee and the pursuer, by introducing the escapee's speed sensitivity factor and obtaining an improved trade-off coefficient. The pursuit module is used to calculate the encirclement speed and pursuit speed of each pursuer based on the improved compromise coefficient, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, so as to realize the coordinated pursuit of the pursuers and the escapee.
[0018] A computer device includes: a processor and a memory, the memory storing a program module, characterized in that the program module runs on the processor to implement any of the methods described above.
[0019] This invention provides a multi-agent cooperative pursuit method based on dynamic trade-off coefficients. It proposes a novel trade-off coefficient design method, significantly enhancing the pursuit algorithm's adaptability to dynamic environments by introducing an escapee velocity factor. By designing different initial escapee formations (gate, triangle, trapezoid, etc. for straight-line escape; equilateral triangle (near / far), pressure movement, trapezoidal encirclement, etc. for repulsive escape), the effectiveness of the improved algorithm under diverse initial layouts is verified. This invention provides a new solution to the multi-robot pursuit problem and offers theoretical guidance and practical reference for multi-robot cooperative pursuit tasks in real-world applications.
[0020] This study addresses the key issue of cooperative pursuit of fast-moving escapees in multi-agent systems, proposing an improved pursuit algorithm based on a dynamic tradeoff coefficient. By introducing an escapee velocity factor to optimize the tradeoff coefficient, the shortcomings of the original algorithm in the tracking and encirclement balance mechanism are effectively addressed, significantly improving the pursuit system's adaptability to dynamic environments.
[0021] The method described in this application achieves performance improvements compared to the original algorithm: capture time is reduced by 30% in flawed encirclement scenarios; total movement distance is reduced by 50% in some cases, and the number of collisions can be reduced to zero. Dynamic adaptability: when the escapee's speed changes, the pursuer adjusts its strategy in real time, improving the stability of the encirclement. Formation compatibility: supports diverse initial layouts, with significant advantages, especially in flawed encirclement scenarios. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of the method; Figure 2 This is a schematic diagram of the Apollonius circle; Figure 3 This is a diagram showing the occupancy angle; Figure 4 This is a diagram showing the coverage angle; Figure 5 These are schematic diagrams of the improved and original methods for gate array formation, where (a) is the method before improvement and (b) is the method after improvement. Figure 6 These are schematic diagrams of the improved and original methods for the triangular array, where (a) is the method before improvement and (b) is the method after improvement. Figure 7 These are schematic diagrams of the improved and original methods for equilateral triangle formations, where (a) is the method before improvement and (b) is the method after improvement. Figure 8 These are schematic diagrams of the improved and original methods for the pressure positioning formation; where (a) is the method before improvement and (b) is the method after improvement. Figure 9 These are schematic diagrams of the improved and original methods for trapezoidal array formation, where (a) is the method before improvement and (b) is the method after improvement. Detailed Implementation
[0024] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Figure 1 This is a flowchart of the method; A multi-agent cooperative pursuit method based on dynamic trade-off coefficients includes the following steps: S1: Obtain the initial position and maximum speed of the pursuer and the escapee, and the capture radius of the pursuer; S2: Determine if the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit and execute S3 when t≥t max The pursuit ended; S3: Obtain the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; S4: Based on the real-time location information of the escapee and the pursuer, a speed-sensitive factor of the escapee is introduced to obtain an improved trade-off coefficient and generate a dynamically adjusted pursuit strategy. S4: Based on the improved compromise coefficient, calculate the encirclement speed and pursuit speed of each pursuer, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, thereby realizing the coordinated pursuit of the escapee by the pursuers. If the distance between any pursuer and the escapee is less than the capture radius dc, the capture is considered successful; if the escapee is not captured within a short period of time after the specified maximum time, the capture is considered a failure.
[0027] Steps S1 / S2 / S3 / S4 are executed sequentially; In the original algorithm, each pursuer is assumed to be a point mass with single integrator dynamics, and its kinematic model can be expressed as:
[0028] in, The position of pursuer i is indicated by , vi is its control input, and Vi is its maximum speed.
[0029] The escapee is also considered as a point mass, and its kinematic model is as follows:
[0030] in, This indicates the position of the escapee, ve is its control input, Ve is its maximum speed, and Ve > Vp; Furthermore, regarding the capture condition: if the distance between the pursuer i and the escapee is less than the non-zero capture radius dc, that is: It is then assumed that pursuer i has successfully captured the escapee.
[0031] The hunters' goal is to work together to confine the escapee to a certain area and eventually capture them; the escapee's goal is to avoid being captured by any hunter for as long as possible.
[0032] Furthermore, the concepts related to the encirclement algorithm, including occupancy angle, coverage angle, overlapping occupancy angle, and escape angle, are detailed below: Figure 2 This is a schematic diagram of the Apollonius circle; In a plane, there exist two points A and B, and a moving point P, such that PA / PB = k (k is a constant and k ≠ 1). When k > 0 and k ≠ 1, the locus of point P is a circle; when k = 1, the locus is the perpendicular bisector of line segment AB. An important property of the Apollonius circle is that it can help solve a class of geometric problems: finding a point whose distance to two given points is a fixed constant in ratio. In a chase game, the locus of the Apollonius circle represents the points that both the pursuer and the fleeing player can reach simultaneously at their current speeds. For the high-speed fleeing player, the pursuers are all inside the circle, while the fleeing player is outside. Any straight path taken by the fleeing player across the Apollonius circle will result in capture.
[0033] Figure 3 This is a diagram showing the occupancy angle; An occupancy angle is used to describe the area that a single pursuer can control or occupy in a chase game to prevent the escapee from escaping in a specific direction.
[0034] First, we define a local time-varying coordinate system centered on the escapee. In this system, the global coordinates of each pursuer are transformed into local coordinates relative to the escapee. For each pursuer i and the escapee, their maximum velocities are denoted as Vi and Ve, respectively, and the velocity ratio λi is defined as...
[0035] Suppose that the pursuer i and the escapee move in a fixed direction at their maximum speeds and meet at some point m in a finite amount of time. Point m must lie on the Apollonius circle, and its equation can be expressed as:
[0036] Where (xie, yie) is the position of the pursuer i in the local coordinate system. As shown in the figure, the two tangents l1 and l2 from the escapee's position to the Apollonius circle define the occupancy angle θi of the pursuer i. This angle is the angle between the two tangents, calculated using the following formula:
[0037] This angle indicates that if the escapee moves in any direction between l1 and l2, the pursuer i can catch the escapee at some point on the circle. The occupancy angle θi depends only on the ratio of the speed of the pursuer i to that of the faster escapee.
[0038] Figure 4 This is a diagram showing the coverage angle; The local coordinate system ^ can be transformed into a local polar coordinate system L, with polar coordinates Pie(ri, ai).
[0039] Where ri is the polar radius, and ai∈[0, 2π) is the polar angle in L. The n pursuers are dispersed counterclockwise in ascending order of polar angle ai (αi+1≥αi). The relationship between adjacent pursuers is shown in the figure. pie and p(i+1)e, their occupancy angles are denoted by θi and θi+1 respectively. Then the coverage angles of adjacent chasers, i, i+1, are defined as...
[0040] Coverage angles i and i+1 are divided into two categories: 1. When If i, i+1 ≤ 0, then there exists an overlapping occupancy angle, which is the portion where the occupancy angles of two adjacent pursuers overlap. When the coverage angles of two pursuers... i, i+1 When i, i+1 is less than or equal to 0, it means that the occupancy angles of the two pursuers overlap. This means that if the escapee tries to move along the direction of the overlapping area, both pursuers can catch it.
[0041] 2. When If i, i+1 > 0, then an escape angle exists. An escape angle is a directional area that the pursuer cannot cover; that is, an area where the escapee can use to escape pursuit. In this case, if the escapee moves along the direction of the escape angle, it can avoid being captured because neither pursuer can cover this angle.
[0042] Coverage angle reflects the ability of two adjacent pursuers to cooperate in capturing an escapee. A larger coverage angle means that the pursuer can better cover the escapee's direction of movement, thus increasing the probability of capture.
[0043] The original algorithm's trade-off mechanism (reflected in its primary reliance on its own control input) In the original algorithm, the chaser's input control is:
[0044] V is and v ih These are the different velocity components of the pursuer in the direction of encirclement and pursuit.
[0045]
[0046] Where, β i The trade-off coefficient determines whether the pursuer is more inclined to surround or chase the fleeing person. β i The calculation formula is:
[0047] Where: δ i and γi These are the surrounding factor and the hunting factor:
[0048]
[0049] Comprehensive v is and v ih The final control input of the pursuer can be expressed as:
[0050] Where: v is It is the encirclement speed, v ih It refers to pursuit speed, r i k is the distance between the pursuer i and the escapee. i and h i These are used to control the speed of the pursuer's encirclement and the speed of the pursuit, respectively. The trade-off coefficient is designed to balance these two objectives, namely:
[0051] When k i When the threshold is high, pursuers tend to form a tight encirclement to limit the escapee's movement space; when the threshold is high, pursuers focus more on quickly approaching the escapee to shorten the capture time. The dynamic adjustment of the trade-off coefficient is achieved by balancing these two objectives to ensure effective capture of the escapee under different circumstances.
[0052] The final control input for the tracker based on the improved trade-off coefficient is as follows:
[0053] Under different initial conditions, the original algorithm exhibits certain limitations in its capture capability. When the escapee's speed is much greater than the pursuer's, even if an encirclement is formed, the capture may fail because the pursuer cannot approach the escapee in time. The shortcomings of the original algorithm in terms of capture capability are mainly reflected in its failure to consider the direction and magnitude of the escapee's speed, resulting in poor adaptability to changes in the escapee's speed. This indicates that further optimization of the trade-off mechanism is needed to improve the algorithm's performance in variable environments.
[0054] The compromise factor β in the original method iIt relies solely on the relative position and coverage angle differences of the pursuers, without considering the dynamic behavior of the escapee, such as the escapee's speed, direction, and magnitude. This means that even if the escapee is accelerating or changing direction, the pursuer cannot dynamically adjust its strategy to cope with these changes. Furthermore, because the compromise coefficient does not take into account the escapee's dynamic behavior, the pursuer cannot prioritize responding to the escapee's escape intentions in key directions, which may lead to the encirclement being broken and ultimately the capture failing. Based on differential game theory, the optimal escape direction of an escapee is usually strongly correlated with its velocity direction (Isaacs, 1965). Therefore, introducing an escapee velocity factor into the tradeoff coefficients allows the pursuer to dynamically adjust their encirclement strategy, improving capture efficiency. Given the limitations of the tradeoff coefficient design in the original algorithm, we aim to ensure that the escapee's velocity not only determines its movement speed but also directly influences the pursuer's tradeoff between forming an encirclement and approaching the escapee. Depending on the escapee's velocity, the new tradeoff coefficients can be dynamically adjusted. For example, when the escapee's velocity is high, the pursuer needs to focus more on forming a stable encirclement to prevent a rapid breakout; while when the escapee's velocity is low, the pursuer can more actively close the distance to achieve rapid capture.
[0055] This dynamic adjustment mechanism based on the escapee's speed can significantly improve the hunter's adaptability and capture success rate in complex and dynamic environments.
[0056] Define an escapee speed sensitivity factor ηi:
[0057] Where: k is the adjustment coefficient, controlling η i The magnitude of the impact; φ i Let V be the angle between the direction of the escapee's velocity and the relative position of the pursuer i, and let Ve be the escapee's velocity. i Then it represents the speed of the pursuer i.
[0058] The improved method for calculating the compromise factor is as follows: The escapee speed sensitivity factor will be added to the original compromise coefficient calculation method, and the improved compromise coefficient is:
[0059] Furthermore, the dynamically adjusted pursuit strategy is as follows: >0: Increase the compromise coefficient after improvement Strengthen the surrounding speed component The pursuit strategy has been adjusted as follows: Increase the density of the encirclement and reduce the distance between adjacent pursuers; Adjust the heading angle to fill the encirclement gap; Reduce the radial velocity of the pursuer to prevent excessive compression from causing the escapee to break through; <0: Reduce the improved compromise factor Prioritize shortening the distance and adjust the pursuit strategy as follows: Increase the radial velocity component of the pursuer ; Maintain minimum enclosure density to prevent escape by detour.
[0060] When =0, the improved compromise coefficient is 0, and the result obtained by the algorithm formula remains unchanged from the result obtained by the original algorithm (formula before improvement). This method is applicable when N pursuers surround and capture one escapee, where N ≥ 2.
[0061] The dynamically adjusted pursuit strategy also includes arranging the pursuit methods of the pursuers, including: Gate-shaped formation: Two pursuers are arranged in a "gate" shape, simulating an interception and encirclement strategy; the "gate" shape is symmetrically distributed on the left and right sides, leaving an escape passage in the middle; Triangular formation: Three pursuers form an acute or obtuse triangle to simulate a "focused" encirclement; the vertices of the acute or obtuse triangle are located on the escapee's likely path. Equilateral triangle formation: Three pursuers form a tight equilateral triangle, simulating a "high-pressure close-in" strategy; the center of the equilateral triangle is close to the initial position of the escapee. Two pursuers pressure the escapee's movement, while one pursuer is in a predetermined trajectory: Two pursuers pressure the escapee's movement on one side, while the other pursuer is in the trajectory in the direction the escapee is being pressured, simulating a "path prediction" strategy. Trapezoidal formation: The four pursuers are arranged in a trapezoidal shape, simulating the "asymptotic compression" strategy.
[0062] A multi-agent cooperative pursuit device based on dynamic trade-off coefficients, comprising: Acquisition Module I: Used to acquire the initial position and maximum speed of both the pursuer and the escapee, as well as the capture radius of the pursuer; Judgment module: Used to determine whether the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit when t≥t max The pursuit ended; Acquisition Module II: Acquire the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; The generation module is used to generate a dynamic adjustment strategy based on the real-time location information of the escapee and the pursuer, by introducing the escapee's speed sensitivity factor and obtaining an improved trade-off coefficient. The pursuit module is used to calculate the encirclement speed and pursuit speed of each pursuer based on the improved compromise coefficient, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, so as to realize the coordinated pursuit of the pursuers and the escapee.
[0063] A computer device includes: a processor and a memory, the memory storing a program module, characterized in that the program module runs on the processor to implement the method as described in any one of the claims.
[0064] Example The experiment aims to verify the effectiveness of the improved algorithm for different escape strategies under different initial hunter formations. Specific objectives include: (1) Verify the performance difference between the improved algorithm and the original algorithm under different initial formation strategies. (2) Investigate the adaptability of the improved algorithm to different escape strategies. The experiment used a custom simulation system developed using PyCharm and Python. Core libraries: NumPy for mathematical calculations, Matplotlib and PyGame for data visualization, and JSON for data recording. Operating environment: Windows 11 system, standard PC with no special hardware The experimental setup (more complex scenarios may be considered for later experiments) is as follows: Map settings: 500x500 pixel 2D bounded plane, obstacle-free settings, Cartesian coordinate system (origin at the top left corner). The maximum speed of the pursuer is Vp = 9, and the maximum speed of the escapee is Ve = 10 (or expressed as a speed ratio). Capture radius dc = 0.25 Termination condition: The distance between any pursuer and the escapee is less than dc (capture successful) or exceeds the maximum number of time steps (capture failed). In our experiments, we compared the original algorithm and the improved algorithm using different escape strategies. The escape strategies involved are described below.
[0065] The strategy of moving in a straight line along a specific angle: Given an initial angle for the escapee, make the escapee move in a straight line along that angle. This strategy is used as a baseline strategy for comparative experiments. Boundary following strategy: This strategy involves moving the escapee close to the map boundary and along it. This strategy utilizes the boundary as a natural barrier to reduce the encirclement angle. The results of the two algorithms are verified when the encirclement angle is reduced.
[0066] Predicting escape strategy: Calculate the average movement direction of the pursuer based on its historical movement direction, and escape in the opposite direction of the pursuer's average movement direction to verify the results of the two algorithms under more complex escape strategies.
[0067] Repulsive escape strategy: Using the repulsive force calculation formula, calculate the direction of the resultant force of the pursuer, and make the escapee run away in the opposite direction of the calculated resultant force. Verify the results of the two algorithms when the escapee generates a more complex nonlinear trajectory.
[0068] Randomized Bézier Curve Strategy: During initialization, the escapee is randomly assigned a point. The escape trajectory is calculated using the Bézier curve formula, allowing the escapee to move smoothly and randomly. The results of both algorithms are verified under unpredictable escape strategies.
[0069] To better analyze and compare the experimental results, we defined three relevant indicators based on the characteristics of the algorithms, and used these three indicators to compare and evaluate the two algorithms.
[0070] Capture time: Definition: The time required from the start of the experiment until the distance between any pursuer and escapee is less than the capture radius dc, or the time of capture failure.
[0071] Significance: Acquisition time is a core indicator reflecting algorithm efficiency; the shorter the acquisition time, the better the algorithm performance.
[0072] Calculation method: The capture time is represented by a time step. The time step is incremented by one each time the algorithm runs. The final recorded result when the algorithm ends is the capture time.
[0073] Total distance traveled: Definition: The total distance traveled by all pursuers.
[0074] Significance: The total distance traveled reflects the energy consumption of the algorithm and the movement efficiency of the pursuer. The smaller the total distance traveled, the more energy-efficient the algorithm is.
[0075] Calculation method: For each pursuer i, record its position change during the experiment, calculate the distance Di it moves, and the total distance moved is... .
[0076] Collisions: Definition: Records the number of collisions between pursuers and between pursuers and escapees. Significance: The number of collisions reflects the safety of the algorithm and the coordination between the pursuers. The fewer the collisions, the safer the algorithm.
[0077] Calculation method: During the experiment, the distance between the pursuers is monitored in real time. If the distance between any two pursuers or between the pursuer and the escapee is less than the safe distance d_safe, then a collision is recorded.
[0078] To compare the algorithms, we used different escaper strategies and set different initial formations for the pursuers according to the corresponding strategies. We then compared the straight-line escape strategy and the repulsive escape strategy and analyzed the experimental results using three metrics.
[0079] (1) Strategy of moving in a straight line along a specific angle Gate-shaped formation: The two pursuers are distributed in a "gate" shape (such as symmetrical distribution on the left and right sides, leaving an escape passage in the middle), simulating an "interception" encirclement strategy.
[0080] (2) The initial positions of the escapee and the pursuer are: e(250, 250), p1(350, 250), and p2(250, 350), respectively. It can be clearly seen from the trajectory diagram that the improved algorithm is more efficient in capturing. Compared with the original algorithm, the improved algorithm shows an encirclement tendency and can quickly change direction according to the escapee.
[0081]
[0082] Figure 5 These are schematic diagrams of the improved and original methods for gate array formation, where (a) is the method before improvement and (b) is the method after improvement. Triangular formation: Three pursuers form an acute / obtuse triangle (the vertex is located on the escapee's inevitable path), simulating a "focused" encirclement.
[0083] The initial positions of the escapee and the pursuer are: e(250,250), p1(400,250), p2(250,400), p3(350,350);
[0084] Regarding this escape strategy, when a pursuer is positioned along the escapee's likely path, the improved algorithm tends to form an encirclement based on the escapee's state, and is not always superior to the original algorithm. Multiple experiments revealed that for a triangular formation (vertices on the escapee's likely path), the original algorithm is better when the pursuer is far from the escapee's flanks and close to the center (forming an obtuse triangle); the improved algorithm is better when the pursuer is close to the escapee's flanks and far from the center (forming an acute triangle). The same logic applies to a trapezoidal encirclement formation.
[0085] Figure 6 These are schematic diagrams of the improved and original methods for the triangular array, where (a) is the method before improvement and (b) is the method after improvement. (3) Repulsive escape strategy Equilateral triangle formation: Based on the complexity of the repulsive escape strategy, the three pursuers form a tight equilateral triangle (with the center close to the escapee's initial position) to simulate the "high-pressure approach" strategy.
[0086] The initial positions of the escapee and the pursuer are: e(250,250), p1(300,250), p2(225,293), p3(225,206); The result obtained is:
[0087] Figure 7 These are schematic diagrams of the improved and original methods for equilateral triangle formations, where (a) is the method before improvement and (b) is the method after improvement. For the escapee repulsion algorithm, the initial position of the pursuer has formed a completely encircling regular polygonal array. There is no significant difference in the capture results between the before and after the improvement of the pursuit algorithm. The improvement effect of the improved algorithm is limited, which may be related to the escape strategy itself. However, it can also be seen from the trajectory diagram that the improved algorithm is better than the previous algorithm in terms of capture trajectory.
[0088] Figure 8 This is a schematic diagram of the improved and original methods of the pressure movement formation; where (a) is the method before improvement and (b) is the method after improvement; two pursuers use pressure movement, and one pursuer is in a predetermined trajectory, referred to as: pressure movement formation. The initial positions of the escapee and the pursuer are: e(250,250), p1(300,250), p2(220,300), p3(150,150). The result is:
[0089] For this initial formation, the improved algorithm is significantly better than the original algorithm.
[0090] Trapezoidal formation: The four pursuers are arranged in a trapezoidal shape to simulate the "asymptotic compression" strategy and verify the algorithm's efficiency in dynamically shrinking multi-directional encirclement. Figure 9 These are schematic diagrams of the improved and original methods for trapezoidal array formation, where (a) is the method before improvement and (b) is the method after improvement.
[0091] The initial positions of the escapee and pursuer are: e(250,250), p1(280,250), p2(230,220), p3(150,320), p4(280,340). The result is:
[0092] The experimental results show that the improved algorithm has improved in terms of performance indicators, and the trajectory graph also shows that the improved algorithm can form a good encirclement and capture the escapee faster.
[0093] Two pursuers pressure the escapee's movement, while one pursuer is positioned on a predetermined trajectory: Two pursuers pressure the escapee's movement on one side, while the other pursuer is positioned on the trajectory in the direction the escapee is being pressured, simulating a "path prediction" strategy.
[0094] The improved algorithm significantly outperforms the original algorithm in capture time, demonstrating its ability to more efficiently restrict escapee movement and achieve capture. The reduction in total movement distance is attributed to the dynamic adjustment of the tradeoff coefficient, which optimizes the pursuer's movement path. The reduced number of collisions highlights the improved algorithm's precise control of safe distances during the pursuit, effectively preventing collisions and showcasing superior coordination. The improved tradeoff coefficient is dynamically adjusted based on the escapee's speed. When the escapee is moving at high speed, the pursuer prioritizes securing the encirclement; when the escapee is moving at low speed, the pursuer quickly approaches the target. This dynamic adjustment mechanism of the tradeoff coefficient allows the pursuer to employ more flexible strategies when facing different escapee speeds, enabling rapid responses to changes in escapee speed and ensuring successful capture.
[0095] Although the improved algorithm is more efficient than the original algorithm in most cases, we can see from the results that it still performs poorly in some strategy cases, and further research and improvement of the algorithm are needed.
[0096] Experimental results show that the improved algorithm enables the pursuers to dynamically adjust their behavior based on the escapee's real-time speed and strategy: when the escapee's speed is high, the pursuers tend to maintain and consolidate the encirclement; while when the target's speed is low, they can quickly shorten the pursuit distance. Furthermore, the improved algorithm significantly enhances the cooperation among pursuers, enabling rapid responses to changes in the escapee's behavior through efficient communication and strategy coordination.
[0097] It is worth noting that when faced with strategies that escape along the map boundaries, the improved algorithm did not significantly outperform the original algorithm. Preliminary analysis suggests that the spatial constraints of the boundaries may limit the full formation of the encirclement, thus affecting the effectiveness of the algorithm. This issue requires further research and optimization.
[0098] Experimental results show that the improved algorithm exhibits significant advantages in several key performance indicators. The main contributions of this study are reflected in three aspects: First, the proposed dynamic trade-off coefficient design method achieves real-time optimization of the pursuit strategy through the escaper speed factor; second, the impact of different initial formations (including gate, triangle, trapezoid, etc.) on the algorithm performance is systematically studied, providing a theoretical basis for formation selection in practical applications; finally, by comparing the algorithm performance under two typical strategies, straight-line escape and repulsive escape, it is confirmed that the improved algorithm has a wider range of adaptability.
[0099] It is worth noting that the experiments also revealed the performance limitations of the algorithm under the boundary following strategy, which points the way for future research. Future work will focus on solving problems such as 3D spatial expansion and sensor noise compensation, and explore the integration of machine learning techniques with this algorithm to further improve the cooperative pursuit capabilities of multi-agent systems in complex dynamic environments.
[0100] The theoretical findings of this study not only enrich the research content in the field of multi-agent cooperative control, but also provide new technical solutions for practical application scenarios such as UAV formation and intelligent security, and have important academic value and engineering application prospects.
[0101] Under different initial formation conditions, if the initial encirclement is relatively complete, the performance difference between the improved algorithm and the original algorithm is small; however, when the initial encirclement has defects, the improved algorithm shows a significant advantage, being able to quickly adjust the formation according to the escapee's state and complete the capture faster.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-agent cooperative pursuit method based on dynamic trade-off coefficients, characterized in that: Includes the following steps: S1: Obtain the initial position and maximum speed of the pursuer and the escapee, and the capture radius of the pursuer; S2: Determine if the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit and execute S3 when t≥t max The pursuit ended; S3: Obtain the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; S4: Based on the real-time location information of the escapee and the pursuer, a speed-sensitive factor of the escapee is introduced to obtain an improved trade-off coefficient and generate a dynamically adjusted pursuit strategy. S4: Based on the improved compromise coefficient, calculate the encirclement speed and pursuit speed of each pursuer, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, thereby realizing the coordinated pursuit of the escapee by the pursuers.
2. The multi-agent cooperative pursuit method based on dynamic trade-off coefficients according to claim 1, characterized in that: The improved compromise coefficient The expression is as follows: in: For the escapee's speed-sensitive factor, As the bounding factor, For pursuit factors; Where: k is the adjustment coefficient. Let V be the angle between the direction of the escapee's velocity and the direction of the pursuer's velocity, and Ve be the escapee's velocity. i Then it represents the speed of the pursuer i.
3. The multi-agent cooperative pursuit method based on dynamic trade-off coefficient according to claim 1, characterized in that: The dynamically adjusted pursuit strategy is as follows: >0: Increase the compromise coefficient after improvement Strengthen the surrounding speed component The pursuit strategy has been adjusted as follows: Increase the density of the encirclement and reduce the distance between adjacent pursuers; Adjust the heading angle to fill the encirclement gap; Reduce the radial velocity of the pursuer to prevent excessive compression from causing the escapee to break through; <0: Reduce the improved compromise factor Prioritize shortening the distance and adjust the pursuit strategy as follows: Increase the radial velocity component of the pursuer ; Maintain minimum enclosure density to prevent escape by detour.
4. The multi-agent cooperative pursuit method based on dynamic trade-off coefficient according to claim 1, characterized in that: This method is applicable when N pursuers surround and capture one escapee, where N ≥ 2.
5. The multi-agent cooperative pursuit method based on dynamic trade-off coefficient according to claim 1, characterized in that: The dynamically adjusted pursuit strategy also includes arranging the pursuit methods of the pursuers, including: Gate-shaped formation: Two pursuers are arranged in a "gate" shape, simulating an interception and encirclement strategy; the "gate" shape is symmetrically distributed on the left and right sides, leaving an escape passage in the middle; Triangular formation: Three pursuers form an acute or obtuse triangle to simulate a "focused" encirclement; the vertices of the acute or obtuse triangle are located on the escapee's likely path. Equilateral triangle formation: Three pursuers form a tight equilateral triangle, simulating a "high-pressure close-in" strategy; the center of the equilateral triangle is close to the escapee's initial position. Two pursuers pressure the escapee's movement, while one pursuer is in a predetermined trajectory: Two pursuers pressure the escapee's movement on one side, while the other pursuer is in the trajectory in the direction the escapee is being pressured, simulating a "path prediction" strategy. Trapezoidal formation: The four pursuers are arranged in a trapezoidal shape, simulating the "asymptotic compression" strategy.
6. A multi-agent cooperative pursuit device based on dynamic trade-off coefficients, characterized in that: include: Acquisition Module I: Used to acquire the initial position and maximum speed of both the pursuer and the escapee, as well as the capture radius of the pursuer; Judgment module: Used to determine whether the pursuit time t of the pursuer has reached the maximum pursuit time t. max Requirement: When t < t max Continue the pursuit when t≥t max The pursuit ended; Acquisition Module II: Acquire the real-time location of the escapee, and then obtain the locations of all adjacent pursuers based on the escapee; The generation module is used to generate a dynamic adjustment strategy based on the real-time location information of the escapee and the pursuer, by introducing the escapee's speed sensitivity factor and obtaining an improved trade-off coefficient. The pursuit module is used to calculate the encirclement speed and pursuit speed of each pursuer based on the improved compromise coefficient, and synthesize the final speed of each pursuer accordingly to obtain the control input signal of each pursuer, so as to realize the coordinated pursuit of the pursuers and the escapee.
7. A computer device, comprising: A processor and a memory, the memory storing a program module, characterized in that the program module runs on the processor to implement claim 1.
5. The method described in any one of the above.