Heterogeneous agent collaborative search system

By designing a dual-mode switching controller and a regional planning module, the problems of environmental adaptability and mission diversity of the UAV-Unmanned Surface Vessel (USV) cooperative search system in complex aquatic environments were solved. This enabled the USV to achieve efficient and accurate tracking and anti-interference capabilities, and improved the robustness and coverage efficiency of the cooperative search system.

CN121900464APending Publication Date: 2026-04-21HUANENG LANCANG RIVER HYDROPOWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUANENG LANCANG RIVER HYDROPOWER CO LTD
Filing Date
2026-01-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing UAV-Unmanned Surface Vessel (USV) collaborative search systems suffer from poor environmental adaptability and inability to meet multi-dimensional mission requirements in complex aquatic environments, leading to control strategy failures and poor collaborative operation results.

Method used

The system employs a dual-mode switching controller and a regional planning module. Through the collaborative work of the first and second controllers, combined with PID and ADRC control algorithms, it achieves precise control of the unmanned surface vessel. The regional planning module constructs a planning system that actively adapts to unknown disturbances through partition design and stationary inspection modes.

Benefits of technology

It has achieved efficient and accurate tracking and anti-interference capabilities for unmanned surface vessels in complex aquatic environments, solved the problems of environmental adaptability and mission diversity, and improved the robustness and coverage efficiency of the collaborative search system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900464A_ABST
    Figure CN121900464A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of marine search, and discloses a heterogeneous agent collaborative search system. The system comprises an agent cluster (a first type of agent is configured with resident inspection and pursuit execution dual modes, and a second type of agent is configured with an annular inspection mode), a task scheduling module, a dynamic pursuit module and a dual-mode switching controller module. The task scheduling module is responsible for intelligent agent screening, mode switching and control reference synchronization after target identification, the dynamic pursuit module generates a pursuit strategy and path parameters, and the controller module switches the first controller and the second controller according to an interference parameter threshold to respectively guarantee the motion precision and the anti-interference capability. According to the dual-mode control, based on the combination of complementarity mechanism modeling and engineering robustness optimization, accurate adaptation of stable and complex dynamic scenes is achieved, and the collaborative search reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of maritime search technology, specifically relating to a heterogeneous intelligent agent cooperative search system. Background Technology

[0002] In the field of heterogeneous intelligent agent collaborative control technology, the division of labor and cooperation among heterogeneous intelligent agents is the core of improving task efficiency: UAVs, with their advantages of high-altitude perspective, high maneuverability and rapid response, can quickly traverse a wide range of work areas and complete large-scale information acquisition tasks such as environmental situational awareness and preliminary target positioning, providing accurate data support for collaborative operations; UAVs, as the core execution unit for surface operations, rely on their surface navigation stability, payload capacity and anti-interference characteristics, and under the guidance of target information transmitted by UAVs, undertake key tasks such as target approach tracking, on-site handling and continuous monitoring. The two work together to build a complete operational closed loop of "air and space exploration - surface execution", effectively making up for the limitations of a single intelligent agent in terms of operational range and environmental adaptability, and significantly improving the reliability and coverage efficiency of complex water search tasks.

[0003] In UAV-Unmanned Surface Vessel (USV) collaborative search systems, the control accuracy and environmental adaptability of USVs directly determine the overall collaborative operation effectiveness, and their control modules are the core components ensuring the smooth progress of the mission. Currently, existing collaborative planning systems generally adopt a single controller architecture for USV control, that is, a single core control unit uniformly generates control signals such as the USV's navigation trajectory, steering commands, and propulsion parameters. This design aims to simplify the system's control logic and development process, adapting to typical simple operating environments such as calm inland lakes and open nearshore areas by pre-setting fixed control algorithms and parameters. In these scenarios with stable environmental parameters and simple mission requirements, a single controller can realize the basic tracking and handling functions of USVs, meeting the basic needs of initial collaborative operations.

[0004] However, as collaborative search operations expand into complex aquatic scenarios (such as open ocean areas with large waves, nearshore areas with dense obstacles, and areas with complex electromagnetic interference), and as the performance requirements of unmanned surface vessels (USVs) increase, the inherent defects of a single controller architecture have become increasingly apparent, making it difficult to meet actual operational needs. On the one hand, it exhibits extremely poor environmental adaptability. The control algorithms and parameters of a single controller are all preset based on specific typical environmental models, while the actual aquatic operating environment is dynamically affected by multiple factors such as meteorology, hydrology, and topography. Sudden large waves can alter the USV's navigation resistance and stability, complex electromagnetic environments can interfere with control signal transmission, and nearshore obstacles increase the difficulty of trajectory planning. These non-preset environmental factors all pose significant challenges. This can lead to the failure of the control strategy of a single controller, resulting in problems such as unmanned surface vessel (USV) tracking trajectory deviation, response delay, and navigation instability. On the other hand, it cannot match the differentiated task requirements. Different operational scenarios have different performance priority requirements for USVs. For example, maritime search and rescue requires priority to ensure tracking speed, water security requires emphasis on positioning accuracy, and resource exploration requires consideration of stability and energy consumption. However, the optimization direction of a single controller is fixed, making it difficult to simultaneously meet multiple performance objectives. When applied across scenarios, a large amount of manual parameter adjustment is required, which is not only cumbersome but also prone to affecting the overall collaborative operation effect due to untimely parameter adaptation. This has become a key bottleneck restricting the promotion and application of UAV-USV collaborative search systems in complex scenarios. Summary of the Invention

[0005] The purpose of this application is to provide a heterogeneous intelligent agent collaborative search system that can flexibly adapt to the needs of multiple environments and multiple tasks.

[0006] The embodiments of this application can be implemented through the following technical solutions: A heterogeneous intelligent agent cooperative search system, comprising: The intelligent agent cluster includes a first type of intelligent agent and a second type of intelligent agent. The first type of intelligent agent is configured with a resident inspection mode and a pursuit execution mode, and the second type of intelligent agent is configured with a ring patrol mode. The task scheduling module is used to select the best-fit first-type intelligent agent, trigger the mode switching command, and synchronize the control baseline requirements after the mode switching to the controller module when the target object is identified; and to trigger the first-type intelligent agent state reset command after the target is handled. The dynamic tracking module is used to generate a tracking strategy and tracking path parameters based on the distance between the first type of intelligent agent and the target object, and to convert the tracking path parameters into motion control requirements and output them to the controller module. The controller module is a dual-mode switching controller, including a first controller and a second controller, used to receive the control baseline requirements of the task scheduling module and the motion control requirements of the dynamic pursuit module, and control the motion state of the first type of intelligent agent in the pursuit execution mode in combination with the collected interference parameters. If the interference parameter is less than the first threshold, the first controller is used to control the motion accuracy of the first type of intelligent agent; if it is greater than or equal to the first threshold, the second controller is used to control the first type of intelligent agent to improve its anti-interference ability.

[0007] Furthermore, the first controller uses the distance deviation and heading deviation between the first type of intelligent agent and the target object as input signals, performs closed-loop calculation through a proportional-integral-derivative control algorithm, and outputs a directional angle control command for adjusting the attitude of the first type of intelligent agent and a thrust control command for adjusting the speed.

[0008] Furthermore, the second controller estimates the internal dynamic disturbances and external environmental disturbances of the first type of intelligent agent in real time during the search process through an extended state observer, and comprehensively compensates for the disturbances and the target distance deviation, and generates the corresponding direction angle and thrust output through a nonlinear state error feedback control law.

[0009] Furthermore, the first threshold is a set environmental disturbance threshold, and its limiting parameter is the external equivalent acceleration acting on the first type of intelligent agent, with a value range of 1.5. ~2.5 ; Alternatively, the limiting parameter of the first threshold is the rate of change of the attitude angle deviation of the first type of intelligent agent, which ranges from 5° / s to 15° / s.

[0010] Preferably, it also includes a region planning module, which divides the target region into partitions based on the number of the first type of intelligent agents, with each partition matched with a first type of intelligent agent to perform a stationary inspection mode, while controlling the second type of intelligent agents to carry out full-domain detection in a ring patrol mode.

[0011] Furthermore, the region planning module also includes an origin verification unit, which is used to perform an avoidance region compatibility verification on the initial center point of the partition corresponding to the first type of intelligent agent. If the verification passes, the initial center point is used as the final patrol origin of the first type of intelligent agent in the corresponding partition; if the verification fails, the partition boundary is expanded and candidate origins are generated, and the optimal patrol origin is selected from the candidate origins.

[0012] Furthermore, when the straight-line distance between the first type of intelligent agent, which has switched to the pursuit execution mode, and the target object is less than the target object's perception radius, the controller module controls the first type of intelligent agent to pursue the target object according to its potential escape route using a pursuit-escape game strategy; otherwise, it directly pursues the target object according to its potential escape route.

[0013] Furthermore, the execution process of the pursuit-escape game strategy is as follows: S24: Establish a two-player zero-sum game model of pursuit and escape, with the first type of intelligent agent in pursuit execution mode as the pursuer and the target object as the escapee; S25: Solve the Nash equilibrium solution for a two-player zero-sum chase-escape game model based on the differential evolution algorithm to obtain the optimal strategy pair for both the chaser and the escapee; S26: Integrate the preset physical constraints into the optimal strategy to optimize the corresponding pursuit strategy. The physical constraints include, but are not limited to, the maximum speed limit, minimum turning radius limit, smooth steering requirements based on heading angular velocity control, and a collision avoidance strategy based on collision radius. S27: Output the optimized tracking strategy of the first type of intelligent agent, and control the first type of intelligent agent to perform a pursuit operation on the target object according to the optimized tracking strategy.

[0014] Furthermore, the execution steps of the advance avoidance strategy are as follows: S261: The straight path from the first type of intelligent agent to the target object is uniformly sampled into several sampling points at a preset sampling interval, and it is determined in turn whether each sampling point is within the safe avoidance range of the obstacle. If so, the first type of intelligent agent continues to pursue the target object according to its potential escape route; otherwise, step S262 is executed. S262: Locate the first detected obstacle on the straight path as the primary obstacle, and generate a preferred waypoint that meets the requirements for safe detour based on the core parameters of the primary obstacle and the current path direction of the first type of agent.

[0015] Preferably, when the first type of intelligent agent is in the stationary inspection mode, a spiral search method is used to carry out area detection; When multiple second-type intelligent agents are in the ring patrol mode, they adopt a reverse circular patrol cooperative method.

[0016] The heterogeneous intelligent agent cooperative search system provided by the embodiments of this application has at least the following beneficial effects: This application designs a dual-mode control scheme for the target pursuit process of a first type of intelligent agent (unmanned surface vessel), in which the first controller and the second controller work together. The core of this scheme is the deep integration of complementary mechanism modeling and engineering robustness optimization. It can not only rely on mature control theory to ensure control accuracy in stable scenarios, but also cope with the disturbance challenges of complex dynamic scenarios through advanced anti-disturbance mechanisms. This application's regional planning module addresses the challenge of planning for unknown environmental interference in unfamiliar target areas. Through a mechanistic design of "balanced zoning - stationary inspection - full-domain collaboration," it constructs a planning system that proactively adapts to unknown interference. Its core logic is to perform grid-based balanced zoning of the unfamiliar target area based on the number of first-type intelligent agents (unmanned surface vessels), ensuring that the geographical range and potential interference coverage dimensions (such as nearshore / open water, shallow / deep water areas, etc.) of each zone are within controllable thresholds. This avoids the problem of delayed perception or untimely response to unknown interference caused by an excessively large coverage area of ​​a single intelligent agent. Simultaneously, it precisely plans for each zone... Matching a Type I intelligent agent to execute the stationary inspection mode, this mode enables the unmanned surface vessel (USV) to form a closed loop of "local environmental perception - interference feature learning - path dynamic optimization" within a fixed zone. Compared with the all-domain roaming inspection, stationary inspection allows the USV to continuously collect unknown environmental interference data such as water flow speed, obstacle distribution, and local wind and waves within the zone, gradually build a zone-level local environment model, accurately capture the spatiotemporal distribution pattern of interference (such as periodic water flow disturbances in specific areas, the location of hidden obstacles, etc.), and dynamically adjust the interference avoidance strategy and control parameter adaptation scheme of the inspection path based on these local interference features. This application designs a collaboratively optimized search mode for the idle stationing and patrol mechanisms of two types of intelligent agents. The detection range of the unmanned surface vessel is circular. When it is in the idle stationing mechanism, it will start a spiral search mode at the center point of the corresponding partition. The detection range of the unmanned aerial vehicle (UAV) is fan-shaped. In order to enhance the integrity of the entire airspace detection, this application adopts a dual-aircraft anti-phase circular patrol collaborative mode to solve the problems of blind spots at partition boundaries, repeated coverage of areas, and long-term missed detection that are easily caused by the detection range characteristics of the first type of intelligent agent (unmanned surface vessel) and the second type of intelligent agent (UAV). To avoid the problem of having no safe origin points to choose from due to the initial partition center point being surrounded by obstacles, the boundaries of risky partitions are expanded outwards. The expansion distance is determined according to preset standards, which ensures that the expanded area is still within the overall control range of the target area, while significantly broadening the selection space of safe candidate patrol origin points and providing sufficient samples for subsequent screening. This application uses different generation logics for safe detour waypoints in the early avoidance strategy based on real-world scenarios to ensure that tracking of the target object is maintained while avoiding obstacles. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of a heterogeneous intelligent agent cooperative search system according to this application; Figure 2 A flowchart of the screening strategy; Figure 3 A flowchart of a pursuit-escape game strategy; Figure 4 A flowchart for an early avoidance strategy; Figure 5 Example of an advance avoidance strategy Figure 1 ; Figure 6 Example of an advance avoidance strategy Figure 2 ; Figure 7 A schematic diagram showing the second type of intelligent agent executing a circular patrol mechanism after the partition allocation is completed and the first type of intelligent agent executes the idle guarding mechanism; Figure 8 A schematic diagram of the trajectory for collaborative search by the first and second type of intelligent agents; Figure 9 The trajectory diagram of the entire tracking game between USV0 and target 0; Figure 10 The trajectory diagram of the entire tracking game between USV0 and Target 1; Figure 11 The trajectory diagram of the entire tracking game between USV2 and Target 2; Figure 12 The trajectory diagram of the entire tracking game between USV3 and Target 3; Figure 13 A flight path diagram for USV0 pursuing target 0; Figure 14 A flight path diagram for USV0's pursuit of target 1; Figure 15 A flight path diagram for USV2 pursuing target 2; Figure 16 A flight path diagram for USV3 pursuing target 3; Figure 17 Distance map for USV0 pursuing target 0; Figure 18 Distance map for USV0 pursuing target 1; Figure 19 Distance map for USV2 pursuing target 2; Figure 20 Distance map for USV3 pursuing target 3; Figure 21 A course deviation diagram for USV0 pursuing target 0; Figure 22 A deviation diagram of the course of USV0 pursuing target 1; Figure 23 A deviation diagram of the course of USV2 pursuing target 2; Figure 24 A heading deviation diagram for USV3 pursuing target 3. Detailed Implementation

[0018] The present application will now be further described based on preferred embodiments and with reference to the accompanying drawings.

[0019] The vocabulary used in this specification is for illustrative purposes and is not intended to limit the scope of this application. Unless otherwise expressly specified and limited, the terms "set," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, a direct connection, or an indirect connection via an intermediate medium; or they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of these terms in this application.

[0020] Furthermore, in the description of the embodiments of this application, various components on the drawings have been enlarged or reduced for ease of understanding, but this is not intended to limit the scope of protection of this application.

[0021] This application provides a heterogeneous intelligent agent cooperative search system (hereinafter referred to as the "system"). The system is designed to meet the cooperative operation requirements of heterogeneous intelligent agents such as UAVs and unmanned surface vessels. It can effectively break through the limitations of the existing single controller framework of unmanned surface vessels, realize precise scheduling and efficient cooperative control among different types of intelligent agents, and thus improve the performance of cooperative search operations in complex water scenarios.

[0022] Figure 1 A schematic diagram of the system is shown, such as Figure 1 As shown, the system includes an intelligent agent cluster, a task scheduling module, a dynamic pursuit module, and a controller module. These components work together to complete the entire collaborative search operation. The intelligent agent cluster comprises two types of intelligent agents: the first type corresponds to unmanned surface vessels (USVs) responsible for close-range target tracking and handling, while the second type corresponds to unmanned aerial vehicles (UAVs) responsible for large-scale information acquisition. To adapt to the needs of different operational phases, the first type of intelligent agent (USV) is configured with a stationary patrol mode and a pursuit execution mode. The stationary patrol mode is used for area monitoring and inspection when no target object is found, while the pursuit execution mode is used for precise tracking and handling after a target object is discovered. The second type of intelligent agent (UAV) is configured with a ring patrol mode to ensure comprehensive and continuous coverage detection of the target area.

[0023] As the core of the system's global decision-making, the task scheduling module's main function is to coordinate the working states of various modules and intelligent agents. Specifically, after the second type of intelligent agent (UAV) and / or the first type of intelligent agent (UAV) completes target object identification and uploads target information, the task scheduling module will select the optimal first type of intelligent agent based on the target location, the working area environment, and the real-time status of each first type of intelligent agent (UAV) (such as remaining battery power, current location, equipment integrity, etc.). Then, it will trigger a mode switching command to the optimal first type of intelligent agent, switching it from the stationary inspection mode to the pursuit execution mode. At the same time, the task scheduling module will synchronize the control baseline requirements after the mode switch to the controller module, providing a basis for the generation of subsequent control signals. In addition, after the first type of intelligent agent completes the target disposal task, the task scheduling module will also trigger a state reset command to the first type of intelligent agent, restoring it to the stationary inspection mode and waiting for the next task command.

[0024] The dynamic tracking module, serving as the core of the system's motion control, employs a dual-mode switching controller. This controller comprises a first controller and a second controller. Its core function is to receive control baseline requirements from the task scheduling module and, in conjunction with various interference parameters collected by the module itself (including external environmental interference parameters such as wind and wave levels, water flow speed, and electromagnetic interference intensity; and internal system interference parameters such as fluctuations in the agent's power system and sensor measurement errors), control the motion state of the first type of intelligent agent (unmanned surface vessel). Specifically, the dual-mode switching logic is as follows: the system presets a first threshold for interference parameters. After the dynamic tracking module collects interference parameters, it compares them with the first threshold. If the interference parameter is less than the first threshold, it indicates that the current operating environment has relatively low interference, and the first controller is used to ensure the motion accuracy of the first type of intelligent agent (unmanned surface vessel). If the interference parameter is greater than or equal to the first threshold, it indicates that the current operating environment has high interference, and the system switches to the second controller to improve the anti-interference capability of the first type of intelligent agent (unmanned surface vessel), ensuring stable completion of tracking and handling tasks even in complex interference environments.

[0025] Furthermore, regarding the specific design of the first controller, in this system, the first controller uses the distance deviation and heading deviation between the first type of intelligent agent (unmanned surface vessel) and the target object as the core deviation signals. By processing the deviation signals, it finally outputs the azimuth angle and thrust control commands used to adjust the motion state of the first type of intelligent agent (unmanned surface vessel).

[0026] In some specific embodiments of this application, the first controller employs a mature and easy-to-debug PID controller. The output of this PID controller in the continuous time domain can be expressed as: ; (1) in, For a moment t The output control quantity of the time controller, To control the deviation (specifically, in this application, the distance deviation and heading deviation between the first type of intelligent agent and the target object, where the distance deviation is the difference between the actual distance between the two and the preset approach distance, and the heading deviation is the difference between the current heading of the first intelligent agent and the optimal heading towards the target). , , These are the proportional gain, integral gain, and derivative gain of the PID controller. These three gain parameters need to be determined by adjusting according to the dynamic characteristics of the specific system (such as the dynamic response speed and sailing resistance characteristics of the unmanned surface vessel) to ensure control accuracy and system stability.

[0027] In some specific embodiments of this application, , , The values ​​are 5.0, 0.1, and 1.5. During debugging, first fix them... =0、 =0, adjust To ensure that the unmanned surface vessel (USV) tracking has no significant steady-state deviation, gradually increase... Suppress overshoot, and finally adjust Eliminate residual bias.

[0028] In some specific embodiments of this application, the heading deviation is calculated by the target position coordinates transmitted by the UAV and the real-time positioning coordinates of the UAV, specifically the difference between the target azimuth angle and the current heading angle of the UAV.

[0029] In practical applications, the PID controller will collect in real time the two-dimensional Euclidean distance (i.e., the basic data of distance deviation) and heading deviation between the first type of intelligent agent (unmanned surface vessel) in the pursuit execution mode and its corresponding target object. The two are used as the core control deviation input to the controller. The heading angle and thrust of the first type of intelligent agent are used as the output of the PID controller. The output control command is continuously updated in the continuous time domain to continuously adjust the unmanned surface vessel's navigation direction and propulsion force until the first type of intelligent agent reaches the preset handling range of its corresponding target object and completes the tracking and positioning task.

[0030] Accordingly, for the specific design of the second controller, the second controller in this system uses the distance between the first type of intelligent agent (unmanned surface vessel) and the target object, the internal disturbance parameters of the system and the external environmental disturbance parameters as a compensation deviation signal. By processing the compensation deviation signal, the accurate output of the azimuth angle and thrust of the first type of intelligent agent (unmanned surface vessel) is achieved, thereby improving the control stability under interference environment.

[0031] In some specific embodiments of this application, the second controller employs an ADRC (Active Disturbance Rejection Controller) with excellent anti-interference performance. This ADRC controller achieves high-precision control of the first type of intelligent agent (unmanned surface vessel) through the coordinated operation of three core modules: the TD tracking differentiator, the LESO linear extended state observer, and the LSEF linear state error feedback. Its specific working mechanism is as follows: First, the TD tracking differentiator arranges a reasonable transition process for the first type of intelligent agent (unmanned surface vessel), effectively solving the control overshoot problem caused by sudden movement or turning of the target object, and providing the first type of intelligent agent with a smooth target tracking trajectory; Second, the LESO linear extended state observer expands the external environmental disturbance parameters (such as wind and waves, water flow interference) and internal system disturbance parameters (such as power system fluctuations, sensor errors) into new state variables, and estimates the operating state and total disturbance magnitude of the first type of intelligent agent in real time and accurately; Finally, the LSEF linear state error feedback module performs linear processing on the estimated state variables, and combines the aforementioned compensation deviation signal to perform disturbance compensation, and finally outputs control commands to adjust the heading angle and thrust of the first type of intelligent agent (unmanned surface vessel), ensuring that the unmanned surface vessel can still stably track the target in a large disturbance environment and avoid problems such as trajectory deviation and navigation instability.

[0032] In some specific embodiments of this application, considering the core motion characteristics of the unmanned surface vessel (mainly manifested as the second-order dynamic response of outputs such as position and heading to inputs such as thrust and servo deflection) when navigating in water, in order to accurately match the design logic of the ADRC controller, we model the unmanned surface vessel as an equivalent second-order system.

[0033] Specifically, the mathematical model of this second-order unmanned surface vessel system can be expressed as: ; (2) It should be noted here that in the formula... y This represents the core output of the unmanned surface vessel (USV), which can be set to the USV's real-time position coordinates (such as longitude and latitude) or heading angle according to actual control requirements. u This corresponds to the input control quantity of the unmanned surface vessel, namely the thruster thrust or servo deflection command used to adjust the motion state. w External disturbances include environmental factors commonly encountered in water operations, such as wind and wave disturbances and fluctuations in water flow resistance; while a and b These are both inherent parameters of the system. Due to factors such as changes in the load of the unmanned surface vessel and hull wear, these two parameters are difficult to measure accurately in actual operations and are therefore unknown.

[0034] To align with the core design principle of ADRC controllers—"eststimulating and compensating for total disturbances"—we further transform the aforementioned second-order system model, specifically rewriting it as follows: ; (3) in, Newly introduced f For the generalized perturbation term, the key design here is that... f It not only includes the original external disturbances w It also integrates internal system disturbances (such as those caused by parameters). a , b (The model error caused by unknown factors, as well as the delay and friction loss of the unmanned surface vessel's power system, etc.) Through this integration, we can unify complex multi-source disturbances into a single generalized disturbance term, thereby simplifying the logic of subsequent disturbance estimation.

[0035] The core idea of ​​the ADRC controller is to use its internal Linear Extended State Observer (LESO) to handle the generalized disturbance. f Perform real-time, accurate estimation, and then use the obtained disturbance estimate. Introducing control laws In this process, active compensation for disturbances is achieved, ultimately transforming the complex controlled system with unknown disturbances into an easily controllable dual-integral model: ; (4) At this point, the disturbance in the formula has become Because LESO can achieve [the following] f The high-precision estimation makes Approaching 0, the equivalent double integral model can be approximated as an ideal linear system without disturbance, which provides convenience for the design of subsequent control algorithms.

[0036] To more clearly illustrate the design logic of the observer, we first transform the original second-order system model (Equation 2) into a state-space expression: ; (5) in, These represent the output position state quantities (such as actual position or heading angle), the rate of change of output state (such as velocity or angular velocity), and the total disturbance expansion state of the system (including the sum of internal dynamic uncertainties and external environmental disturbances) of the first type of intelligent agent (such as unmanned surface vessel). h This represents the first derivative of the total disturbance, i.e., the rate of change of the disturbance. When designing observers, it is usually assumed that this value is bounded. In this state-space model, we will... Introduced as a new state variable (its physical meaning is the output of the unmanned surface vessel). y The rate of change (such as velocity, angular velocity), and This corresponds to the external location disturbance mentioned earlier. Based on this state-space model, we design a state observer to estimate the system state and total disturbance in real time, thereby improving the system's anti-interference capability and robustness. The state-space form of this observer is as follows: ; (6) in, .

[0037] Furthermore, we convert the above state-space observer into a linear extended state observer (LESO) form adapted to this system, specifically as follows: ; (7) Here L Let be the gain vector of the observer, and its expression is: As will be explained later, the parameters of this gain vector can be directly determined by the observer bandwidth, without the need for repeated adjustments.

[0038] As mentioned earlier, through LESO disturbance compensation, the controlled system is equivalent to a disturbance-free double integral model. At this point, we can use the simple and reliable PD control algorithm to adjust this ideal model. The corresponding control law design is as follows: ; (8) In the formula r To control the set value, that is, the target position or target course that the unmanned surface vessel needs to track, The differential gain of the PD controller, The proportional gain of the PD controller, For the Linear Extended State Observer (LESO) to the output quantity Real-time estimates, For LESO's rate of change of output The real-time estimate. It should be noted that we use [the specified value] here. (That is, adding a new state variable) Observations (observations) rather than directly using The core purpose of participating in the control law calculation is to ensure that the closed-loop transfer function of the system is a pure second-order form without zeros. This design can effectively avoid instability problems such as system overshoot and oscillation caused by zeros, significantly optimize the dynamic response characteristics of the system, and make the tracking motion of the unmanned surface vessel more stable and accurate.

[0039] To simplify the controller parameter tuning process (reduce the difficulty of on-site debugging), we will link the PD controller parameters with the controller bandwidth. Related, specifically designed as follows: ; (9) in, Let be the closed-loop transfer function of the system. s For the Laplace operator in the complex frequency domain, The controller bandwidth is a core tuning parameter of the entire control algorithm, and its physical meaning is the controller's response speed to input commands. The larger the value, the faster the response, but this may cause system oscillations. The smaller the value, the more stable the system, but the response latency will increase. (The proportional gain of the PD controller...) and differential gain All by Directly determined, the specific relationship is as follows With this design, on-site commissioning personnel do not need to adjust separately. and The two parameters only need to be adjusted according to the dynamic characteristics of the unmanned surface vessel (such as maximum speed and turning response time). With just one parameter, the debugging process is greatly simplified, and the engineering practicality of the controller is improved.

[0040] Finally, we make Combining the PD control law (Equation 8) and the LESO model (Equation 7) from the previous text, the complete ADRC controller output expression can be derived: ; (10) in, The system matrix of the observer error system is... E This is the coefficient vector of the rate of change of the disturbance. e Let be the error vector between the observer's state estimate and the system's true state, where . The characteristic polynomial is: ; (11) Where, in the formula , The observer bandwidth is a core tuning parameter of LSEO. The larger the LESO value, the better it is for generalized perturbations. f A faster estimation response time may result in greater sensitivity to sensor measurement noise; conversely, a slower response time leads to a smoother estimation, but with increased latency. It should be noted that the parameters in the LESO gain vector L at this point... Only and Specifically, the parameters can be directly calculated using the pole placement method of the characteristic polynomial, which also eliminates the need for complex manual debugging, further simplifying the entire controller parameter tuning process.

[0041] In some specific embodiments of this application, the first threshold is a set environmental interference threshold, the core function of which is to determine whether the external environmental interference has reached a level that requires triggering a controller mode switch. Specifically, the limiting parameters of the first threshold can be flexibly selected according to the actual application scenario to adapt to different types of environmental interference monitoring needs.

[0042] In a preferred embodiment, the limiting parameter of the first threshold is the external equivalent acceleration acting on the first type of intelligent agent. Here, "external equivalent acceleration" refers to the acceleration component acting on the first type of intelligent agent (e.g., unmanned surface vessel) caused by external environmental factors (such as wind, waves, ocean currents, gusts, etc.). This parameter directly reflects the degree of influence of the external environment on the agent's motion state. In this embodiment, the value range of the first threshold is set to 1.5. ~2.5 The reason for choosing this range is that when the external equivalent acceleration is less than 1.5... When environmental interference is relatively weak, a PID controller can meet the motion control accuracy requirements for the first type of intelligent agent; however, when the external equivalent acceleration exceeds 2.5... At this point, environmental interference is already quite significant, which may cause the agent's motion trajectory to deviate or its posture to become unstable. In this case, it is necessary to switch to the ADRC controller to enhance the anti-interference capability.

[0043] In another preferred embodiment, the limiting parameter of the first threshold is the attitude angle deviation change rate of the first type of intelligent agent. The "attitude angle deviation change rate" refers to the rate at which the deviation between the actual attitude angle and the target attitude angle of the first type of intelligent agent changes over time. This parameter reflects the attitude stability of the first type of intelligent agent during motion. In this embodiment, the value range of the first threshold is set to 5° / s to 15° / s. This range is chosen because when the attitude angle deviation change rate is less than 5° / s, the agent's attitude change is relatively gentle, and the PID controller can maintain its stable motion; while when the attitude angle deviation change rate exceeds 15° / s, the agent's attitude change is more drastic, which may affect its target tracking accuracy or lead to loss of control. In this case, it is necessary to switch to the ADRC controller to quickly adjust its motion state.

[0044] By using the two different limiting parameters and value ranges mentioned above, this application can flexibly determine whether to switch the controller mode based on the specific type and severity of environmental interference, thereby effectively improving the system's anti-interference capability and robustness while ensuring control accuracy.

[0045] This application addresses the target pursuit process of a first-type intelligent agent (unmanned surface vessel), designing a dual-mode control scheme with a first controller and a second controller working collaboratively. The core of this scheme is the deep integration of complementary mechanism modeling and engineering robustness optimization. It leverages mature control theory to ensure control accuracy in stable scenarios while employing advanced disturbance rejection mechanisms to address the challenges of complex dynamic scenarios. Taking the PID controller and ADRC controller used in practice as examples, the two form a precise functional complementarity: the PID controller is built based on classical control theory using frequency domain stability analysis. Its advantages lie in its simple control logic and mature engineering implementation. It can effectively suppress steady-state errors through the synergistic effect of proportional-integral-derivative (PID) controllers, ensuring the system's steady-state control accuracy. In practical applications, its control parameters are specifically optimized based on core dynamic parameters such as the unmanned surface vessel's maximum speed, steering response time, and propulsion system power, ensuring that the unmanned surface vessel can accurately track the target trajectory during stable navigation without significant disturbances, avoiding problems such as heading deviation and accumulated distance errors; while the ADRC controller… The RC controller relies on the core mechanism of the Extended State Observer (ESO). By expanding internal disturbances (such as power system delays and parameter drift caused by load changes) and external disturbances (such as wind and wave impacts and ocean current interference) during the unmanned surface vessel's navigation into new state variables, it achieves real-time and accurate estimation of the system state and total disturbance. Then, through error feedback linearization processing, it transforms the complex nonlinear disturbance system into an easily controllable linear system. Theoretically, it can achieve stable control under unknown system states and complex disturbance scenarios. This characteristic can also provide technical support for operations in unfamiliar target areas mentioned below, effectively alleviating the anti-disturbance control problem caused by the unknown environment of unfamiliar areas. In actual pursuit operations, the dual-mode control dynamically switches according to the scenario: during the stable navigation phase when the target's movement is steady and the sea conditions are calm, the system prioritizes PID control to ensure tracking accuracy and control efficiency; when encountering severe environmental changes such as sudden escape of the target, large waves, or strong currents that cause strong disturbances, making it difficult for PID control to balance tracking accuracy and system stability, the system will seamlessly switch to ADRC control mode. Through ESO, it estimates and compensates for complex disturbances such as ocean current impacts in real time, quickly responding to changes in the target's motion state. This effectively solves the core contradiction between rapid responsiveness and anti-interference during unmanned surface vessel pursuit under complex sea conditions, ensuring the continuous and stable progress of the pursuit mission.

[0046] In heterogeneous intelligent agent collaborative search operations, some scenarios are planned and designed based on known operation areas. In these scenarios, efficient collaboration of intelligent agents can be achieved by relying on preset geographic information, fixed routes, and mature supply point layouts. However, in actual operations, there are still a large number of unfamiliar operation areas (such as unexplored areas in the open sea, unknown waters after sudden disasters, and cross-regional emergency rescue sites). The planning and operation in these areas face several prominent challenges: First, refueling is difficult. Unfamiliar areas lack preset supply stations, limiting the endurance of intelligent agents (especially unmanned surface vessels), requiring dynamic balancing of task coverage and energy consumption during operations. Second, environmental adaptability is highly demanding. Environmental parameters such as meteorology (such as sudden storms), hydrology (such as undercurrents and shoals), and topography (such as underwater obstacles and complex nearshore landforms) in unfamiliar areas are completely unknown and dynamically changing, requiring intelligent agents to quickly adapt to the complex interference of the unknown environment. Third, there is a lack of preset route references. Planning in known areas can optimize paths based on historical route data, but there is no route planning basis in unfamiliar areas, requiring real-time exploration of safe and efficient operation routes. However, the design of existing heterogeneous intelligent agent collaborative control systems is mostly based on known environments. Their collaborative perception algorithms rely on preset environment models, dynamic task allocation logic is optimized based on fixed area characteristics, and precise control strategies are debugged for known interference scenarios. They do not fully consider the unknown and dynamic characteristics of unfamiliar environments, which makes it difficult for the system to achieve efficient collaborative perception (unable to quickly identify key information and potential risks in unknown environments), reasonable dynamic task allocation (difficult to adjust task division according to the real-time status of the agent and environmental changes), and stable and precise control (unable to adapt to unknown interference, resulting in a decrease in control accuracy) in unfamiliar scenarios. These core problems seriously restrict the promotion and application of heterogeneous intelligent agent collaborative search systems in unfamiliar work areas.

[0047] In some preferred embodiments of this application, the system further includes a region planning module, which divides the target area into partitions based on the number of first-type intelligent agents. Each partition is matched with a first-type intelligent agent to perform a resident inspection mode, while controlling a second-type intelligent agent to carry out full-area detection in a ring patrol mode. The regional planning module addresses the challenge of planning for unknown environmental interference in unfamiliar target areas. Through a mechanistic design of "balanced zoning - stationary inspection - full-domain collaboration," it constructs a planning system that proactively adapts to unknown interference. Its core logic is to perform grid-based balanced zoning of the unfamiliar target area based on the number of Type I intelligent agents (unmanned surface vessels), ensuring that the geographical range and potential interference coverage dimensions (such as nearshore / open water, shallow / deep water areas, etc.) of each zone are within controllable thresholds. This avoids problems such as delayed perception or untimely response to unknown interference caused by an excessively large coverage area of ​​a single intelligent agent. Simultaneously, a Type I intelligent agent is precisely matched to each zone to execute a stationary inspection mode. This mode allows the unmanned vessel to form a closed loop of "local environmental perception - interference feature learning - dynamic path optimization" within a fixed zone. Compared to full-domain roaming inspection, stationary inspection allows the unmanned vessel to continuously collect water flow velocity data within the zone. By collecting data on unknown environmental disturbances such as obstacle distribution and local wind and waves, a zone-level local environment model is gradually constructed to accurately capture the spatiotemporal distribution patterns of disturbances (such as periodic water flow disturbances in specific areas and the location of hidden obstacles). Based on these local disturbance characteristics, the disturbance avoidance strategy and control parameter adaptation scheme of the inspection path are dynamically adjusted. At the same time, the regional planning module synchronously controls the second type of intelligent agent (UAV) to carry out full-domain detection in a ring patrol mode, forming a synergistic effect of "UAV rapid full-domain scanning - UAV zone-level fine perception". When the UAV detects a target or a sudden disturbance in the whole domain (such as a large-scale wind and waves passing through), it can quickly synchronize the information to the UAV in the corresponding zone. Since the UAV has been stationed in the zone and accumulated local disturbance data, it can quickly adjust the response strategy based on the existing local environment model to avoid trajectory deviation or mission delay caused by unknown disturbances. This planning model breaks down unknown interference in an unfamiliar global environment into local controllable interference in multiple zones. It relies on stationary inspections to achieve accurate perception and dynamic adaptation of local interference, and then achieves information sharing and resource scheduling optimization through global collaboration. This fundamentally solves the problems of blind planning and delayed interference response in unknown environments, and significantly improves the system's planning and adaptation capabilities and overall operational robustness to interference in unfamiliar areas and unknown environments.

[0048] In some preferred embodiments of this application, in order to solve the problems of blind spots at the intersection of partition boundaries, repeated coverage of areas, and long-term missed detection that are easily caused by the detection range characteristics of the first type of intelligent agent (unmanned surface vessel) and the second type of intelligent agent (unmanned aerial vehicle), a collaboratively optimized search mode is designed for the idle stationing and patrol mechanism of the two types of intelligent agents. The detection range of the unmanned surface vessel (USV) is circular. When it is in an idle stationary mode, it will initiate a spiral search mode at the center point of the corresponding zone. This mode moves in a circle with the center point of the zone as the origin and a fixed length as the initial radius. The radius of the circle fluctuates periodically with time according to a sinusoidal law. At the same time, a fixed orthogonal phase offset is added to the search path of the USV in adjacent zones. The path phase difference avoids excessive overlap of the search areas of adjacent USVs, which greatly improves the coverage uniformity and search efficiency within the zone. The detection range of the unmanned aerial vehicle (UAV) is fan-shaped. In order to enhance the integrity of the entire airspace detection, this application adopts a dual-aircraft anti-phase circular cruise cooperative mode. That is, both UAVs use the geometric center point of the unfamiliar area as the center and a preset fixed length as the cruise radius to perform circular cruise flight at the maximum angular velocity. The cruise phase of the two UAVs always maintains a 180° difference. Through phase complementarity, the coverage gap of the fan-shaped detection range of a single UAV is filled. Together with the spiral search of the USV, a sea-air cooperative full-coverage detection network is formed, which effectively avoids the problems of missed detection and duplicate coverage.

[0049] Furthermore, in some preferred embodiments of this application, unfamiliar areas often contain fixed or temporary obstacles such as reefs, shipwrecks, and no-navigation zones. These avoidance areas directly affect the patrol safety and zoning coverage integrity of the first type of intelligent agent (unmanned surface vessel). To ensure that the unmanned surface vessel can accurately cover the corresponding zone and strictly avoid the risk of obstacle collision when executing the idle stationing mechanism, the area planning module also includes an origin verification unit, which is used to perform avoidance area compatibility verification on the initial center point of the corresponding zone of the first type of intelligent agent. If the verification passes, the initial center point is used as the final patrol origin point of the first type of intelligent agent in the corresponding zone; if the verification fails, the zone boundary is expanded and candidate origin points are generated, and the optimal patrol origin point is selected from the candidate origin points.

[0050] Furthermore, the origin verification unit uses the geometric center point of each partition as a candidate initial patrol origin point, and simulates the movement trajectory of the unmanned surface vessel (USV) initiating a spiral search mode with that point as the center using environmental modeling tools. The key verification step is to check whether the trajectory touches the avoidance areas (obstacles) within and around the partition: if the simulation results show that the USV maintains a safe distance from the avoidance area throughout the complete spiral search (including the radius sinusoidal fluctuation process), then the center point of that partition is directly determined as the final patrol origin point without further adjustment; if the trajectory poses a risk of touching the avoidance area, then the origin optimization and screening process begins.

[0051] Furthermore, to avoid the problem of having no safe origin point to choose from due to the initial partition center point being surrounded by obstacles, the boundaries of the partitions with risks are expanded outwards. The expansion distance is determined according to a preset standard (this preset distance is not less than the sum of the minimum turning radius of the unmanned surface vessel and the safe buffer distance), which ensures that the expanded area is still within the overall control range of the target area, and also significantly broadens the selection space of safe candidate patrol origin points, providing sufficient samples for subsequent screening.

[0052] Within the expanded partitioned area, a grid-based point placement algorithm is used to generate uniformly distributed candidate origins. The grid density needs to be set in conjunction with the partition area and the detection accuracy of the unmanned surface vessel (USV) to ensure that the distance between adjacent candidate origins is no greater than half of the USV's detection range, avoiding the omission of optimal origins due to overly sparse candidate point distribution; at the same time, the number of grids is controlled within the computational capacity to improve the efficiency of subsequent screening. Each generated grid point serves as a potential patrol origin candidate.

[0053] A dual screening process is performed on all candidate origin points: the first step is safety filtering, eliminating candidate points whose distance from the avoidance zone is smaller than a second preset range (the sum of the UAV's emergency braking distance and obstacle buffer distance), ensuring that the remaining candidate origin points are all within the safe zone; the second step is optimal selection, calculating the straight-line distance between each safe candidate origin point and the original zone center point, and selecting the point closest to the original center point as the final patrol origin point. This selection criterion maximizes the match between the UAV's patrol range and the initial zone, reduces coverage deviation, and balances operational efficiency with safety redundancy.

[0054] After the patrol origin is determined by the regional planning module, the unmanned surface vessel can start a spiral search mode based on the origin, and cooperate with the second type of intelligent agent (unmanned aerial vehicle) for the opposite circular cruise to form an initial deployment pattern of "sea-air collaboration and blind-spot safety".

[0055] When a target object is detected in the partition corresponding to any first-type intelligent agent, the second-type intelligent agent, based on the real-time information of the target object, selects the first-type intelligent agent that is suitable for tracking the target object through a filtering strategy, and switches the working state of the first-type intelligent agent from the idle guarding mechanism to the pursuit mechanism until the first-type intelligent agent captures or disposes of the target object.

[0056] This step outlines the agent-based collaborative response mechanism after target detection. Its core lies in building an efficient "any detection - unified scheduling" response link. When a first-type agent (unmanned surface vessel) detects a target object within its corresponding zone, or a second-type agent (unmanned aerial vehicle) discovers a target object within any zone during its patrol, the second-type agent acts as the scheduling core. Based on the target object's real-time information (including position coordinates, speed, heading angle, and surrounding environmental data), it uses a preset filtering strategy to select the first-type agent suitable for the tracking task. Specifically, if the target is detected by a first-type agent, it immediately synchronizes the target information to the second-type agent via a communication link. The second-type agent then assesses the information's real-time performance and completeness, determines the optimal first-type agent, sends a state switching command, and establishes a real-time information synchronization channel. If the target is detected by a second-type agent, it can directly complete the selection and scheduling of first-type agents based on its own accurately collected target information. The selected first-class intelligent agent will immediately switch from the idle stationary mechanism to the pursuit mechanism, and perform the tracking task with the support of the target dynamic information continuously provided by the second-class intelligent agent, until the target object is captured or disposed of.

[0057] Specifically, such as Figure 2 As shown, the screening strategy includes the following steps: S21: Collect the state information of all Type I agents in the idle standby mechanism. The state information includes, but is not limited to, the real-time movement speed, remaining energy, payload capacity and straight-line distance relative to the target object of each Type I agent. The core logic of this step is as follows: On the one hand, the pursuit capability of unmanned surface vessels (USVs) is a comprehensive reflection of multiple factors such as speed, energy, and payload. Focusing only on USVs that are close at hand may result in them having to turn back midway due to insufficient energy, and considering only USVs that are fast may result in them being unable to complete the task due to the lack of capture devices. Therefore, it is necessary to collect status information in all dimensions. On the other hand, collecting information on USVs with idle stationary mechanisms can avoid including other intelligent agents that are performing other tasks in the screening process, reduce invalid calculations, and improve screening efficiency.

[0058] S22: Based on the environmental constraints of the target area and the real-time information of the target object, predict the potential escape route of the target object. The environmental constraints of the target area include, but are not limited to, the water flow velocity, wave level, and the distribution of obstacles within the area. The core logic of this step is that environmental factors in the target waters have a significant impact on the trajectory of the target object. When moving downstream, the target's escape velocity is superimposed on the water flow velocity, and obstacles force the target object to change course. If the route is predicted solely based on the target object's current state, it is prone to deviation from the actual escape trajectory, leading to inefficient tracking by the unmanned surface vessel. Therefore, environmental constraints need to be incorporated as important parameters into the prediction model to make the escape route more closely resemble the actual scenario.

[0059] S23: Based on the potential escape routes of the target object, select the first type of intelligent agent with the shortest expected arrival time on the potential escape route, and identify it as the intelligent agent to perform the target object tracking task.

[0060] This step selects an agent capable of quickly intercepting the target object as the agent to perform the pursuit and tracking task. The determination of this agent is not simply based on "speed / time", but rather incorporates environmental influences and detour costs to ensure the authenticity and accuracy of the results.

[0061] Furthermore, once the first type of agent is identified to carry out the pursuit task, its operating state switches from the idle standby mechanism to the pursuit mechanism. The corresponding pursuit strategy is as follows: If the straight-line distance between the first type of agent and the target object is greater than the target object's perception radius, the agent will directly pursue along the target object's potential escape route. If the initial distance between the first type of agent and the target object is relatively close when switching to the pursuit mechanism, or if the target object enters the range that the first type of agent can perceive (i.e., the target object can perceive the distance between itself and the first type of agent carrying out the pursuit task in real time) as the distance between the two gradually decreases during the pursuit, the target object will enter the range that the first type of agent can perceive (i.e., the target object can perceive the distance between itself and the first type of agent carrying out the pursuit task in real time). In this case, direct pursuit based solely on the predicted escape route can no longer meet the accuracy and timeliness requirements of the pursuit task. Therefore, a pursuit-escape game strategy is required to carry out the pursuit operation based on the target object's potential escape route in order to ensure the pursuit effect.

[0062] Furthermore, such as Figure 3 As shown, the execution process of the pursuit-escape game strategy is as follows: S24: Establish a two-player zero-sum game model in which the first type of intelligent agent, which switches to the pursuit mechanism, is the pursuer and the target object is the escapee.

[0063] Step S24 visualizes the pursuit-escape adversarial relationship between "Type I intelligent agents—dynamic target objects" through a game theory model, using the adversarial characteristics of zero-sum games to characterize the essential conflict of interest between the two parties. Unlike the classic one-sided optimal control problem, in this two-person zero-sum game model, both the pursuer and the escapee have their own dedicated performance index functions (i.e., payoff functions).

[0064] Specifically, the payment function for the first type of intelligent agent performing the pursuit mission is: The payoff function for the target object is Its expression is as follows: ; Where distance represents the real-time straight-line distance between the first type of intelligent agent and the target object. i For a first-class intelligent agent, the specific policy is one of the available policy options. j The specific strategy in the set of optional strategies for the target object.

[0065] From this definition, it is easy to see that for a two-player zero-sum game, the payoff functions of both parties satisfy... That is, an increase in the gains of one party will necessarily correspond to a decrease in the gains of the other party, which is consistent with the conflicting interests characteristic of the pursuit and confrontation.

[0066] S25: Solve the Nash equilibrium solution for a two-player zero-sum game model based on the differential evolution algorithm to obtain the optimal strategy pair for both the pursuer and the pursuer; The Nash equilibrium here refers to a stable combination of strategies in which neither the pursuer nor the fleeing player can improve their own payoff by unilaterally changing their own strategies, given the opponent's strategy. Its expression is as follows: ; In this study, the game payoff matrix is ​​constructed using the distance parameter between the first type of agent and the target object. V The final payoffs for both sides after the game reaches equilibrium. J This represents the payoff calculation function in the game process. The set of all possible policies for the first type of agent. n The number of policies for the first type of agent. The set of all possible strategies for the target object. m The number of strategies for the target object.

[0067] In the specific solution, the differential evolution algorithm performs differential mutation operations (generating mutation vectors based on the differences between individuals in the population) and crossover operations (recombining the mutation vectors with the original individuals) on the policy variables of the first type of agent and the target object. Combined with selection operations, it retains the best individual and converges to the optimal policy pair corresponding to the Nash equilibrium solution after iterative optimization.

[0068] S26: Integrate the preset physical constraints into the optimal strategy to optimize the corresponding pursuit strategy. The physical constraints include, but are not limited to, the maximum speed limit, minimum turning radius limit, smooth steering requirements based on heading angular velocity control, and a collision avoidance strategy based on collision radius deflection heading collision advance avoidance strategy for the first type of intelligent agent.

[0069] In some specific embodiments of this application, such as Figure 4 As shown, the specific steps for implementing the advance avoidance strategy are as follows: S261: The straight path from the first type of intelligent agent to the target object is uniformly sampled into several sampling points at a preset sampling interval, and each sampling point is determined to be within the safe avoidance range of the obstacle. If so, the first type of intelligent agent continues to pursue the target object according to its potential escape route; otherwise, step S262 is executed.

[0070] In some specific embodiments of this application, in step S261, the sampling interval can be set according to the movement speed of the first type of intelligent agent and the distribution density of obstacles, and can be dynamically adjusted based on actual needs, and each sampling point records corresponding three-dimensional coordinate information.

[0071] In some specific embodiments of this application, in step S261, the safe avoidance range of the obstacle is a spatial area preset based on the size and type of the obstacle (such as static obstacles and dynamic obstacles) and safety redundancy requirements (for example, for fixed obstacles, the safe avoidance range is a circular area with the center of the obstacle as the center, the radius of the obstacle plus the radius of the first type of intelligent agent itself plus a safety distance of 2-3 meters).

[0072] S262: Locate the first detected obstacle on the straight path as the primary obstacle, and generate a preferred waypoint that meets the requirements for safe detour based on the core parameters of the primary obstacle and the current path direction of the first type of agent.

[0073] In some specific embodiments of this application, the reason for prioritizing the handling of the first obstacle in step S262 is that it directly obstructs the current pursuit path, and timely avoidance can prevent subsequent path planning from becoming passive.

[0074] In some specific embodiments of this application, in step S262, the core parameters of the main obstacle include the three-dimensional coordinates of the center point, physical radius information, and real-time location information.

[0075] It should be noted that the generation of safe detour waypoints in step S262 is divided into two scenarios based on the actual situation. One is a normal tracking scenario where the target object is not obscured by obstacles and only the agent's straight-line tracking path is blocked. The other is a dynamic target tracking scenario where the target object is directly blocked by obstacles and the agent needs to detour around the obstacles while maintaining tracking of the target object. The two scenarios correspond to different waypoint generation logic: Let's look at the first scenario: when the target object is not blocked by obstacles, the specific generation method is as follows: Using the center point of the main obstacle as the center, and the obstacle's safe radius (i.e., the sum of the obstacle's own radius and the preset safe redundancy distance) as the reference, an circumscribed square is drawn. The four vertices of this circumscribed square are the initial candidate detour points. These four points are chosen because they are furthest away from the core area of ​​the obstacle while covering the main detour directions around the obstacle, ensuring the safety and diversity of the detour path. Then, based on the preset first filtering condition, these four candidate points are screened to finally determine the optimal detour point.

[0076] Furthermore, the first filtering conditions include: 1. The straight-line distance between the candidate point and the current position of the first type of intelligent agent is not less than 100 meters. This distance threshold is set in combination with the control response speed of the intelligent agent and the accuracy of the detour path planning. The purpose is to avoid generating invalid waypoints that are too close to the current position, which would cause the detour actions to be too hasty or frequent; 2. The planned path from the current position of the first type of intelligent agent to the candidate point must be collision-free, that is, all sampling points on the path must not fall within the safe avoidance range of any obstacle; 3. The planned path from the candidate point to the current position of the target object must be collision-free, ensuring that the subsequent pursuit path can be quickly connected after the detour without affecting the pursuit efficiency. From the candidate points that simultaneously meet these three conditions, the one closest to the target object is selected as the optimal detour point. This can minimize the length of the pursuit path after the detour while ensuring safety, thereby improving the pursuit efficiency.

[0077] In the second scenario, when the path to the target object is blocked by an obstacle, it's necessary to plan a detour route based on the target object's dynamic motion state to avoid losing the target simply by bypassing the obstacle. The specific steps are: First, the agent's perception module collects historical motion data of the target object (including historical speed, heading angle, etc.). After filtering and denoising, a forward predicted path trajectory for the target object over a future period (e.g., 5-10 seconds, dynamically adjusted based on the target's speed) is generated based on a preset motion prediction model (e.g., Kalman filter model, particle filter model, etc.). Simultaneously, to handle the extreme case where the entire forward predicted path is obscured by obstacles, a reverse motion path trajectory (i.e., a reverse extension trajectory based on historical motion trends) is also generated as an alternative detour direction. Next, points on both the forward predicted path trajectory and the reverse movement path trajectory are sampled and then filtered according to preset effective waypoint selection criteria. These criteria include: 1. The sampled point itself is not within the safe avoidance range of any obstacle, ensuring the safety of the waypoint itself; 2. The sampled point must be located within a preset target area—an area defined by the target object's current position and a preset pursuit operation radius (e.g., 500 meters), ensuring that the target object can still be tracked after detour; 3. The planned path from the current position of the first-type agent to the sampled point must be collision-free, meaning the path will not conflict with any obstacles throughout. After filtering, the optimal detour point is selected from the effective waypoints of the forward predicted path trajectory that are closest to the current position of the first-type agent, prioritizing the continuity and efficiency of the pursuit; if no effective waypoint meets the criteria in the forward predicted path trajectory, the optimal detour point is selected from the effective waypoints of the reverse movement path trajectory that are closest to the current position of the first-type agent, ensuring safe avoidance and maintaining tracking of the target object even in extreme cases.

[0078] The following will combine Figure 5 and Figure 6 The specific implementation methods of the advance avoidance strategy of this application are described in detail.

[0079] like Figure 5As shown, this corresponds to step S261 in the advance avoidance strategy of this application, specifically the path collision prediction scenario for the first type of intelligent agent (i.e., the pursuing subject performing the pursuit task). In this embodiment, there is a maximum-sized obstacle (i.e., the main obstacle) on the initial straight-line pursuit path between the first type of intelligent agent and the target object (i.e., the escaping subject). By uniformly sampling and verifying the initial straight-line pursuit path, it can be seen that some sampling points on the straight-line path fall within the safe avoidance range of the core obstacle. The aforementioned safe avoidance range is a spatial area preset based on the physical contour size of the core obstacle and environmental safety requirements. Sampling points falling into this area indicates that there is a collision risk on the initial straight-line path, and subsequent detour waypoint planning steps need to be executed. Further combined with scene characteristics, it is determined that the core obstacle only blocks the initial straight-line pursuit path of the first type of intelligent agent, and the target object is not obscured by the obstacle. Therefore, the application scenario corresponding to this embodiment is a normal tracking scenario where "the target object is not blocked by the obstacle", and the following safe detour waypoint generation logic is applicable.

[0080] like Figure 6 The diagram illustrates the process of generating the optimal waypoint in a typical tracking scenario. The specific technical implementation is as follows: First, based on the inherent parameters of the core obstacle (including the three-dimensional coordinates of its center point and its own radius), a preset safety redundancy distance is superimposed (this redundancy distance is determined based on the motion mobility parameters of the first type of intelligent agent and the collision warning response threshold, used to ensure a buffer space for avoidance operations), thus constructing a safe avoidance zone for the core obstacle. Figure 6 The ring-shaped space area formed by the "obstacle body + green ring" is defined in the middle; secondly, based on the center point of the core obstacle, a circumscribed square of the aforementioned safety avoidance area is drawn. Figure 7 The green-bordered square), the four vertices of the circumscribed square ( Figure 6 The points marked with green pentagrams are the initial candidate detour points. Selecting the vertices of the circumscribed square as candidate points ensures that all candidate points are outside the safe avoidance zone and cover the main detour directions around the core obstacle, providing sufficient samples for subsequent optimal path selection. Next, based on the preset first filtering condition and the distance to the target object, the candidate point with the shortest straight-line distance to the target object (i.e., the point marked with a green pentagram) is selected from the candidate points that have passed the compliance verification. Figure 6 The green pentagram in the upper left corner is selected as the optimal detour waypoint. This selection logic can minimize the length of the pursuit path after the detour while ensuring the safety of the detour, thus ensuring the efficiency of the overall pursuit mission.

[0081] S27: Output the optimized tracking strategy of the first type of intelligent agent, and control the first type of intelligent agent to perform a pursuit operation on the target object according to the optimized tracking strategy.

[0082] When the first type of agent completes its pursuit of the target object (e.g., successfully locks onto the target, or the target is handed over to the subsequent processing unit), or when the pursuit task termination trigger condition occurs (e.g., the target leaves the tracking range and exceeds the preset search time limit, or the system issues a task abort command), the system will switch the first type of agent's mechanism back to the idle standby mechanism. Simultaneously, the system will access the agent's information and locate its corresponding dedicated partition. Then, based on the relative position of the first type of agent's current location and its corresponding partition, combined with real-time environmental data, the system will generate the optimal return path using a dynamic path planning algorithm. During the first type of agent's return to its corresponding partition, it enters the set of selectable agents for the target object's reappearance. That is, the first type of agent may be selected to pursue other target objects during its return to its corresponding partition, or it may be selected to pursue other target objects after returning to its corresponding partition.

[0083]

Example

[0084] like Figure 8The diagram shows the preset trajectories and initial motion logic design for the UAVs and USVs: The UAV's preset trajectory is a circular trajectory with a fixed radius centered on the center of the work area. This design allows the UAV to perform a uniform coverage scan of the work area, ensuring timely detection of target objects at different locations within the area. The USV's preset trajectory is a double-helix trajectory with the centers of the four quadrants of the work area as references. This trajectory design enables the USV to perform detailed searches of each quadrant area in a flexible and comprehensive manner, improving the detection capability of dispersed targets. Initially, all UAVs and USVs are evenly distributed along the left boundary of the work area, and the system controls each device based on A... The algorithm moves to its respective preset trajectory. After reaching the preset trajectory, the UAV and USV continue to move along their respective preset trajectories to maintain the area detection state.

[0085] like Figures 9-12 The diagram illustrates the entire pursuit and maneuver process between a USV and a target object, from detection to disposal. The core control logic is as follows: During the pursuit, the system calculates the straight-line distance between the USV and the target object in real time. If this distance exceeds the target object's detection radius (400m), the target object cannot perceive the USV's pursuit. In this case, the USV does not need to consider the target's countermeasures and directly determines the target object's real-time status (if the target is within the UAV / USV detection area, its position within a preset time frame is predicted based on its current position and velocity; if the target leaves the detection area, its current position is predicted based on the last detected position and velocity information). The system then uses A... The algorithm plans a path and moves towards the target object. If the distance is less than or equal to 400m, the USV enters the target object's perception area, and a two-player zero-sum game mechanism is initiated. The target object's payoff is the distance d between the two, and the USV's payoff is -d. The system uses a differential evolution algorithm to solve the Nash equilibrium solution of this game model, obtains the optimal pursuit strategy for the USV, and then plans the target waypoint to achieve accurate capture of the target object.

[0086] like Figures 13-16The diagram illustrates the application process of the control algorithm for multi-USV collaborative pursuit of a target object: Upon the appearance of the target object, the system employs a PID-ADRC adaptive switching control algorithm to achieve collaborative control of four USVs. This algorithm dynamically switches controller types through a built-in disturbance perception mechanism: when external disturbances (such as waves, wind loads, etc.) are strong, the system automatically activates the more robust ADRC controller to enhance the ability to suppress disturbances; when external disturbances are weak and the environment is stable, the system switches to the faster-responding and more energy-efficient PID controller to improve the USV's navigation efficiency and control accuracy. During the actual pursuit process, the system fully considers the impact of external disturbances such as waves on USV navigation. The trajectory results show that the USVs can stably navigate along the reference desired trajectory, with only slight deviations occurring in a few short periods, and can quickly recover stability, fully demonstrating the control system's good adaptability and real-time response capability to complex disturbance environments.

[0087] like Figures 17-20 As shown, this is the dynamic change curve of the distance between the USV and the target object during the pursuit process; Figures 21-24 As shown, the curves represent the evolution of the deviation of the USV from the target object's desired course. These two sets of error curves intuitively reflect the control effect of the proposed control system on path tracking accuracy and course stability in actual pursuit missions: throughout the pursuit process, the course error remains stable, and the distance between the USV and the target object continues to decrease; even under extreme environmental disturbances such as strong winds and waves, the system significantly enhances the suppression effect on sudden disturbances through the PID-ADRC adaptive switching control algorithm, enabling the USV to maintain good path tracking performance and course maintenance capability under complex sea conditions, fully verifying the comprehensive technical advantages of the control algorithm in terms of anti-interference and stability.

[0088] The above embodiments of this application present preferred embodiments and specific implementation methods in the order of their technical implementation. While the dual-mode switching control scheme for the first type of intelligent agent (unmanned surface vessel) is described later, this scheme possesses independent and complete technical implementation logic. Even without relying on the assistance of other preferred embodiments such as the area planning module mentioned above, it can still achieve excellent control performance. This is because the core of dual-mode switching control lies in the dynamic threshold judgment based on interference parameters. An independent control system is constructed through the functional complementarity of the first and second controllers. No other modules are required for prior support; it only needs to receive the control baseline requirements of the task scheduling module and the motion control requirements of the dynamic pursuit module to autonomously complete precision control in stable scenarios and anti-disturbance adjustment in complex disturbance scenarios. Through the deep integration of complementary mechanism modeling and engineering robustness optimization, it effectively solves the core contradiction between accuracy and anti-disturbance in target pursuit, ensuring the stable and efficient advancement of collaborative search operations in different environments.

[0089] The specific embodiments of this application have been described in detail above. For those skilled in the art, several improvements and modifications can be made to this application without departing from the principle of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A heterogeneous intelligent agent cooperative search system, characterized in that, include: The intelligent agent cluster includes a first type of intelligent agent and a second type of intelligent agent. The first type of intelligent agent is configured with a resident inspection mode and a pursuit execution mode, and the second type of intelligent agent is configured with a ring patrol mode. The task scheduling module is used to select the best-fit first-type intelligent agent, trigger the mode switching command, and synchronize the control baseline requirements after the mode switching to the controller module when the target object is detected; and to trigger the first-type intelligent agent state reset command after the target is handled. The dynamic tracking module is used to generate a tracking strategy and tracking path parameters based on the distance between the first type of intelligent agent and the target object, and to convert the tracking path parameters into motion control requirements and output them to the controller module. The controller module is a dual-mode switching controller, including a first controller and a second controller, used to receive the control baseline requirements of the task scheduling module and the motion control requirements of the dynamic pursuit module, and control the motion state of the first type of intelligent agent in the pursuit execution mode in combination with the collected interference parameters. If the interference parameter is less than the first threshold, the first controller is used to control the motion accuracy of the first type of intelligent agent; if it is greater than or equal to the first threshold, the second controller is used to control the first type of intelligent agent to improve its anti-interference ability.

2. The heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: The first controller uses the distance deviation and heading deviation between the first type of intelligent agent and the target object as input signals, performs closed-loop calculation through a proportional-integral-derivative control algorithm, and outputs a direction angle control command for adjusting the attitude of the first type of intelligent agent and a thrust control command for adjusting the speed.

3. The heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: The second controller estimates the internal dynamic disturbances and external environmental disturbances of the first type of intelligent agent in real time during the search process through an extended state observer, and comprehensively compensates for the disturbances and the target distance deviation, and generates the corresponding direction angle and thrust output through a linear state error feedback control law.

4. The heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: The first threshold is a set environmental disturbance threshold, and its limiting parameter is the external equivalent acceleration acting on the first type of intelligent agent, with a value range of 1.

5. ~2.5 ; Alternatively, the limiting parameter of the first threshold is the rate of change of the attitude angle deviation of the first type of intelligent agent, which ranges from 5° / s to 15° / s.

5. The heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: It also includes a regional planning module, which divides the target area into partitions based on the number of the first type of intelligent agents. Each partition is matched with a first type of intelligent agent to perform a stationary inspection mode, while controlling the second type of intelligent agents to carry out full-area detection in a ring patrol mode.

6. The heterogeneous intelligent agent cooperative search system according to claim 5, characterized in that: The region planning module also includes an origin verification unit, which is used to perform avoidance region compatibility verification on the initial center point of the partition corresponding to the first type of intelligent agent. If the verification passes, the initial center point is used as the final patrol origin of the first type of intelligent agent in the corresponding partition. If the verification fails, the partition boundary is expanded and candidate origins are generated, and the optimal patrol origin is selected from the candidate origins.

7. The heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: When the straight-line distance between the first type of agent, which has switched to the pursuit execution mode, and the target object is less than the target object's perception radius, the controller module controls the first type of agent to pursue the target object according to its potential escape route using a pursuit-escape game strategy; otherwise, it directly pursues the target object according to its potential escape route.

8. A heterogeneous intelligent agent cooperative search system according to claim 7, characterized in that, The execution process of the pursuit and escape game strategy is as follows: S24: Establish a two-player zero-sum game model of pursuit and escape, with the first type of intelligent agent in pursuit execution mode as the pursuer and the target object as the escapee; S25: Solve the Nash equilibrium solution for a two-player zero-sum chase-escape game model based on the differential evolution algorithm to obtain the optimal strategy pair for both the chaser and the escapee; S26: Integrate the preset physical constraints into the optimal strategy to optimize the corresponding pursuit strategy. The physical constraints include, but are not limited to, the maximum speed limit, minimum turning radius limit, smooth steering requirements based on heading angular velocity control, and a collision avoidance strategy based on collision radius. S27: Output the optimized tracking strategy of the first type of intelligent agent, and control the first type of intelligent agent to perform a pursuit operation on the target object according to the optimized tracking strategy.

9. A heterogeneous intelligent agent cooperative search system according to claim 8, characterized in that, The execution steps of the advance avoidance strategy are as follows: S261: The straight path from the first type of intelligent agent to the target object is uniformly sampled into several sampling points at a preset sampling interval, and it is determined in turn whether each sampling point is within the safe avoidance range of the obstacle. If so, the first type of intelligent agent continues to pursue the target object according to its potential escape route. Otherwise, proceed to step S262; S262: Locate the first detected obstacle on the straight path as the primary obstacle, and generate a preferred waypoint that meets the requirements for safe detour based on the core parameters of the primary obstacle and the current path direction of the first type of agent.

10. A heterogeneous intelligent agent cooperative search system according to claim 1, characterized in that: When the first type of intelligent agent is in the stationary inspection mode, it uses a spiral search method to carry out area detection; When multiple second-type intelligent agents are in the ring patrol mode, they adopt a reverse circular patrol cooperative method.