Unmanned ship cluster control method under sensing and communication limited conditions and related equipment

By acquiring the local state information of unmanned surface vessels (USVs) and the adaptive speed obstacle area reward function, obstacle avoidance and encirclement control of USVs under conditions of limited perception and communication were achieved, improving the collision avoidance capability and encirclement success rate of heterogeneous USV formations.

CN121957136APending Publication Date: 2026-05-01CENT SOUTH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-03-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Under conditions of limited perception and communication, heterogeneous unmanned surface vessel (USV) formations face challenges such as limited communication, poor stability and robustness of collision avoidance control during encirclement and capture missions.

Method used

By acquiring the local state information of each pursuit-type unmanned surface vessel (USV), its current situation is determined, and the USV is controlled to surround or attack based on the determination result. The adaptive speed obstacle area and the surround/attack reward function are used for obstacle avoidance and capture, reducing the dependence on high-quality communication and host computer.

Benefits of technology

It improves the obstacle avoidance capability and capture success rate of unmanned surface vessels, is suitable for complex aquatic environments, and solves the problems of limited communication and poor stability and robustness of collision avoidance control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121957136A_ABST
    Figure CN121957136A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned ship cluster control method and related equipment under the condition of limited perception and communication, and relates to the technical field of cooperative control of multiple unmanned ships, and the method comprises the steps: obtaining the local state information of each pursuit type unmanned ship in an unmanned ship cluster; according to the local state information of each pursuit type unmanned ship, the current situation of each pursuit type unmanned ship is judged, and a judgment result is obtained; according to the judgment result, each pursuing type unmanned ship is controlled to surround or attack the surrounding target; the problems that the communication is limited and the collision avoidance control stability and robustness are poor in the process of executing the hunting task by the heterogeneous unmanned ship formation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-unmanned surface vessel (USV) cooperative control technology, and in particular to a method and related equipment for USV swarm control under conditions of limited perception and communication. Background Technology

[0002] Unmanned surface vessels (USVs) are small surface vessels capable of autonomous navigation and mission completion in the ocean. In recent years, with the continuous development and exploration of the ocean globally, USVs have seen rapid development due to their small size, low cost, high speed, and flexibility, making them increasingly valuable for various applications.

[0003] The pursuit-escape game problem is a typical issue in unmanned surface vessel (USV) swarm research. In recent years, with the rapid development of artificial intelligence technology, learning-based game theory methods have been increasingly applied to the pursuit-escape game problem of USVs. This method utilizes data, knowledge, and rules, combined with machine learning, to establish and optimize decision-making models for various action entities during the game, providing reliable data support for USV applications.

[0004] During unmanned surface vessel (USV) swarm operations, obstacles may interfere with the normal operation of USVs, disrupt their formation, or even cause damage to them. Therefore, obstacle avoidance and mutual avoidance are research hotspots in USV swarm operations and are key to ensuring mission success.

[0005] With breakthroughs in technologies such as satellite communication and 5G, problems in wireless communication, such as limited bandwidth, time delay, transmission noise, and intermittent interruptions, have been improved. Communication constraints in multi-agent cooperative control are gradually being relaxed, making the implementation of multi-unmanned surface vessel (USV) cooperative control possible. Compared to land and near-shore communication, long-range wireless communication is still highly susceptible to weather and environmental factors, and therefore, multi-USV cooperative strategies must explicitly consider the communication topology.

[0006] Current research on collision avoidance path planning for multiple unmanned surface vessels (USVs) has been widely reported in terms of optimizing swarm performance through collaborative path planning, task allocation, and resource allocation. The control systems primarily employ centralized and distributed control. Centralized control, given known system environment information, allows the controller to flexibly coordinate multiple USVs to complete predetermined tasks, mainly applied to system formation control. However, with the increase in the number of USVs, the computational resources required for centralized control grow exponentially, placing higher demands on the controller's computing power. Furthermore, communication delays between each USV and the central controller further affect system stability. In addition, centralized control methods focus on calculating the time-optimal navigation trajectory in environments with static obstacles. For complex environments containing dynamic obstacles, centralized planning systems struggle to scale to handle a large number of individuals and lack robustness to motion errors and USV malfunctions. Distributed control, on the other hand, distributes control tasks to multiple USVs in the system. Each USV independently makes decisions based on data acquired from sensors installed on its surface, collaboratively completing the work of the entire system. Distributed strategies are suitable for USV swarm control systems, offering relatively lower computational complexity, higher scalability and flexibility, and robustness to individual USV navigation errors and USV malfunctions within the swarm system. In practical applications, the number of unmanned surface vessels (USVs) can be flexibly adjusted within a certain range, without being limited by distributed control systems. Therefore, with the increasing development of unmanned systems, distributed control systems will become an effective means of swarm control. However, USV swarms in real-world mission environments often consist of multiple heterogeneous USVs undertaking different tasks. Furthermore, during mission execution, USV swarms may encounter static or dynamic obstacles such as other vessels unrelated to the mission. In addition, most current multi-USV swarm algorithms integrate collision avoidance rewards into the main task, training the same policy network, which often leads to difficulties in policy training convergence and problems such as unclear priorities during USV operations. Summary of the Invention

[0007] This invention provides a method and related equipment for controlling unmanned surface vessel (USV) swarms under conditions of limited perception and communication. Its purpose is to solve the problems of limited communication, poor stability and robustness of collision avoidance control in heterogeneous USV swarms during the execution of encirclement and capture missions.

[0008] To achieve the above objectives, this invention provides a method for controlling an unmanned surface vessel (USV) swarm under conditions of limited perception and communication. The USV swarm includes multiple pursuit-type USVs, and the USV swarm control method includes: Step 1: Obtain the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel swarm; Step 2: Based on the local state information of each pursuing unmanned surface vessel, determine the current situation of each pursuing unmanned surface vessel and obtain the judgment result; Step 3: Based on the judgment result, control each pursuit-type unmanned surface vessel to surround or attack the target.

[0009] Furthermore, step 1 includes: For each pursuing unmanned surface vessel (USV) in the USV swarm, the absolute speed, absolute position, and collision radius of the pursuing USV are obtained through sensors on the pursuing USV itself. The absolute speed, absolute position, and collision radius of other pursuing unmanned surface vessels (USVs) in the vicinity are obtained through the communication equipment on the pursuing USV. The relative position, relative speed, and collision radius of the target and obstacles within the radar's perception range are obtained by the radar on the pursuit-type unmanned surface vessel. The absolute speed, absolute position, and collision radius of the pursuing unmanned surface vessel (USV), the absolute speed, absolute position, and collision radius of other pursuing USVs near the pursuing USV, and the relative position, relative speed, and collision radius of the target and obstacles within the radar perception range are used as local state information.

[0010] Furthermore, step 2 includes: Based on the relative position, relative speed, and collision radius of other pursuing unmanned surface vessels (USVs) near the pursuing USV, as well as the relative position, relative speed, and collision radius of obstacles within the radar sensing range, it is determined whether each pursuing USV is in the danger zone, and a first judgment result is obtained. Based on the relative positions of other pursuit unmanned surface vessels (USVs) near the pursuit USV and the relative position of the target being pursued, a second judgment result is obtained by determining whether the target is within the encirclement.

[0011] Furthermore, the center of the danger zone is the center of the obstacle, and the radius of the danger zone is 1.25 times the collision radius of the obstacle.

[0012] Furthermore, step 3 includes: When the first judgment result indicates that the pursuit-type unmanned surface vessel is in the danger zone, control the pursuit-type unmanned surface vessel to move away from the danger zone; When the first judgment result is that the pursuit unmanned surface vessel is not in the danger zone, and the second judgment result is that the target of the encirclement is not in the encirclement, the adaptive speed obstacle area and the expected collision time are calculated based on the relative position, relative speed, collision radius of each pursuit unmanned surface vessel, the relative position, relative speed, and collision radius of obstacles within the radar perception range. Based on the relative position, relative speed, collision radius, relative position, relative speed, collision radius, adaptive speed obstacle area, and expected collision time of each pursuing unmanned surface vessel, an encirclement reward function is constructed. The control command corresponding to the maximum function value of the encirclement reward function is used to control each pursuit-type unmanned surface vessel to encircle the target. When the first judgment result is that the pursuing unmanned surface vessel is not in the danger zone, and the second judgment result is that the target is in the encirclement, the encirclement reward function is constructed based on the relative position, relative speed, collision radius of each pursuing unmanned surface vessel, and the relative position, relative speed, and collision radius of obstacles within the radar perception range. The control commands corresponding to the maximum function value of the besiege reward function are used to control each pursuit-type unmanned surface vessel to besiege the target.

[0013] Furthermore, the calculation expression for the adaptive velocity obstacle region is: ; in, Indicates the adaptive speed obstacle region. Indicates the actual speed obstacle area. Indicates the first A pursuit-type unmanned surface vessel, This indicates obstacles detected by the pursuit-type unmanned surface vessel. This indicates the absolute speed of the pursuit-type unmanned surface vessel. This indicates the speed within the adaptive speed obstacle zone.

[0014] Furthermore, the encirclement reward function is: ; in, Indicates the first The encirclement reward obtained by the pursuit-type unmanned surface vessel. This indicates that the surrounding area forms a reward. Indicates a collision penalty. Indicates the first The sum of obstacle avoidance rewards between a pursuit-type unmanned surface vessel and all obstacles within its perception range. This indicates the number of obstacles detected by the pursuit-type unmanned surface vessel.

[0015] Furthermore, the siege reward function is: ; in, This represents the siege reward function value. Indicates distance penalty, Indicates punishment for attack. This indicates that the encirclement will maintain the reward. This indicates an obstacle or punishment.

[0016] This invention also provides a swarm control device for unmanned surface vessels (USVs) under conditions of limited perception and communication. The USV swarm includes multiple pursuit-type USVs, and the USV swarm control device includes: The acquisition module is used to acquire the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel cluster; The judgment module is used to determine the current situation of each pursuing unmanned surface vessel based on the local state information of each pursuing unmanned surface vessel, and obtain the judgment result; The control module is used to control each pursuit-type unmanned surface vessel to surround or attack the target based on the judgment result.

[0017] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for controlling unmanned surface vessels swarms under conditions of limited perception and communication.

[0018] The above-described solution of the present invention has the following beneficial effects: This invention acquires the local state information of each pursuing UAV in an unmanned surface vessel (USV) swarm; based on the local state information of each pursuing USV, it determines the current situation of each pursuing USV and obtains a judgment result; based on the judgment result, it controls each pursuing USV to surround or attack the target. Compared with the prior art, this invention controls each pursuing USV to surround or attack the target based on the current situation of each pursuing USV, avoiding excessive pursuit behavior of USVs before the encirclement is formed, improving the obstacle avoidance capability of USVs, and attacking the target after the encirclement is formed, thus increasing the success rate of the encirclement. It does not rely on water environment parameters, high-quality communication, or unified decision-making by a host computer, and can be applied to complex water environments. It solves the problems of limited communication, poor stability and robustness of collision avoidance control in heterogeneous USV swarms during the execution of encirclement missions.

[0019] Other beneficial effects of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating an embodiment of the present invention; Figure 2 This is a schematic diagram of the enclosing circle in an embodiment of the present invention; Figure 3 This is a schematic diagram of the absolute collision velocity region in an embodiment of the present invention; Figure 4 This is a comparison chart of the average collision avoidance success rate in embodiments of the present invention; Figure 5 This is a schematic diagram of the unmanned surface vessel swarm control device in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of the unmanned surface vessel cluster control device in an embodiment of the present invention. Detailed Implementation

[0021] To make the technical problems, solutions, and advantages of this invention clearer, a detailed description will be provided below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0023] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0024] This invention addresses existing problems by providing a method and related equipment for controlling unmanned surface vessels (USVs) swarms under conditions of limited perception and communication.

[0025] like Figure 1 As shown, embodiments of the present invention provide a method for controlling an unmanned surface vessel (USV) swarm under conditions of limited perception and communication. The USV swarm includes multiple pursuit-type USVs, and the USV swarm control method includes: Step 1: Obtain the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel swarm; Step 2: Based on the local state information of each pursuing unmanned surface vessel, determine the current situation of each pursuing unmanned surface vessel and obtain the judgment result; Step 3: Based on the judgment result, control each pursuit-type unmanned surface vessel to surround or attack the target.

[0026] Specifically, step 1 includes: For each pursuing unmanned surface vessel (USV) in the USV swarm, the absolute speed, absolute position, and collision radius of the pursuing USV are obtained through sensors on the pursuing USV. The absolute speed, absolute position, and collision radius of other pursuing unmanned surface vessels (USVs) in the vicinity are obtained through the communication equipment on the pursuing USV. The relative position, relative speed, and collision radius of the target and obstacles within the radar's perception range are obtained by the radar on the pursuit-type unmanned surface vessel. The absolute speed, absolute position, and collision radius of the pursuing unmanned surface vessel (USV), the relative position, relative speed, and collision radius of the target and obstacles within the radar's sensing range, are taken as local state information. This local state information can be represented as follows: ; in, Indicates the first The absolute speed of the pursuit-type unmanned surface vessel itself. Indicates the first The absolute position of the pursuit-type unmanned surface vessel itself. Indicates the first The collision radius of a pursuit-type unmanned surface vessel. This indicates that other nearby pursuit unmanned surface vessels and the first A set of relative positions of several pursuit-type unmanned surface vessels. This indicates that other nearby pursuit unmanned surface vessels and the first The set of relative velocities of a pursuit-type unmanned surface vessel. This represents the set of collision radii of other nearby pursuing unmanned surface vessels. Indicates the first The relative positions of obstacles within the perception range of a pursuit-type unmanned surface vessel. Indicates the first The relative speed of obstacles within the perception range of a pursuit-type unmanned surface vessel. Indicates the first A set of collision radii for obstacles within the perception range of a pursuit-type unmanned surface vessel. Indicates the target of the encirclement and capture and the first The relative positions of the pursuit-type unmanned surface vessel. Indicates the target of the encirclement and capture and the first The relative speed of the pursuing unmanned surface vessel. This indicates the collision radius of the target being pursued.

[0027] In this embodiment of the invention, the sensors on the pursuit-type unmanned surface vessel (USV) can be a global positioning system (GPS) and an inertial measurement unit (IMU), or other common status data acquisition devices. The collision radius of the pursuit-type USV itself is fixed data, which can be written into the USV at the time of manufacture. The radar mentioned in this embodiment of the invention refers to lidar, which is used to sense the absolute speed, absolute position, and collision radius of other USVs near the pursuit-type USV within a limited range. In this embodiment of the invention, the absolute speed, absolute position, and collision radius of other USVs near the pursuit-type USV can also be obtained through a communication module with limited performance.

[0028] It should be noted that encirclement and capture missions typically involve multiple pursuit drones and one target. The pursuit drones need to use their numerical advantage to surround the target and prevent it from escaping through the gaps between them.

[0029] In this embodiment of the invention, the pursuing unmanned surface vessel (USV), other pursuing USVs adjacent to the pursuing USV, the target being pursued, and obstacles within the perception range are all located on the same two-dimensional plane, and their absolute positions are represented by rectangular coordinates, as shown in the following formula: ; in, Indicates absolute position. Indicates the first The coordinates of the unmanned surface vessel in the east-west direction are used in this embodiment of the invention, with due east as the positive direction. Indicates the first The coordinates of the unmanned surface vessel in the north-south direction are taken as positive, with due north as the positive direction.

[0030] Specifically, step 2 includes: Based on the relative position, relative speed, and collision radius of other pursuing unmanned surface vessels (USVs) near the pursuing USV, as well as the relative position, relative speed, and collision radius of obstacles within the radar sensing range, it is determined whether each pursuing USV is in the danger zone, and a first judgment result is obtained. Based on the relative positions of other pursuit unmanned surface vessels (USVs) near the pursuit USV and the relative position of the target being pursued, a second judgment result is obtained by determining whether the target is within the encirclement.

[0031] It should be noted that the center of the danger zone is the center of the obstacle, and the radius of the danger zone is 1.25 times the collision radius of the obstacle; the encirclement is defined as a polygon formed by the pursuing unmanned surface vessel and other nearby pursuing unmanned surface vessels as vertices, such as... Figure 2 As shown.

[0032] Specifically, step 3 includes: When the first judgment result indicates that the pursuit-type unmanned surface vessel is in the danger zone, control the pursuit-type unmanned surface vessel to move away from the danger zone; When the first judgment result is that the pursuit unmanned surface vessel is not in the danger zone, and the second judgment result is that the target of the encirclement is not in the encirclement, the adaptive speed obstacle area and the expected collision time are calculated based on the relative position, relative speed, collision radius of each pursuit unmanned surface vessel, the relative position, relative speed, and collision radius of obstacles within the radar perception range. Based on the relative position, relative speed, collision radius, relative position, relative speed, collision radius, adaptive speed obstacle area, and expected collision time of each pursuing unmanned surface vessel, an encirclement reward function is constructed. The control command corresponding to the maximum function value of the encirclement reward function is used to control each pursuit-type unmanned surface vessel to encircle the target. When the first judgment result is that the pursuing unmanned surface vessel is not in the danger zone, and the second judgment result is that the target is in the encirclement, the encirclement reward function is constructed based on the relative position, relative speed, collision radius of each pursuing unmanned surface vessel, and the relative position, relative speed, and collision radius of obstacles within the radar perception range. The control commands corresponding to the maximum function value of the besiege reward function are used to control each pursuit-type unmanned surface vessel to besiege the target.

[0033] Specifically, the calculation expression for the adaptive velocity obstacle region is: ; in, Indicates the adaptive speed obstacle region. Indicates the actual speed obstacle area. Indicates the first A pursuit-type unmanned surface vessel, This refers to obstacles sensed by the pursuing unmanned surface vessel (USV), including other nearby pursuing USVs and other static and dynamic obstacles. This indicates the absolute speed of the pursuit-type unmanned surface vessel. This indicates the speed within the adaptive speed obstacle zone.

[0034] Specifically, the derivation process of the adaptive velocity obstacle region calculation expression is as follows: First, the speed obstacle region is constructed using the speed obstacle method, expressed as: ; in, Indicates the speed obstacle area. Indicates the number of the obstacle. Indicates the current position of the pursuit-type unmanned surface vessel. This indicates the pursuit-type unmanned surface vessel and the obstacles it detects. The relative velocity between them can be obtained from the set of relative velocities. and get, Indicates the current position from the pursuit-type unmanned surface vessel, along... Rays in direction, This indicates obstacles detected by the pursuit-type unmanned surface vessel. This represents the Minkowski vector sum operation. It is represented as a circle on a two-dimensional plane with radius ER as the sum of the collision radii of the pursuing unmanned surface vessel and the obstacle; Since the sector area of ​​the velocity obstacle region is one of the main factors affecting velocity selection in the velocity obstacle method, and the radius ER is the main factor affecting the area, the adaptive expansion radius AER is used instead of the radius ER. The radius collision coefficient REC and the adaptive expansion radius AER are shown in the following formulas: ; ; in, Indicates the first A pursuit-type unmanned surface vessel and obstacles The Euclidean distance between them can be determined by... and Calculations show that express The collision radius of a pursuit-type unmanned surface vessel. Indicates obstacles The collision radius; After introducing the radius collision coefficient REC, the adaptive expansion radius AER is used instead of the radius ER, thus improving the velocity obstacle method into the Adaptive Velocity Obstacles (AVO) method, which yields the actual velocity obstacle area. ; Finally, to address the motion oscillation problem, this embodiment of the invention improves the adaptive velocity obstacle method into the adaptive reciprocal velocity obstacle method (ARVO), where both parties share half of the obstacle avoidance responsibility. This leads to the derivation of the adaptive velocity obstacle region calculation expression: ; in, This represents the speed within the adaptive speed obstacle zone, i.e., the speed within the absolute collision speed zone, such as... Figure 3 As shown.

[0035] However, it should be noted that real-world environments often contain multiple obstacles; therefore, adaptive speed obstacle zones are necessary. It is the superposition of all absolute collision velocity regions.

[0036] Specifically, the expected collision time is calculated from the relative position and relative velocity, as shown in the following expression: ; in, Indicates the expected collision time. Represents relative velocity, taken from the corresponding set of relative velocities. and , Represents relative position, taken from the corresponding set of relative positions. and .

[0037] Specifically, the encirclement reward function is: ; in, Indicates the first The encirclement reward obtained by the pursuit-type unmanned surface vessel. This indicates that the surrounding area forms a reward. ,when When the value is 0, it indicates that the target has not been surrounded or captured. A value of 10 indicates that the target has been surrounded and captured. Indicates a collision penalty. ,when A value of 0 indicates that no collision occurred. A value of -10 indicates that a collision has occurred. Indicates the first The sum of obstacle avoidance rewards between a pursuit-type unmanned surface vessel and all obstacles within its perception range. , It is a negative number, used as a direction penalty coefficient. This represents the average distance between all edges of the encirclement and the target. A positive number represents the base reward when the collision risk is low. A negative number represents the collision time penalty coefficient. When the current speed of the pursuing unmanned surface vessel is within the ACVZ range, it will be subject to a collision time penalty. When the expected collision time... When the time is less than 0.1s, the collision time penalty coefficient will revert to its original value. The base reward will be multiplied; at the same time, the unmanned surface vessel will lose its base reward. , This indicates the number of obstacles detected by the pursuit-type unmanned surface vessel.

[0038] In this embodiment of the invention, the control instruction corresponding to the maximum function value surrounding the reward function depends on the policy network. Influenced by the training level of the policy network, the policy network will select the control action corresponding to the currently explored known maximum function value. The specific process is as follows: Step i: Initialize the policy network and evaluation network of the encirclement algorithm; Step ii, will As an external sensing input, a bidirectional gated recurrent network is used to obtain the final state sum, which is then combined into a fixed-length vector. This vector is concatenated with the state information of the unmanned surface vessel and used as the current state of the pursuit type. This state will then be used as the input of the neural network. Step iii: Use layer normalization to transform the input data into normalized data, thereby improving the stability of training; Step iv: The normalized input data is input into the policy network. The policy network generates and executes control actions, the environment undergoes a state transition, and the pursuing unmanned surface vessel obtains its own state information and external sensory quantities at the next moment and records the encirclement reward obtained. The normalized input data is then fed into the evaluation network to calculate the corresponding state value. The pursuit-type unmanned surface vessel saves experience data to the experience pool in the following format: ,in, This represents the policy network parameters of the current encirclement algorithm. Next, sampling selection action The logarithmic probability; Step v: The pursuit-type unmanned surface vessel continues to explore, record, and save experience data; Step vi: After each round, update the policy network and evaluation network using all the empirical data of the pursuit unmanned surface vessels. Each piece of empirical data is used fifty times and then discarded. After the update is completed, enter a new round until the network parameters of the policy network and the evaluation network reach their optimal values.

[0039] It should be noted that the policy network of the encirclement algorithm consists of two fully connected layers with 256 neurons each, using ReLU as the activation function, and outputting a two-dimensional normal distribution mean of acceleration, activated by the hyperbolic tangent function. It also includes trainable parameters defined using torch.nn.Parameter to represent the standard deviation of the normal distribution. The evaluation network of the encirclement algorithm also consists of two fully connected layers with 256 neurons each, using ReLU as the activation function, and outputting a one-dimensional state value.

[0040] Specifically, the siege reward function is: ; in, This represents the siege reward function value, used to characterize the average reward of all pursuing unmanned surface vessels (USVs) in the USV swarm. Indicates distance penalty, Indicates punishment for attack. This indicates that the encirclement will maintain the reward. This indicates an obstacle or punishment.

[0041] Specifically, when the pursuit-type unmanned surface vessel (USV) is far from the target, it will be subject to a significant distance penalty, prompting it to move closer to the target. The calculation expression for the distance penalty for each pursuit-type USV is as follows: ; in, This represents the distance penalty coefficient. This represents the Euclidean distance between the pursuit-type unmanned surface vessel and the target being pursued, with a distance penalty. It is a negative value, and its absolute value decreases as the pursuer gets closer to the escapee; The attack penalty is designed to incentivize the pursuit UAVs to attack the encirclement target, encouraging them to attack only after surrounding the target. The calculation expression for the attack penalty for each pursuit UAV is as follows: ; in, This represents the attack penalty coefficient, indicating the penalty incurred by the pursuing unmanned surface vessel for each collision with the target. This indicates the penalty coefficient for encirclement attacks, encouraging pursuit-type unmanned surface vessels to attack only after encircling and capturing the target. The penalty for attacking is... for At that time, that is, colliding within the encirclement, and the attack is punished. for At that time, it crashes outside the encirclement, and the attack is punished. A value of 0 indicates that no collision occurred; The encirclement maintenance reward is used to incentivize pursuit drones to maintain an encirclement of the target by leveraging their numerical advantage. Its calculation formula is as follows: ; in, This indicates that the reward coefficient is maintained within the encirclement. This indicates that the encirclement is close to the reward coefficient. This represents the distance between the target and the encirclement, which is the straight-line distance from the target to the nearest edge of the encirclement. The reward is given when the encirclement is maintained. for When the target is within the encirclement, a reward is given for maintaining the encirclement. for When the target is outside the encirclement, it indicates that the target is outside the encirclement. The formula for calculating obstacle penalty is: ; in, This represents the obstacle penalty coefficient, indicating the severity of the penalty incurred by a pursuing unmanned surface vessel (USV) when approaching an obstacle. This indicates the distance between the obstacle and the pursuing unmanned surface vessel (USV), and the penalty for obstacles. for When the unmanned surface vessel is in a danger zone, an obstacle will trigger a penalty. A value of 0 indicates that the unmanned surface vessel is outside the danger zone.

[0042] In this embodiment of the invention, the control instruction corresponding to the maximum function value of the siege reward function also depends on the policy network. Influenced by the training level of the policy network, the policy network will select the control action corresponding to the currently explored known maximum function value. The specific process is as follows: Step i: Initialize the policy network, evaluation network, target policy network, and target evaluation network of the siege algorithm. The parameters of the target policy network and target evaluation network of the siege algorithm are consistent with those of the initial policy network and evaluation network of the siege algorithm. Step ii: The basic local state information of each pursuing unmanned surface vessel is input into the policy network of its respective besieging algorithm. The policy network of the besieging algorithm outputs a definite instruction. Then, two-dimensional random noise that conforms to a normal distribution is added to increase the exploration capability of the reinforcement learning algorithm. Step iii: Based on the maneuverability of the pursuit-type unmanned surface vessel (USV), the determined command is converted into the actual action executed by the USV. The actual action executed by the USV is limited by the upper limit of acceleration, as shown in the following formula: ; in, The actual actions performed by the pursuit-type unmanned surface vessel. The upper limit of acceleration for pursuit-type unmanned surface vessels, This represents an action superimposed with two-dimensional random noise; Step iv: All pursuing unmanned surface vessels (USVs) interact with the environment, causing state transitions and obtaining their respective rewards and basic local state information for the next state. The experience data of each pursuing USV is then stored. Add to the experience pool; Step v: After the experience pool is full, randomly extract all experience data from the experience pool at the same time. All Concatenate them into global state information S, and then concatenate S and The input is fed into the evaluation network of the siege algorithm to obtain the evaluation network's assessment of the actions in the current state. Rating ; Step vi, will The input is fed into the target policy network of the siege algorithm to obtain the action at the next time step. ; The global state information at the next moment and its actions The input is fed into the target evaluation network of the siege algorithm to obtain the target evaluation network's evaluation of the action. Rating The TD-error is calculated as shown in the following formula: ; in, For TD-error, Reward retention rate represents the current value of future rewards; Step vii, will As the loss function for evaluating the network, gradient descent is used to update the evaluation network of the Siege algorithm, adjusting all parameters and changing its output score. This makes the TD-error as close to zero as possible, thereby minimizing the loss function; Step viii involves updating the policy network of the siege algorithm, adjusting all parameters so that the actions derived from the current basic local state information can achieve greater success in the evaluation network of the siege algorithm. Using the policy network of the siege algorithm, the output action in the new state is selected, and the gradients of all parameters in the policy network with respect to the action and the gradients of the action with respect to the evaluation network of the siege algorithm are calculated. The gradient of the parameter can be used to estimate the value of the parameter pair according to the chain rule. The gradient.

[0043] It should be noted that the target policy network and target evaluation network of the siege algorithm are soft-updated after each round of training, and the policy network and evaluation network of the siege algorithm have the same architecture, both consisting of four fully connected layers.

[0044] The unmanned surface vessel (USV) swarm control method provided in this embodiment of the invention and the USV formation encirclement control method using the interactive speed obstacle method are used to control a pursuit-type USV. The average collision avoidance success rate of the two methods when performing encirclement missions is as follows: Figure 4 As shown, by Figure 4 It is evident that the unmanned surface vessel (USV) swarm control method under conditions of limited perception and communication in the embodiments of the present invention can effectively improve the collision avoidance capability of heterogeneous USV swarms and enhance their safe navigation capability.

[0045] This invention acquires the local state information of each pursuing unmanned surface vessel (USV) in a USV swarm; based on the local state information of each pursuing USV, it determines the current situation of each pursuing USV and obtains a judgment result; based on the judgment result, it controls each pursuing USV to surround or attack the target. Compared with the prior art, this invention controls each pursuing USV to surround or attack the target based on the current situation of each pursuing USV, avoiding excessive pursuit behavior of USVs before the encirclement is formed, improving the obstacle avoidance capability of USVs, and attacking the target after the encirclement is formed, thus increasing the success rate of the encirclement. It does not rely on water environment parameters, high-quality communication, or unified decision-making by a host computer, and can be applied to complex water environments. It solves the problems of limited communication, poor stability and robustness of collision avoidance control in heterogeneous USV swarms during the execution of encirclement missions.

[0046] Corresponding to the unmanned surface vessel swarm control method under perception and communication constraints described in the above embodiments, such as Figure 5 As shown, this embodiment of the invention also provides an unmanned surface vessel (USV) swarm control device 100 under conditions of limited perception and communication. The USV swarm includes multiple pursuit-type USVs, and the USV swarm control device 100 includes: The acquisition module 101 is used to acquire the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel cluster; The judgment module 102 is used to judge the current situation of each pursuing unmanned surface vessel based on the local state information of each pursuing unmanned surface vessel, and obtain the judgment result; The control module 103 is used to control each pursuit-type unmanned surface vessel to surround or attack the target based on the judgment result.

[0047] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0048] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0049] This invention also provides an unmanned surface vessel swarm control device, such as... Figure 6 As shown, the unmanned surface vessel (USV) swarm control device in this embodiment includes a terminal device D10, a communication device D11, and multiple remote execution devices D12. Figure 6 Only one remote execution device is shown in the diagram), wherein the terminal device D10 includes: at least one processor D100 ( Figure 6 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the above-described unmanned surface vessel swarm control method under conditions of limited perception and communication.

[0050] Each remote execution device D12 includes: a shipborne controller D120, an inertial sensor D121, a satellite locator D122, a radar D123, an electronic speed controller D124, a left motor D125, and a right motor D126. It should be noted that the remote execution device D12 is a pursuit-type unmanned surface vessel. It achieves fusion positioning through the satellite locator D122 and the inertial sensor D121 to obtain its own position information and motion state. Then, it obtains the relative position, collision radius, and relative speed of other nearby pursuit-type unmanned surface vessels and obstacles through the radar D123, thereby constructing its own basic local state information.

[0051] The terminal device D10 can be a desktop computer, laptop, handheld computer, server, server cluster, or cloud server, etc. This terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art will understand that... Figure 3This is merely an example of terminal device D10 and does not constitute a limitation on terminal device D10. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0052] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0053] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0054] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0055] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0056] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for controlling unmanned surface vessel swarms under conditions of limited perception and communication, characterized in that, The unmanned surface vessel (USV) swarm comprises multiple pursuit-type USVs, and the USV swarm control method includes: Step 1: Obtain the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel cluster; Step 2: Based on the local state information of each of the pursuing unmanned surface vessels (USVs), determine the current situation of each USV and obtain the determination result. Step 3: Based on the judgment result, control each pursuit-type unmanned surface vessel to surround or attack the target.

2. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 1, characterized in that, Step 1 includes: For each pursuing unmanned surface vessel (USV) in the USV swarm, the absolute speed, absolute position, and collision radius of the pursuing USV are obtained through sensors on the pursuing USV. The absolute speed, absolute position, and collision radius of other pursuing unmanned surface vessels (USVs) near the pursuing USV are obtained through the communication equipment on the pursuing USV. The relative position, relative speed, and collision radius of the target and obstacles within the radar's sensing range are obtained through the radar on the pursuit-type unmanned surface vessel. The absolute speed, absolute position, and collision radius of the pursuing unmanned surface vessel (USV), the absolute speed, absolute position, and collision radius of other pursuing USVs adjacent to the USV, and the relative position, relative speed, and collision radius of the target and obstacles within the radar's sensing range are used as local state information.

3. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 2, characterized in that, Step 2 includes: Based on the relative position, relative speed, and collision radius of other pursuing unmanned surface vessels (USVs) adjacent to the pursuing USV, as well as the relative position, relative speed, and collision radius of obstacles within the radar sensing range, it is determined whether each pursuing USV is within the danger zone, and a first judgment result is obtained. Based on the relative positions of other pursuit unmanned surface vessels (USVs) adjacent to the pursuit USV and the relative position of the target being pursued, it is determined whether the target being pursued is within the encirclement, thus obtaining a second determination result.

4. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 3, characterized in that, The center of the danger zone is the center of the obstacle, and the radius of the danger zone is 1.25 times the collision radius of the obstacle.

5. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 3, characterized in that, Step 3 includes: When the first determination result indicates that the pursuit-type unmanned surface vessel is within a dangerous area, the pursuit-type unmanned surface vessel is controlled to leave the dangerous area; When the first judgment result is that the pursuing unmanned surface vessel is not in the danger zone, and the second judgment result is that the target to be captured is not in the encirclement, the adaptive speed obstacle area and the expected collision time are calculated based on the relative position, relative speed, collision radius of each pursuing unmanned surface vessel, the relative position, relative speed, and collision radius of obstacles within the radar perception range. A bounding reward function is constructed based on the relative position, relative speed, collision radius of each pursuing unmanned surface vessel, the relative position, relative speed, collision radius of obstacles within the radar perception range, the adaptive speed obstacle area, and the expected collision time. Each of the pursuing unmanned surface vessels is controlled to surround the target using a control command corresponding to the maximum function value of the encirclement reward function. When the first judgment result is that the pursuing unmanned surface vessel is not in the danger zone, and the second judgment result is that the target to be captured is within the encirclement, an encirclement reward function is constructed based on the relative position, relative speed, collision radius of each pursuing unmanned surface vessel, and the relative position, relative speed, and collision radius of obstacles within the radar perception range; Each of the pursuing unmanned surface vessels is controlled to attack the target by using control commands corresponding to the maximum function value of the attack reward function.

6. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 5, characterized in that, The calculation expression for the adaptive velocity obstacle region is: ; in, Indicates the adaptive speed obstacle region. Indicates the actual speed obstacle area. Indicates the first A pursuit-type unmanned surface vessel, This indicates obstacles detected by the pursuit-type unmanned surface vessel. This indicates the absolute speed of the pursuit-type unmanned surface vessel. This indicates the speed within the adaptive speed obstacle zone.

7. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 5, characterized in that, The encirclement reward function is: ; in, Indicates the first The encirclement reward obtained by the pursuit-type unmanned surface vessel. This indicates that the surrounding area forms a reward. Indicates a collision penalty. Indicates the first The sum of obstacle avoidance rewards between a pursuit-type unmanned surface vessel and all obstacles within its perception range. This indicates the number of obstacles detected by the pursuit-type unmanned surface vessel.

8. The unmanned surface vessel swarm control method under perception and communication constraints according to claim 5, characterized in that, The siege reward function is: ; in, This represents the siege reward function value. Indicates distance penalty, Indicates punishment for attack. This indicates that the encirclement will maintain the reward. This indicates an obstacle or punishment.

9. A swarm control device for unmanned surface vessels under conditions of limited perception and communication, characterized in that, The unmanned surface vessel (USV) swarm comprises multiple pursuit-type USVs, and the USV swarm control device includes: The acquisition module is used to acquire the local state information of each pursuit-type unmanned surface vessel in the unmanned surface vessel cluster; The judgment module is used to determine the current situation of each of the pursuing unmanned surface vessels based on the local state information of each vessel, and to obtain the judgment result. The control module is used to control each pursuit-type unmanned surface vessel to surround or attack the target based on the judgment result.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the unmanned surface vessel swarm control method under perception and communication constraints as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-boat two-layer cooperative automatic control contraction pursuing method

    CN114489080A

  • Multi-unmanned ship cooperative hunting method, computer equipment and storage medium

    CN118092447A

  • Multi-unmanned ship cooperative hunting training method based on bidirectional deep reinforcement learning

    CN118626867A

  • Cooperative hunting control method for unmanned surface vehicle

    CN121349077A