Unmanned ship intelligent hunting control method, device and equipment
Through the hierarchical policy network and the intelligent round-up control method of unmanned boats optimized by reinforcement learning, the problems of single coordination mechanism, insufficient state input and weak strategy generalization capabilities in the intelligent round-up control of unmanned boats are solved, and efficient round-up in complex maritime environments are achieved.
Patent Information
- Application Number
- CN202510756891.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the existing intelligent round-up control technology of unmanned boats, there are problems such as lack of flexible role division of the coordination mechanism, inability to adapt to dynamic entity changes, weak strategy generalization ability and sparse reward functions, resulting in a low round-up success rate.
The hierarchical policy network is used to combine reinforcement learning, and a local coordinate system is built by obtaining unmanned boat state information and local environmental data, a structured state input vector is constructed using distance sorting and filling and cutting mechanisms, a navigation speed and direction control instruction is generated, a structured reward mechanism is designed to optimize the policy network, and the feature fusion of multi-source state information and task role and behavior decisions are realized.
The stability and flexibility of multi-boat collaboration strategy have been improved, and the success rate of roundup missions in complex marine dynamic target environments has been significantly improved.
Smart Images

Figure CN120276451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control of unmanned boats, and particularly to an intelligent encirclement control method, device and equipment for unmanned boats. Background Technique
[0002] In recent years, the technology of unmanned surface vehicles (USVs) has been gradually studied in the fields of marine dynamic target monitoring and intelligent interception tasks. Especially in the field of multi-boat encirclement control, through the division of labor and cooperation among multiple boats, unmanned boats can efficiently encircle and suppress dynamic targets. However, the existing intelligent encirclement control technology for unmanned boats still faces the following challenges: First, the cooperation mechanism still mainly relies on static formations or centralized control rules, lacking a flexible role division mechanism for dynamic target tasks; Second, the input state design is generally simplified, and the target behavior, environmental disturbances and multi-boat situation information are not effectively integrated, resulting in poor adaptability of the strategy to environmental changes; Third, the modeling of the functional differences of each role in the encirclement process is insufficient, and the strategy generalization ability is weak, making it difficult to deal with targets with different numbers and behavior patterns; Fourth, during the training process, the reward function often has sparse problems and insufficient cooperative guidance, which restricts the formation of high-quality strategies and affects the success rate of the encirclement task. Summary of the Invention
[0003] The present invention provides an intelligent encirclement control method, device and equipment for unmanned boats, which solves the problems existing in the existing intelligent encirclement control method for unmanned boats, such as single role allocation of the encirclement strategy, inability of the state input to adapt to dynamic entity changes, and low success rate of encirclement.
[0004] To solve the above technical problems, the technical solution of the present invention is as follows:
[0005] An embodiment of the present invention proposes an intelligent encirclement control method for unmanned boats, including:
[0006] Obtain the state information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to represent the state information, so as to obtain the sensed state information corresponding to the current unmanned boat;
[0007] Based on the sensed state information, construct a structured state input vector by using a distance sorting and filling and cropping mechanism;
[0008] Input the structured state input vector into a hierarchical policy network, and the high-level of the hierarchical policy network generates task roles and behavior decisions, and the low-level combines the high-level output and the environmental state in the structured state input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism using reinforcement learning;
[0009] Control the current unmanned boat to perform a surrounding mission on the target boat according to the sailing speed and direction control instruction.
[0010] Optionally, obtain the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to represent the status information, so as to obtain the sensed status information corresponding to the current unmanned boat, including:
[0011] Obtain the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat;
[0012] Establish a local coordinate system with the position of the target boat as the origin, and uniformly represent the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat in the local coordinate system, so as to obtain the sensed status information corresponding to the current unmanned boat; wherein, the local environment data includes: the status of the target boat, the status of the cooperative unmanned boat, the status of the obstacle, and the status of the marine environment.
[0013] Optionally, through a reward mechanism, use reinforcement learning to optimize a preset policy network to obtain a hierarchical policy network, including:
[0014] Construct a typical task scenario in the unmanned boat surrounding training environment and generate training status data;
[0015] Based on the training status data, in view of the dynamic changes in the number of cooperative unmanned boats and obstacles, use a distance sorting, filling and cropping mechanism to organize the status information sensed by the current unmanned boat into an input vector in a unified format;
[0016] Input the input vector into the input layer of the preset policy network, and perform feature fusion on the input vector through a multi-head attention layer to generate a fused representation of multi-source states;
[0017] The decision-making layer of the preset policy network generates task roles and behavior decisions based on the output of the multi-head attention layer, and outputs high-level discrete action instructions;
[0018] The basic layer of the policy network combines the output of the decision-making layer and the output of the multi-head attention layer to generate continuous control instructions;
[0019] Execute the continuous control instructions and interact with the unmanned boat surrounding training environment to update the unmanned boat status;
[0020] Based on a structured reward function, calculate the immediate reward according to the updated unmanned boat status and continuous control instructions, and optimize the preset policy network through a multi-agent proximal policy optimization algorithm to obtain a hierarchical policy network.
[0021] Optionally, construct a typical task scenario in the unmanned boat surrounding training environment and generate training status data, including:
[0022] In the training environment of unmanned boat round-up, a target boat, a current unmanned boat and multiple cooperative unmanned boats are set, and their interactions with each other and with the environment are simulated to construct a typical round-up task scenario;
[0023] Training state data is generated through the typical round-up task scenario, and the training state data includes the state data of the target boat, the current unmanned boat, the cooperative unmanned boats and the obstacles.
[0024] Optionally, based on the training state data, in response to the dynamic changes in the number of cooperative unmanned boats and obstacles, a distance sorting, padding and cropping mechanism is adopted to organize the state information sensed by the current unmanned boat into an input vector in a unified format, including:
[0025] Based on the training state data, the state objects sensed by the current unmanned boat in the training state data are sorted according to the relative distance from the current unmanned boat, and the state objects include variable numbers of cooperative unmanned boat state objects and obstacle state objects sensed by the current unmanned boat;
[0026] When the number of state objects exceeds the preset maximum value, cropping is performed, and when it does not reach, it is filled with zero vectors, and the sorted, cropped and filled state data is constructed into an input vector in a unified format.
[0027] Optionally, feature fusion processing is performed on various types of state information through a multi-head attention layer to generate a fused representation of multi-source states, including:
[0028] Embed various types of state information in the input vector in a unified format into a feature space of the same dimension to obtain an input matrix;
[0029] Based on the input matrix, calculate the outputs of multiple attention heads, and splice the results of each attention head and then perform a linear mapping to generate a fused representation of multi-source states.
[0030] Optionally, the structured reward function includes:
[0031] A striker role reward function for the striker unmanned boat to quickly approach the target boat and block it frontally; a flank role reward function for the flank unmanned boats to assist in surrounding from both sides and enhance deterrence cooperation; a defender role reward function for the defender unmanned boat to implement suppression from behind the target boat and deter and intervene in a high-threat state.
[0032] An embodiment of the present invention also provides an intelligent round-up control device for unmanned boats, including:
[0033] An acquisition module, configured to acquire the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to uniformly represent the status information, so as to obtain the sensed status information corresponding to the current unmanned boat;
[0034] A processing module, configured to construct a structured state input vector based on the sensed status information by using a distance sorting and filling and cropping mechanism;
[0035] A control module, configured to input the structured state input vector into a hierarchical policy network, where the high level of the hierarchical policy network generates task roles and behavior decisions, and the low level combines the high-level output and the environmental state in the structured state input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism and using reinforcement learning;
[0036] An execution module, configured to control the current unmanned boat to perform a pursuit task on the target boat according to the navigation speed and direction control instructions.
[0037] An embodiment of the present invention further provides a computing device, including: a processor and a memory, where the memory stores a computer program, and when the program runs on the processor, the computer is enabled to execute the above method.
[0038] An embodiment of the present invention further provides a computer-readable storage medium, storing instructions, and when the instructions run on a computer, the computer is enabled to execute the above method.
[0039] The above solution of the present invention has at least the following beneficial effects:
[0040] The intelligent pursuit control method of the unmanned boat of the present invention acquires the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat, constructs a local coordinate system based on the target boat to uniformly represent the status information, and obtains the sensed status information corresponding to the current unmanned boat; based on the sensed status information, constructs a structured state input vector by using a distance sorting and filling and cropping mechanism; inputs the structured state input vector into a hierarchical policy network, where the high level of the hierarchical policy network generates task roles and behavior decisions, and the low level combines the high-level output and the environmental state in the structured state input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism and using reinforcement learning; controls the current unmanned boat to perform a pursuit task on the target boat according to the navigation speed and direction control instructions. The present invention effectively improves the stability, flexibility and generalization ability of the multi-boat cooperation strategy, especially in a complex marine dynamic target environment, and significantly improves the success rate of the pursuit task. Description of the Drawings
[0041] Figure 1It is a schematic flowchart of the intelligent hunting control method for an unmanned boat of the present invention;
[0042] Figure 2 It is a schematic structural diagram of the hierarchical policy network in the intelligent hunting control method for an unmanned boat of the present invention;
[0043] Figure 3 It is a schematic flowchart of the training process of the policy network in the intelligent hunting control method for an unmanned boat of the present invention;
[0044] Figure 4 It is a schematic diagram of the role distribution and cooperation strategy in the unmanned boat hunting task of the intelligent hunting control method for an unmanned boat of the present invention;
[0045] Figure 5 It is a schematic module structure diagram of the intelligent hunting control device for an unmanned boat of the present invention. Detailed implementation manners
[0046] The exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various ways and should not be limited to the embodiments described herein. These embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the technical scope of the present invention to those skilled in the art.
[0047] As Figures 1 to 4 shown, an embodiment of the present invention provides an intelligent hunting control method for an unmanned boat, including:
[0048] Step 11: Obtain the state information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to represent the state information, so as to obtain the sensed state information corresponding to the current unmanned boat;
[0049] Step 12: Based on the sensed state information, construct a structured state input vector by using a distance sorting and filling and cropping mechanism;
[0050] Step 13: Input the structured state input vector into a hierarchical policy network. The high level of the hierarchical policy network generates task roles and behavior decisions, and the low level combines the high-level output and the environmental state in the structured state input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism and using reinforcement learning. Specifically, it is obtained by offline training of the preset policy network by using reinforcement learning with a reward mechanism designed in combination with role division and environmental state;
[0051] Step 14: Control the current unmanned boat to perform a hunting task on the target boat according to the navigation speed and direction control instructions.
[0052] In this embodiment, each unmanned boat completes decision-making and control based on the perceived local environmental information. All state information is uniformly represented in a local coordinate system with the position of the target boat as the origin, and is organized into a structured state input vector, which is input into a policy network with high and low-level decoupling. The policy network consists of a decision-making layer and a basic layer: among them, the decision-making layer generates high-level discrete action instructions according to the input state vector, including task role assignment (frontier, flank, rear guard) and behavior selection (no deterrence, non-lethal deterrence, lethal deterrence), which are used to guide the tactical mode. The task roles have different divisions of labor: the frontier role is responsible for quickly approaching the target and blocking its escape path frontally; the flank role is responsible for interfering from both sides and controlling the maneuvering space of the target; the rear guard role is responsible for suppressing from behind the target boat and undertaking deterrence and defense control tasks. The basic layer combines the environmental state and the output result of the decision-making layer to generate control instructions including the target speed and heading angle, which are used to guide the actual movement of the unmanned boat. The target speed is a normalized control value, and the range is , corresponding to the maximum reverse speed to the maximum forward speed; the target heading angle is The angle between, indicating the target direction in the local coordinate system; through the design of "policy layering + action decoupling", it supports the policy network to dynamically generate individual behaviors according to the local situation during task execution, improving the multi-boat cooperation efficiency, structural flexibility and control robustness during the encirclement process. In this embodiment, the intelligent encirclement control method of the unmanned boat is applicable to task scenarios such as offshore monitoring, rapid interception, and formation encirclement control for dynamic targets, and can significantly improve the intelligent cooperative control level of surface unmanned boat clusters in complex sea conditions;
[0053] In this embodiment, the structured state input vector is composed of the following five types of sub-vectors spliced together: the current state of the unmanned boat, the state of the target boat, the state of the cooperative unmanned boat, the state of the obstacle, and the state of the marine environment. To keep the input format consistent and adapt to the policy network, each state is first serialized into a unified vector, defined as: ; among them, Is the structured state input vector, Is the current state of the unmanned boat, Is the state of the target boat, ( )Is the state of the cooperative unmanned boat; ( )Is the state of the obstacle, Is the environmental state; And Respectively represent the observed number of cooperative unmanned boats and obstacles. The above two types of states are sorted in ascending order of the relative distance from the current unmanned boat to facilitate subsequent cropping and filling operations;
[0054] In an alternative embodiment of the present invention, step 11 may include:
[0055] Step 111: Obtain the status information of the current unmanned boat and the local environment data it senses.
[0056] Step 112: Establish a local coordinate system with the position of the target boat as the origin (eastward as the axis and northward as the axis). Represent the status information of the current unmanned boat and the local environment data it senses in this coordinate system uniformly to obtain the sensed status information corresponding to the current unmanned boat. Among them, the local environment data includes: the status of the target boat, the status of the cooperative unmanned boats, the status of obstacles, and the status of the marine environment (environmental status).
[0057] In this embodiment, by constructing a local coordinate system with the position of the target boat as the origin to uniformly represent multi-source status information, the spatial consistency of the input data can be enhanced; the specific types of status information include: the status vector of the current unmanned boat, where is the position of the current unmanned boat in the local coordinate system, is the speed of the current unmanned boat, is the heading angle of the current unmanned boat, is the mission role of the current unmanned boat, 0 represents the vanguard, 1 represents the flank, and 2 represents the rear guard is the behavior status of the current unmanned boat, 3 represents no deterrence behavior, 4 represents non-lethal deterrence, and 5 represents lethal deterrence; the status vector of the target boat, where is the speed of the target boat, is the heading angle of the target boat, is the behavior status of the target boat, 6 represents stopping the boat, 7 represents escaping, and 8 represents attacking; the status vector of the cooperative unmanned boats, where is the position of the th cooperative unmanned boat, is the speed of the th cooperative unmanned boat, is the heading angle of the th cooperative unmanned boat, is the mission role of the th cooperative unmanned boat, is the behavior status of the th cooperative unmanned boat; the status vector of the obstacles, where is the position of the center of the obstacle, is the speed of the obstacle, is the moving direction of the obstacle, is the radius of the equivalent circle of the obstacle; the status vector of the marine environment, where is the wind speed, is the wind direction, is the wave height, is the wave direction.
[0058] In an alternative embodiment of the present invention, through a structured reward mechanism, reinforcement learning is used to optimize a preset policy network to obtain a hierarchical policy network, including:
[0059] Step 131, construct a typical task scenario in the unmanned boat pursuit training environment and generate training state data; the training state data includes: state data of the target boat, the current unmanned boat, cooperative unmanned boats and obstacles, as well as environmental disturbance data; the typical task scenario includes one target boat, one current unmanned boat and multiple cooperative unmanned boats; the scenario includes one target boat, one current unmanned boat and multiple cooperative unmanned boats, and obstacle and environmental disturbance factors are set;
[0060] Step 132, based on the training state data, for the dynamic changes in the number of cooperative unmanned boats and obstacles, use a distance sorting, filling and cropping mechanism to organize the state information sensed by the current unmanned boat into an input vector in a unified format for subsequent network processing;
[0061] Step 133, input the input vector into the input layer of the preset policy network, and perform feature fusion processing on various types of state information through a multi-head attention layer to generate a fused representation of multi-source states;
[0062] Step 134, the decision-making layer of the preset policy network generates task roles and behavior decisions based on the fused representation, and outputs high-level discrete action instructions, where the high-level discrete action instructions include role assignment (such as forward, flank, rear guard) and behavior selection (such as no deterrence, non-lethal deterrence, lethal deterrence);
[0063] Step 135, the basic layer of the policy network combines the fused representation and the output of the decision-making layer to generate continuous control instructions, including the normalized target speed and heading angle, for guiding the actual movement behavior of the unmanned boat;
[0064] Step 136, execute the control instructions and interact with the training environment (unmanned boat pursuit training environment) to update the state of the current unmanned boat;
[0065] Step 137, according to the structured reward function, calculate the immediate reward based on the updated state and actions (continuous control instructions), and use the multi-agent proximal policy optimization algorithm (MAPPO) to optimize and train the policy network, and finally obtain a control strategy with good task adaptation ability.
[0066] In this embodiment, the unmanned boat capture training environment refers to an overall simulation platform for developing and evaluating capture control algorithms. The environment integrates a hydrodynamic physics engine, a sensor and sea state disturbance model, an action-state interaction interface, a reward evaluation interface, and a log visualization tool, and can simulate the real dynamic response of target boats, unmanned boats, and obstacles and their interaction with environmental factors such as wind and waves on a computer with high fidelity. The training environment provides a large-scale "virtual test field" that is configurable, repeatable, and can be run without a real boat, providing a closed-loop data flow for the reinforcement learning algorithm: the policy network outputs actions, and the environment returns the next state and immediate rewards, thereby supporting iterative optimization of the policy and performance verification; the typical task scenario is a specific task configuration loaded once in the above training environment, which is used to define the participating entities and initial conditions of a capture round. The scene contains a fixed target boat, a current unmanned boat (trained subject) and multiple coordinated unmanned boats. Static or dynamic obstacles can be deployed as needed, and sea conditions such as wind speed, wave height, and wave direction can be set. Each scene also specifies the starting position, speed, role label, and the end of the roundup judgment rule, so that the training environment can generate continuous time series state data and reward signals during operation. By batch generating different typical mission scenarios, it can cover a variety of target maneuvers, the number of coordinated boats, and complex sea conditions, thereby improving the generalization and robustness of the policy network.
[0067] like Figure 2 As shown, in this embodiment, the preset strategy network adopts a multi-level structure with high- and low-level decoupling, which is responsible for high-level strategy generation and low-level control instruction output respectively. The high-level strategy network (decision layer) and the low-level control network (base layer) share a set of multi-head attention modules to extract multi-source fusion feature representation from the structured state input vector, unify the state information related to the modeling task, and improve the feature expression efficiency and the overall stability of the network.
[0068] The decision layer is used to generate high-level discrete action instructions, including task role assignment and behavior selection; role types include forward (0), wing (1), and defender (2); behavior types include no deterrence (3), non-lethal deterrence (4), and lethal deterrence (5). The output discrete action vector is recorded as: ,in, Indicates the role task currently performed by the unmanned boat. Indicates the behavior instructions of the current unmanned boat; the decision layer network includes two fully connected hidden layers, containing 256 and 128 neurons respectively, and the activation function is ReLU. The output layer uses a softmax structure to normalize the probability of the role and the behavior action respectively, and selects the action corresponding to the maximum probability as the execution instruction;
[0069] The base layer (low-level control network) generates a continuous control action vector based on the received fused features and high-level outputs, which is used to adjust the navigation behavior of the unmanned boat. The output action is defined as: , where represents the normalized target speed, represents the maximum reverse, represents the maximum forward, represents the target heading angle in the local coordinate system; the input vector of the base layer is the fused state feature and the output action of the decision-making layer, and the concatenation result is expressed as: , where , is the attention output dimension; the base layer network also includes two fully connected hidden layers, which contain 256 and 128 neurons respectively, the activation function is ReLU, and the final output control action guides the unmanned boat to execute the target navigation speed and heading;
[0070] As Figure 3 shown, in this embodiment, the off-line reinforcement learning of the policy network is completed iteratively in the unmanned boat encirclement training environment, and the specific process is as follows:
[0071] Initialize the encirclement scenario, set a target boat, a current unmanned boat and several cooperative unmanned boats; establish a local coordinate system with the target boat as the origin, synchronously obtain the states of the current unmanned boat, cooperative unmanned boats, obstacles and sea conditions and uniformly express them; on this basis, generate a unified format input vector, sort the states of the cooperative unmanned boats and obstacles according to the distance from the current unmanned boat, clip if the quantity exceeds the limit, and fill with zero vectors if it is insufficient to keep a fixed dimension; the input vector fuses multi-source information through the multi-head attention module (the current unmanned boat state as the query vector, and the other states as the key / value), and outputs hidden features for subsequent networks to use; the high-level policy network probabilistically gives discrete actions such as role assignment and deterrence behavior according to the fused features, and the low-level control network combines the high-level actions and environmental features to generate the normalized speed and heading angle; each unmanned boat sails according to its own instructions, and the target boat and the environment evolve synchronously; calculate the immediate reward according to the structured reward function (role reward + environmental safety reward + task level reward), and use the multi-agent proximal policy optimization algorithm (MAPPO) to adjust the network parameters; if the continuous encirclement score reaches the threshold, assign a termination reward and end the current round, otherwise enter the next loop. Through this "high-low layer decoupling + action decoupling" architecture, the cooperative unmanned boats with different roles can form differentiated and transferable encirclement strategies during training, significantly improving the encirclement success rate in complex sea conditions.
[0072] In an optional embodiment of the present invention, step 131 may include:
[0073] Step 1311: In the unmanned boat pursuit training environment, construct a pursuit scenario including a target boat, a current unmanned boat, and multiple cooperative unmanned boats, and synchronously generate obstacles and ocean environmental disturbances to simulate their interactions with each other and with the environment, thus constructing a typical pursuit task scenario.
[0074] Step 1312: Generate training state data through the typical pursuit task scenario. The training state data includes the state data of the target boat, the current unmanned boat, the cooperative unmanned boats, and the obstacles; that is, record the initial states of the target boat, the current unmanned boat, the cooperative unmanned boats, the obstacles, and the ocean environment, and start the pursuit training round.
[0075] In a preferred embodiment, to improve the fidelity of the pursuit training environment, the following key scenario elements can be further introduced, including: Obstacle modeling: Set static obstacles (such as fixed facilities, navigation buoys) and dynamic obstacles (such as moving boats, floating objects); Environmental disturbance modeling: Parametrically model sea state disturbances such as wind speed, wind direction, wave height, and wave direction, and input them as state variables into the policy network; Game confrontation simulation: The target boat executes evasion and breakthrough maneuvers according to the anti-pursuit strategy to test the robustness of the strategy.
[0076] In an alternative embodiment of the present invention, step 132 may include:
[0077] Step 1321: Based on the training state data, sort the state objects perceived by the current unmanned boat in the training state data according to the relative distance from the current unmanned boat. The state objects include a variable number of cooperative unmanned boat state objects and obstacle state objects perceived by the current unmanned boat; these two types of objects are heterogeneous targets with dynamically changing quantities in the pursuit task and need to be sorted separately according to the relative distance from the current unmanned boat to unify the input format. Among them, the state of the cooperative unmanned boat (such as position, speed, heading, role, etc.); the state of the obstacle (such as position, speed, direction, radius, etc.).
[0078] Step 1322: When the number of state objects exceeds the preset maximum value, perform cropping; when it does not reach the maximum value, fill it with zero vectors, and construct the sorted, cropped, and filled state data into an input vector in a unified format.
[0079] In this embodiment, the cooperative unmanned boats and the obstacles are sorted in ascending order of the distance from the current unmanned boat: If the quantities respectively exceed the upper limits 、 , then only retain the state information of the nearest cooperative unmanned boats and obstacles; if they are insufficient, fill them with zero vectors at the end; the processing results are concatenated to obtain an input vector with a fixed dimension , where is the current state of the unmanned boat, is the target state, and are the states of the cooperative unmanned boat and the obstacle after sorting / cropping / padding respectively, is the state of the marine environment; through this preprocessing, the input dimension can be kept consistent when the number of entities changes dynamically, ensuring the stable operation of the network and improving the system scalability.
[0080] In an optional embodiment of the present invention, step 133 may include:
[0081] Step 1331, embed various state information in the input vector of unified format into a feature space of the same dimension to obtain an input matrix;
[0082] Step 1332, based on the input matrix, adopt a multi-head attention mechanism for feature fusion, specifically including: calculating the outputs of multiple attention heads (using the current unmanned boat state information as the query vector, and the remaining states as the key and value, calculating the attention weights through multiple independent attention heads and obtaining the outputs), and concatenating the results of each attention head and then performing a linear mapping to generate a fused feature representation as the input of the subsequent module of the policy network.
[0083] In this embodiment, to enhance the modeling ability of the policy network for multi-source state information and improve its perception and decision-making ability in complex encirclement tasks, the present invention introduces a customized multi-head attention mechanism into the network structure; the multi-head attention mechanism is embedded in the policy network as a state aggregation module, using the structured input vector for multi-source information fusion to improve the expression ability for heterogeneous situations, and the specific operations are as follows:
[0084] The input vector , where , , , , represent the state vectors of the current unmanned boat, the target boat, the cooperative unmanned boat, the obstacle, and the marine environment respectively, , are the maximum numbers of the cooperative unmanned boat and the obstacle, and the input dimension consistency is maintained through cropping and padding; embed the above five types of state data into a feature space of a unified dimension to construct an input matrix: , , where represents the total number of state sources, corresponding to the current unmanned boat, the target boat, the cooperative unmanned boat, the obstacle, and the marine environment in sequence; based on the scaled dot-product attention mechanism, use the embedded representation of the current unmanned boat state as the query vector, and the remaining states as the context information (key and value), and construct the attention calculation process as follows: , , , where represents the embedded features of the current unmanned boat; represents the remaining state inputs except the current unmanned boat, , , are trainable linear transformation matrices, corresponding to the transformation parameters of the query, key, and value vectors respectively, and are the dimensions of the query / key vector and the value vector respectively; the above vectors pass through the attention function: to calculate the attention weights of the current unmanned boat for each state source and the aggregated feature representation; to enhance the expression ability for different feature subspaces, parallel attention heads are introduced, and each head independently calculates the attention output: ; the outputs of each attention head are concatenated and linearly mapped to obtain the final fused feature output: , where is the output mapping matrix, is the dimension of the final fused feature; this fused feature is used as the shared input of the decision-making layer and the basic layer of the policy network, enabling the unmanned boat to dynamically focus on the target position, cooperative formation, obstacle threat, and sea condition disturbance according to the real-time situation, thereby improving the perception accuracy and cooperative control performance in the encirclement scenario.
[0085] In an optional embodiment of the present invention, the structured reward function includes role-specific reward items designed based on task division of labor, including:
[0086] The striker role reward function for the striker unmanned boat to quickly approach the target boat and block it frontally; the flank role reward function for the flank unmanned boat to assist in surrounding from both sides and enhance deterrence cooperation; the defender role reward function for the defender unmanned boat to implement suppression from behind the target boat and deter and intervene in a high-threat state.
[0087] As Figure 4 shown, in this embodiment, for the behavioral objectives of the three types of task roles of striker, flank, and defender, role-specific reward functions are designed respectively, including:
[0088] Striker role reward function ; where represents the relative distance change between the striker and the target boat, used to evaluate the immediate approach efficiency; represents the relative distance change between the striker and the target boat in the current frame relative to the initial moment, used to measure the overall approach trend; represents the front-blocking effect reward, defined as: , where the blocking angle The interception ability of the forward located in the escape direction of the target boat (such as Figure 4 in ) is defined as: , and are the two-dimensional position vectors of the target boat and the forward respectively, is the velocity vector of the target boat; is the weighting coefficient of the forward role reward item;
[0089] Flank role reward function ; where, represents the degree of cooperation forming an enclosing angle with the forward, is the angle formed by the forward and the flank with the target boat as the vertex (such as Figure 4 in and ), is the ideal enclosing angle (such as ), controls the angle sensitivity; measures the influence of the distance between the flank and the target boat on the deterrence effect, is the distance between the flank and the target boat, is the maximum distance for the deterrence to take effect; is the weight of the flank role reward item;
[0090] Defender role reward function ; where, the suppression situation reward , is the suppression angle quantifying whether the defender is behind the escape direction of the target boat (such as Figure 4 in ), is the two-dimensional position vector of the defender; the deterrence response matching reward , indicates whether the target boat is in a threatened state (such as escape / attack), is the real-time distance between the defender and the target boat, is the optimal deterrence distance, controls the distance sensitivity of the deterrence response; is the current deterrence behavior level of the defender, is defined as follows: , 0 indicates no deterrence, represent the reward intensities of non-lethal and lethal deterrence respectively; the dynamic distance maintenance reward , and are the acceptable minimum and maximum tracking distances respectively, is the recommended tracking distance; are the weighting coefficients of the suppression, deterrence response and distance maintenance rewards respectively.
[0091] In a preferred embodiment, the structured reward function further sets an environment interaction related reward item to enhance the obstacle avoidance ability and path stability of the unmanned boat under complex sea conditions, mainly including two sub-modules: collision avoidance reward and collision penalty;
[0092] Define the total environmental safety reward as: ;
[0093] Among them, the obstacle collision avoidance reward item is used to encourage the unmanned boat to actively avoid surrounding obstacles and maintain a safe distance, and is represented by an exponential function: , where, represents the Euclidean distance between the current unmanned boat and the nearest obstacle, is the perception range scale parameter, is a small constant to prevent the denominator from being zero, is the weight coefficient of the collision avoidance reward item;
[0094] The collision penalty item is used to suppress the unmanned boat from entering high-risk areas and is represented by a continuous risk function based on the Sigmoid function: , where, is the maximum penalty intensity, is the collision risk determination distance threshold, is the response slope coefficient.
[0095] In a preferred embodiment, the structured reward function further sets a task level reward item, and through a task termination reward mechanism based on continuous situation scoring, guides the strategy to converge stably to an effective encirclement solution, specifically including:
[0096] Comprehensively measure the distance proximity and direction suppression between the individual and the target, and construct an encirclement success scoring function , where, is the number of effective unmanned boats participating in the encirclement currently, is the th unmanned boat's distance from the target boat, is the distance threshold for determining an effective encirclement, is used to truncate negative values, is the weight of the distance term and the direction term, represents the azimuth angle between the target boat's movement direction and the th unmanned boat; The role direction suppression function gives differential rewards according to the role type: Among them, indicates that the current role is a forward, indicates that the current role is a flank, indicates that the current role is a defender, and is an angle-sensitive parameter used to adjust the precision of flanking and the precision of suppressing the full-back.
[0097] In this embodiment, to improve the stability of the determination of the success of the encirclement, the sliding window method is used to smooth the score fluctuations; let the window length At time calculate the average score within the window where is the instantaneous score of the th frame; when the condition is met, that is, the sliding average encirclement score is not lower than the set encirclement success determination threshold , it is determined that the encirclement is successful and the current round ends;
[0098] This embodiment further sets a task-level reward to give global feedback according to the encirclement result at the end of the round. Its core logic includes: when the sliding average encirclement score satisfies , that is, when it is determined that the encirclement is successful, a termination reward is issued according to the linear proportion of the encirclement score, where is the maximum reward value that can be obtained when the task is successful; if the current round reaches the maximum number of steps , but the encirclement score is still lower than the task success threshold , it is determined that the task fails, and a fixed penalty term is imposed, where is the penalty value when the task fails; the task termination reward is settled uniformly at the end of the round: ; indicates that if the encirclement is successful, indicates that if the encirclement fails, 0 indicates other situations; in this embodiment, the total instantaneous reward at each time step is , where is the character-specific reward (forward, flanker, full-back), is the environmental interaction reward (collision avoidance reward and collision penalty), is the above-mentioned task termination reward; through the "character-environment-task" three-level reward architecture, the algorithm can not only maintain individual division of labor and cooperative suppression during the training process, but also take into account safe trajectories and task orientation, thus significantly improving the encirclement success rate and strategy robustness in complex sea conditions.
[0099] As Figure 5 shown, an embodiment of the present invention also provides an intelligent encirclement control device 50 for an unmanned boat, including:
[0100] An acquisition module 51, configured to acquire the state information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to represent the state information, so as to obtain the sensed state information corresponding to the current unmanned boat;
[0101] A processing module 52, configured to construct a structured state input vector based on the perception state information by using a distance sorting and filling and cropping mechanism;
[0102] A control module 53, configured to input the structured state input vector into a hierarchical policy network, where a high layer of the hierarchical policy network generates task roles and behavior decisions, and a low layer combines the output of the high layer and the environmental state in the structured state input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism by using reinforcement learning;
[0103] An execution module 54, configured to control a current unmanned boat to perform a pursuit task on a target boat according to the navigation speed and direction control instructions.
[0104] Optionally, obtaining the state information of the current unmanned boat and the local environmental data sensed by the current unmanned boat, and constructing a local coordinate system based on the target boat to uniformly represent the state information, so as to obtain the perception state information corresponding to the current unmanned boat, including:
[0105] Obtaining the state information of the current unmanned boat and the local environmental data sensed by the current unmanned boat;
[0106] Establishing a local coordinate system with the position of the target boat as the origin, and uniformly representing the state information of the current unmanned boat and the local environmental data sensed by the current unmanned boat in the local coordinate system to obtain the perception state information corresponding to the current unmanned boat; wherein, the local environmental data includes: the state of the target boat, the state of the cooperative unmanned boat, the state of the obstacle, and the state of the ocean environment.
[0107] Optionally, optimizing a preset policy network through a reward mechanism by using reinforcement learning to obtain a hierarchical policy network, including:
[0108] Constructing a typical task scenario in an unmanned boat pursuit training environment, and generating training state data;
[0109] Based on the training state data, in view of the dynamic changes in the number of cooperative unmanned boats and obstacles, using a distance sorting, filling and cropping mechanism to organize the state information sensed by the current unmanned boat into an input vector in a unified format;
[0110] Inputting the input vector into an input layer of a preset policy network, and performing feature fusion on the input vector through a multi-head attention layer to generate a fused representation of multi-source states;
[0111] A decision layer of the preset policy network generates task roles and behavior decisions based on the output of the multi-head attention layer, and outputs high-level discrete action instructions;
[0112] The basic layer of the policy network combines the output of the decision-making layer and the output of the multi-head attention layer to generate continuous control instructions;
[0113] Execute the continuous control instructions and interact with the unmanned boat pursuit training environment to update the state of the unmanned boat;
[0114] Based on the structured reward function, calculate the immediate reward according to the updated state of the unmanned boat and the continuous control instructions, and optimize the preset policy network through the multi-agent proximal policy optimization algorithm to obtain the hierarchical policy network.
[0115] Optionally, construct a typical task scenario in the unmanned boat pursuit training environment and generate training state data, including:
[0116] In the unmanned boat pursuit training environment, set a target boat, a current unmanned boat and multiple cooperative unmanned boats, simulate their interactions with each other and with the environment, and construct a typical pursuit task scenario;
[0117] Generate training state data through the typical pursuit task scenario, and the training state data includes the state data of the target boat, the current unmanned boat, the cooperative unmanned boats and the obstacles.
[0118] Optionally, based on the training state data, for the dynamic changes in the number of cooperative unmanned boats and obstacles, adopt a distance sorting, padding and cropping mechanism to organize the state information perceived by the current unmanned boat into an input vector of a unified format, including:
[0119] Based on the training state data, sort the state objects perceived by the current unmanned boat in the training state data according to the relative distance from the current unmanned boat, and the state objects include the variable number of cooperative unmanned boat state objects and obstacle state objects perceived by the current unmanned boat;
[0120] When the number of state objects exceeds the preset maximum value, crop it, fill it with zero vectors when it does not reach, and construct the sorted, cropped and filled state data into an input vector of a unified format.
[0121] Optionally, perform feature fusion processing on various types of state information through the multi-head attention layer to generate a fused representation of multi-source states, including:
[0122] Embed various types of state information in the input vector of the unified format into a feature space of the same dimension to obtain an input matrix;
[0123] Based on the input matrix, calculate the outputs of multiple attention heads, and splice the results of each attention head and then perform a linear mapping to generate a fused representation of multi-source states.
[0124] Optionally, the structured reward function includes:
[0125] A striker role reward function for a striker unmanned boat to quickly approach a target boat and conduct a frontal blockade; a flanker role reward function for a flanker unmanned boat to assist in surrounding from both sides and enhance deterrence cooperation; a defender role reward function for a defender unmanned boat to suppress from behind the target boat and deter intervention in a high-threat state.
[0126] It should be noted that this device corresponds to the above method, and all implementation manners in the above method are applicable to the embodiments of this device and can also achieve the same technical effects.
[0127] An embodiment of the present invention further provides a computing device, including: a processor and a memory. The memory stores a computer program. When the program runs on the processor, it enables the computer to execute the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0128] An embodiment of the present invention further provides a computer-readable storage medium storing instructions. When the instructions run on a computer, it enables the computer to execute the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
[0129] The above are the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An intelligent pursuit control method for an unmanned boat, characterized in that, Including: Obtain the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to uniformly represent the status information, so as to obtain the sensed status information corresponding to the current unmanned boat; Based on the sensed status information, construct a structured status input vector by using a distance sorting and filling and cropping mechanism; Input the structured status input vector into a hierarchical policy network. The high-level of the hierarchical policy network generates task roles and behavior decisions, and the low-level combines the high-level output and the environmental status in the structured status input vector to generate navigation speed and direction control instructions; wherein, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism using reinforcement learning; Control the current unmanned boat to perform a pursuit task on the target boat according to the navigation speed and direction control instructions.
2. The intelligent encirclement control method for the unmanned boat according to claim 1, wherein, Obtain the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat, and construct a local coordinate system based on the target boat to uniformly represent the status information, so as to obtain the sensed status information corresponding to the current unmanned boat, including: Obtain the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat; Establish a local coordinate system with the position of the target boat as the origin, and uniformly represent the status information of the current unmanned boat and the local environment data sensed by the current unmanned boat in the local coordinate system, so as to obtain the sensed status information corresponding to the current unmanned boat; wherein, the local environment data includes: the status of the target boat, the status of the cooperative unmanned boat, the status of the obstacle, and the status of the marine environment.
3. The intelligent encirclement control method for the unmanned boat according to claim 1, characterized in that, Optimizing a preset policy network through reinforcement learning by a reward mechanism to obtain a hierarchical policy network includes: Construct a typical task scenario in the unmanned boat pursuit training environment and generate training status data; Based on the training status data, in view of the dynamic changes in the number of cooperative unmanned boats and obstacles, use a distance sorting, filling and cropping mechanism to organize the status information sensed by the current unmanned boat into an input vector in a unified format; Input the input vector into the input layer of the preset policy network, and perform feature fusion on the input vector through a multi-head attention layer to generate a fused representation of multi-source status; The decision layer of the preset policy network generates task roles and behavior decisions based on the output of the multi-head attention layer and outputs high-level discrete action instructions; The basic layer of the policy network combines the output of the decision layer and the output of the multi-head attention layer to generate continuous control instructions; Execute the continuous control instructions and interact with the unmanned boat pursuit training environment to update the unmanned boat status; Based on a structured reward function, calculate the immediate reward according to the updated unmanned boat status and the continuous control instructions, and optimize the preset policy network through a multi-agent proximal policy optimization algorithm to obtain a hierarchical policy network.
4. The intelligent encirclement control method for the unmanned boat according to claim 3, wherein Construct a typical task scenario in the unmanned boat pursuit training environment and generate training status data, including: In the unmanned boat pursuit training environment, set a target boat, a current unmanned boat and multiple cooperative unmanned boats, simulate their interactions with each other and with the environment, and construct a typical pursuit task scenario; Generate training state data through a typical encirclement mission scenario, where the training state data includes the state data of the target boat, the current unmanned boat, the cooperative unmanned boat, and the obstacles.
5. The intelligent encirclement control method for the unmanned boat according to claim 3, wherein Based on the training state data, in response to the dynamic changes in the number of cooperative unmanned boats and obstacles, use a distance sorting, padding, and cropping mechanism to organize the state information perceived by the current unmanned boat into an input vector in a unified format, including: Based on the training state data, sort the state objects perceived by the current unmanned boat in the training state data according to the relative distance from the current unmanned boat. The state objects include a variable number of cooperative unmanned boat state objects and obstacle state objects perceived by the current unmanned boat. When the number of state objects exceeds the preset maximum value, crop them; when it does not reach the maximum value, fill them with zero vectors, and construct the sorted, cropped, and filled state data into an input vector in a unified format.
6. The intelligent pursuit control method for an unmanned boat according to claim 3, wherein Perform feature fusion processing on various types of state information through a multi-head attention layer to generate a fused representation of multi-source states, including: Embed various types of state information in the input vector in a unified format into a feature space of the same dimension to obtain an input matrix. Based on the input matrix, calculate the outputs of multiple attention heads, and splice the results of each attention head and then perform a linear mapping to generate a fused representation of multi-source states.
7. The intelligent encirclement control method for the unmanned boat according to claim 3, characterized in that The structured reward function includes: A striker role reward function for the striker unmanned boat to quickly approach the target boat and block it frontally; a flanker role reward function for the flanker unmanned boats to assist in the encirclement from both sides and enhance deterrence and cooperation; a defender role reward function for the defender unmanned boat to suppress from behind the target boat and deter and intervene in a high-threat state.
8. An intelligent control device for unmanned boats to conduct intelligent encirclement and capture, characterized in that, It includes: An acquisition module for acquiring the state information of the current unmanned boat and the local environment data perceived by the current unmanned boat, and constructing a local coordinate system based on the target boat to uniformly represent the state information, and obtaining the perceived state information corresponding to the current unmanned boat. A processing module for constructing a structured state input vector based on the perceived state information by using a distance sorting and padding / cropping mechanism. A control module for inputting the structured state input vector into a hierarchical policy network. The high level of the hierarchical policy network generates task roles and behavior decisions, and the low level combines the high-level output and the environmental state in the structured state input vector to generate navigation speed and direction control instructions. Among them, the hierarchical policy network is obtained by optimizing a preset policy network through a structured reward mechanism using reinforcement learning. An execution module for controlling the current unmanned boat to perform an encirclement mission on the target boat according to the navigation speed and direction control instructions.
9. A computing device, characterized in that, It includes: A processor and a memory. The memory stores a computer program, and when the program runs on the processor, it executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A storage instruction, when the instruction runs on a computer, causes the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Task distribution method of MAUVS round up
CN107562074A
Centerless robot cluster surrounding task control method and system
CN112527012A
Construction method and application of unmanned ship cluster hunting control model
CN116400700A
Multi-machine hunting method and device for hierarchical collaborative learning, electronic equipment and medium
CN117350326A
Multi-agent encompassing reinforcement learning method based on skill learning and self-attention
CN118569066A
Cited By
Foot type robot cluster cooperative control method fusing ground acting force perception
CN120821273A
Gas separation regulation and control system based on multi-agent cooperation
CN121541456A
Unmanned aerial vehicle cluster autonomous collaborative countering method
CN121613948A