An intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC

By introducing a mask attention mechanism into the SAC algorithm, the unmanned boat collision avoidance algorithm is improved, which solves the problem of insufficient adaptability caused by the change in the number of dynamic target ships, and achieves more efficient and safe collision avoidance decisions.

CN120010526BActive Publication Date: 2025-06-27JIMEI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510484418.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-27
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing unmanned boat collision avoidance algorithm based on deep reinforcement learning is difficult to dynamically adapt when facing the dynamically changing target ship count, resulting in insufficient adaptability and generalization capabilities of collision avoidance strategies.

Method used

The mask attention mechanism is used to adaptively process the variable-length target ship information, and reinforcement learning training is carried out in combination with the SAC algorithm to generate dimensionally unified fusion state features, which are used to dynamically adjust collision avoidance decisions.

Benefits of technology

It effectively resolves the input dimension conflict caused by the dynamic changes in the number of target ships, improves the adaptability and generalization of collision avoidance strategies of unmanned boats in complex sea conditions, and improves navigation safety and decision-making efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010526B_ABST
    Figure CN120010526B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent collision avoidance method for an unmanned boat based on a masked attention mechanism and SAC, including: Dynamic target ship information processing: adaptively fuse features of variable-length target ship information obtained in real time through a masked attention mechanism, dynamically generate masked attention weights based on the real-time collision risk score of the target ship, and mask invalid target ship data through a padding or truncation strategy to generate a fusion state feature with a unified dimension; SAC framework collaborative optimization: input the fusion state feature into the Actor network to generate a continuous collision avoidance action strategy; input the fusion state feature into the Critic network, estimate the state-action value function through a double Q-network structure combined with an entropy regularization term, and balance policy exploration and exploitation; End-to-end decision execution: optimize network parameters in the training phase and output collision avoidance control instructions in real time in the deployment phase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of collision avoidance of unmanned boats, reinforcement learning, etc., and particularly relates to an intelligent collision avoidance method for unmanned boats based on a masked attention mechanism and SAC. Background Art

[0002] As the core carrier of intelligent marine equipment, Unmanned Surface Vehicles (USVs) play an irreplaceable role in fields such as marine environmental monitoring, maritime search and rescue, port inspection, and military reconnaissance. However, with the continuous improvement of the complexity of marine tasks, traditional collision avoidance methods (such as the artificial potential field method and the dynamic window method) have gradually shown limitations in scenarios such as dynamic obstacle interaction and multi-target cooperative collision avoidance, including poor environmental adaptability and insufficient computational real-time performance. In recent years, collision avoidance methods based on Deep Reinforcement Learning (DRL) have provided new solutions to the USV collision avoidance problem with their adaptability to complex dynamic environments and end-to-end decision-making advantages, and have become a research hotspot in this field.

[0003] Currently, USV collision avoidance algorithms based on DRL are mainly divided into two categories:

[0004] Value-based algorithms (such as DQN, Double DQN): Such algorithms learn the state-action value function through the Q-learning framework with the goal of maximizing long-term rewards. Their advantage lies in the clear theoretical convergence, but there are problems of overestimation bias, and they are limited by the discrete action space, making it difficult to meet the control requirements of continuous steering and speed change of USVs.

[0005] Policy-based algorithms (such as PPO, SAC): Such algorithms directly optimize the policy function and update parameters through policy gradients, avoiding the complexity of value function calculation. Policy-based algorithms support continuous action output and show stronger flexibility and robustness in complex and continuous environments, especially suitable for the high-dimensional and dynamic state space in the unmanned boat collision avoidance task.

[0006] However, policy-based algorithms still face challenges in practical applications. When the number of target ships changes dynamically, traditional methods need to fix the state space dimension, resulting in deficiencies in generalization and adaptability. Existing technologies usually adopt padding or truncation strategies to handle variable numbers of target ships, but padding may introduce redundant noise and reduce training efficiency; truncation may lead to the loss of key obstacle information and increase navigation risks.

[0007] Related Prior Art

[0008] The prior art closest to the present invention includes:

[0009] The patent with the authorization announcement number CN117168468B proposes a multi-unmanned boat collaborative navigation method based on proximal policy optimization;

[0010] The patent with the authorization announcement number CN110658829B discloses a ship collision avoidance decision-making method based on generative adversarial imitation learning;

[0011] The patent with the publication number CN117523925A designs a collision avoidance method for a group of unmanned boats combining deep reinforcement learning and LSTM neural network;

[0012] The patent with the publication number CN116954232A proposes a multi-ship collision avoidance decision-making system based on reinforcement learning.

[0013] The above technologies have not effectively solved the contradiction between the dynamic change of the number of target ships and the fixed state space requirement of the SAC algorithm, resulting in limited adaptability and generalization ability of the collision avoidance strategy. In actual navigation, the number of target ships changes dynamically with the sea area, time, and navigation conditions, while existing methods usually only model a fixed number of target ships and adapt the input dimension by padding or truncating. This processing method has significant defects: padding redundant information may interfere with model training, while truncating key target ship information may pose safety hazards in high-density navigation scenarios. Summary of the Invention

[0014] To solve the problem that the existing solutions do not consider the dynamic change of the number of target ships and the fixed state space dimension constraint of the SAC algorithm, resulting in insufficient adaptability of the collision avoidance strategy of unmanned boats in complex sea conditions and poor generalization ability, and being unable to well meet the dynamic collision avoidance requirements in different density navigation environments, the present invention proposes an intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC. By using the masked attention mechanism to adaptively process the variable-length target ship information and combining with the SAC algorithm for reinforcement learning training, it can effectively solve the problems that existing methods cannot dynamically adapt to the change of the number of target ships, the information loss caused by padding or truncating strategies, and the influence of fixed state dimension constraints on decision-making stability, thereby improving the intelligent level and navigation safety of the unmanned boat collision avoidance system.

[0015] Its core design is as follows:

[0016] An intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC, including:

[0017] Dynamic target ship information processing: Adaptive feature fusion of the real-time obtained variable-length target ship information through the masked attention mechanism, dynamically generating masked attention weights based on the real-time collision risk score of the target ship, and shielding invalid target ship data through padding or truncating strategies to generate a unified dimension fusion state feature;

[0018] SAC Framework Collaborative Optimization:

[0019] Input the fused state feature into the Actor network to generate a continuous collision avoidance action strategy;

[0020] Input the fused state feature into the Critic network, and estimate the state-action value function by combining the double Q-network structure with the entropy regularization term to balance policy exploration and exploitation;

[0021] End-to-End Decision Execution: Optimize the network parameters during the training phase and output collision avoidance control instructions in real time during the deployment phase.

[0022] The core design of the present invention lies in embedding the masked attention mechanism into the SAC framework to solve the input dimension conflict caused by the change in the number of dynamic target ships. As Figure 3 shown, the fused state feature (f in ) is processed by the masked attention generation module (step S3, see the embodiment part later, the same below), and then input into the Actor network (step S4.2) and the Critic network (step 6) simultaneously. During the training phase (steps 1-8), the model optimizes the policy network parameters through entropy regularization and realizes robust learning by combining the dynamic masking mechanism; during the deployment phase (steps S1-S6), the trained Actor network is loaded to output collision avoidance actions in real time. This design verifies its effectiveness through the dynamic priority rule (step S2.4) and the end-to-end decision-making process (step S5) in the "Detailed Implementation".

[0023] Among them, the masked attention generation module: corresponding to steps S3.2-S3.9 ( Figure 2 process), generates the fused feature f in through separate feature fusion (the own ship state and the target ship state are processed independently);

[0024] Actor-Critic Collaborative Optimization: Steps S4.2 (Actor network) and steps 7.2-7.5 (Critic network) clearly take f in as the input, and balance exploration and exploitation through entropy regularization (Equations 24-27).

[0025] End-to-End Decision-Making Process:

[0026] Steps S5 (action mapping) and S6 (real-time control) verify the effectiveness during the deployment phase.

[0027] Furthermore, the generation rule of the mask includes:

[0028] Full Masking Processing:

[0029] When the number of target ships exceeds the preset threshold, select the N target ships with the largest collision risk scores;

[0030] Completely shield the invalid target ship by filling with 0s or truncating the extra targets;

[0031] Gradient retention mechanism:

[0032] During the training phase, apply a minimum value to the attention scores corresponding to the invalid target ships to retain gradient propagation.

[0033] Among them, the training phase (gradient retention):

[0034] Step S3.6: "Set the attention score of the invalid target ship to 10 -9 , avoid participating in the Softmax calculation but retain gradient propagation" (corresponding to Equation 9 - Equation 10).

[0035] Deployment phase (complete shielding):

[0036] Step S2.4: "When the number of target ships exceeds the threshold, fill with 0s or truncate the extra targets" (corresponding to the state construction logic).

[0037] Dynamic priority rule:

[0038] Step S2.4: "Select the N target ships with the highest collision risk", and the scoring is based on Equation 20.

[0039] Furthermore, the construction of the fused state features includes:

[0040] Separate the state s of the own ship OS from the state s of the target ship TS , and generate the query vector Q and the key - value vectors {K, V} through independent linear transformations;

[0041] Concatenate the query vector Q with the key vector K of each target ship i and input it into the additive attention network to generate the dynamic attention weight α i ;

[0042] According to the dynamic attention weight α i weight - sum the value vectors V i to generate the fused feature f of the target ship TS , and concatenate it with the own - ship feature Q to form the final input vector f in .

[0043] Furthermore, the input of the Critic network is the fused state feature f generated by the masked attention mechanism in, the value function is optimized by combining the double Q-network with the entropy term. The target Q-value calculation combines the state features generated by the masked attention mechanism and the negative logarithm term of the policy entropy to dynamically balance exploration and exploitation; the Actor network updates the policy parameters by maximizing the entropy-regularized value weighted by the masked attention, and the parameters of the Critic network are softly updated by the mixing coefficient τ.

[0044] Among them, the design of the double Q-network:

[0045] Step 7.2: "The Critic network adopts a double Q structure, and optimizes the value function by minimizing the mean square error of the outputs of the two Q-networks."

[0046] Entropy regularization term:

[0047] Equation 24 (target Q-value calculation) includes the entropy term, and the adaptive temperature coefficient α is optimized by Equation 26.

[0048] Soft update mechanism:

[0049] Equation 27 (update of target network parameters), mixing coefficient τ = 0.005.

[0050] Furthermore, the output of the collision avoidance action policy includes:

[0051] The normalized action parameters output by the Actor network are converted into thrust and torque through linear mapping;

[0052] According to the propeller dynamics model, the thrust is converted into a rotational speed command, and according to the rudder servo model, the torque is converted into a rudder angle command.

[0053] Furthermore, the reward function in the training stage is a multi-objective weighted sum, including:

[0054] Distance reward term: Guide the unmanned boat to approach the target point quickly;

[0055] Course reward term: Keep the course towards the target point;

[0056] Collision risk penalty term: Calculate the dynamic threat score based on the closest point of approach distance and time of encounter between the target ship and the unmanned boat;

[0057] Rule compliance penalty term: Enforce the steering and speed constraints in compliance with the International Regulations for Preventing Collisions at Sea.

[0058] Furthermore, the priority rule for the dynamic target ship information is: when the number of detected target ships exceeds the preset threshold, select the N target ships with the greatest collision risk calculated based on the closest point of approach distance and the closest time of encounter.

[0059] Furthermore, the input of the Critic network of the SAC framework is the fused state features generated by the masked attention mechanism.

[0060] Further, the unmanned boat dynamics model in the training stage is a three-degree-of-freedom MMG model, including surge, sway, and yaw motion equations.

[0061] In addition, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.

[0062] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0063] Although the embodiment takes SAC as the core framework, those skilled in the art can consider equivalently replacing it with TD3 or PPO based on the "generality of the Actor-Critic architecture". However, the SAC framework solution of the embodiment of the present invention is still the best implementation solution.

[0064] Compared with the prior art, the intelligent collision avoidance method of the present invention and its preferred solution deeply integrate the masked attention mechanism with the SAC framework, and achieve the following outstanding technical effects in dynamic and complex navigation scenarios:

[0065] Improved adaptability to dynamic environment:

[0066] In the training stage, the invalid target ship gradient is retained by the minimum value, and the weight allocation rule of the effective target is learned;

[0067] In the deployment stage, redundant information interference is avoided through complete masking processing, and key threat targets are dynamically focused on.

[0068] Enhanced decision-making efficiency and safety:

[0069] The Actor network generates continuous collision avoidance actions based on linear mapping, realizing precise control of the course and thrust;

[0070] The Critic network optimizes the value estimation through the double Q structure and the entropy regularization term, balancing exploration and exploitation.

[0071] Guarantee of rule compliance:

[0072] The design of the reward function enforces that the collision avoidance strategy complies with the International Regulations for Preventing Collisions at Sea, improving navigation safety and compliance. Description of the Drawings

[0073] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0074] Figure 1 It is the flowchart of the collision avoidance decision-making of the unmanned boat in the embodiment of the present invention.

[0075] Figure 2 This is the flowchart of the mask attention mechanism processing in the embodiment of the present invention.

[0076] Figure 3 This is the structural diagram of the intelligent collision avoidance algorithm for unmanned boats based on the mask attention mechanism and SAC in the embodiment of the present invention. Detailed implementation manners

[0077] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail as follows:

[0078] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0079] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0080] The embodiment of the present invention proposes an intelligent collision avoidance method for unmanned boats based on the mask attention mechanism and SAC. This method effectively alleviates the high dependence of existing algorithms on the fixed state space dimension by introducing the mask attention mechanism to adaptively model the variable-length target ship information. By dynamically adjusting the attention weights, this method can maintain stable collision avoidance decision-making ability under different numbers of target ships, improving the adaptability and safety of unmanned boats in complex navigation environments.

[0081] The collision avoidance decision-making process provided in this embodiment ( Figure 1 ) is the online deployment application process of the trained model. Its core modules (mask attention, Actor / Critic network) optimize the parameters through the subsequent training process (Steps 1-8):

[0082] Step S1: Collect sensor data;

[0083] Step S2: Data preprocessing;

[0084] Step S3: Use the mask attention mechanism to process the data again and calculate the input of the collision avoidance model;

[0085] Step S4: The collision avoidance model calculates the decision;

[0086] Step S5: The unmanned boat control system executes the collision avoidance decision;

[0087] Step S6: Repeat steps S1 - S5 until the collision avoidance behavior is completed.

[0088] For step S1, the types of information required in the collision avoidance task of the unmanned ship include the navigation status information of the own ship (OS) and the information of the target ships (TS). Among them, the navigation status of the own ship is mainly obtained through the Global Navigation Satellite System (GNSS), including information such as the real - time position, speed, and heading of the own ship; the navigation status of the target ships is mainly obtained through millimeter - wave radar detection, including information such as the current position, speed, and heading of the target ships. These information provide important data support for the dynamic collision avoidance decision - making of the unmanned ship.

[0089] For step S2, data pre - processing is an important link to achieve collision avoidance, aiming to process the originally acquired navigation data to meet the requirements of the algorithm input. As the preferred solution of this embodiment, the specific steps are as follows:

[0090] Step S2.1, Use the sliding window filtering method to identify and remove outliers in the position, speed, and heading information provided by GNSS and millimeter - wave radar, ensuring the authenticity and reliability of the data.

[0091] Step S2.2, Convert the relative coordinate data obtained by the millimeter - wave radar into longitude - latitude coordinates according to the longitude - latitude information of the own ship obtained by GNSS. The conversion formula is:

[0092] (1)

[0093] In the formula, r is the average radius of the earth, and are the longitude - latitude positions of the own ship and the target ship respectively, is the relative position of the target ship detected by the millimeter - wave radar.

[0094] Subsequently, use the Mercator projection to convert the longitude - latitude positions of the own ship and the target ship into the Cartesian coordinate system for facilitating the input of the subsequent calculation model. The Mercator projection formula is shown as formula (2) below:

[0095] (2)

[0096] In the formula is the accuracy of the central meridian, are the latitude and longitude of the position to be converted respectively.

[0097] Step S2.3, Use the linear interpolation method to align the data of GNSS and millimeter - wave radar, ensuring that the timestamps of all input data are consistent. The specific calculation formula is shown as formula (3) below:

[0098] (3)

[0099] Where t is the time to be interpolated, are two adjacent times and satisfy , are respectively The data at time. Through this interpolation method, the consistency of all sensor data at the same timestamp can be guaranteed.

[0100] Step S2.4, construct the masked attention input according to the processed data. The input consists of three parts: , and the calculation formula is shown in Equation (4):

[0101] (4)

[0102] Where represent the speed, heading, and turning speed of the OS respectively. represents the relative bearing and distance from the OS to the target point. are respectively composed of the information of each target ship. The of the i-th target ship is calculated as follows:

[0103] (5)

[0104] Where represents the distance between the OS and . represents the relative bearing and relative heading from the OS to . is the speed. represents the encounter situation between the OS and , while represents the collision risk between the OS and . is The safety area radius. In order to adapt to the existing mainstream code interfaces, a sufficiently large number n of TSs will be defined during the calculation of . When the number of TSs detected within the range of the OS sensor is less than n, 0 will be filled; when the number of TSs detected within the range of the OS sensor is greater than n, the n TSs with the greatest collision risk will be selected to calculate .

[0105] For step S3, the attention mechanism used is additive attention. By using the masked attention mechanism to dynamically weight multiple TS information, the collision avoidance model can automatically identify key target information, capture the state features that best reflect the impact of TSs on the OS navigation decision, and improve the effectiveness of collision avoidance decisions. The processing flow of this step is as Figure 2As shown in the figure, as a preferred solution of this embodiment, the specific implementation steps are as follows:

[0106] Step S3.1: Divide the data s preprocessed in step S2 into two parts, namely, representing the state characteristics of the target ship and .

[0107] Step S3.2: Perform a linear transformation on to obtain a query vector Q. The calculation formula is as follows:

[0108] (6)

[0109] where , are trainable parameters, and is the hidden feature dimension.

[0110] Step S3.3: Perform a linear transformation on each target ship state feature to obtain a key vector and a value vector , where :

[0111] (7)

[0112] where , are trainable parameters.

[0113] Step S3.4: After copying and expanding the query vector Q to , splice it with the key vector to form a comprehensive feature vector :

[0114] (8)

[0115] Step S3.5: Calculate the attention score of each target ship through an additive attention network (composed of a linear layer and a Tanh activation function). The calculation formula is as follows:

[0116] (9)

[0117] where , , and are all trainable parameters.

[0118] Step S3.6: Use the masked attention mechanism to set the attention scores corresponding to invalid target ships to to avoid participating in the Softmax calculation.

[0119] Step S3.7: Perform Softmax calculation on the attention scores after masking processing to obtain the normalized attention weights of each target ship:

[0120] (10)

[0121] And satisfy 。

[0122] Step S3.8: Use the attention weights calculated in Step S3.7 to perform dynamic weighted fusion on the target ship features

[0123] (11)

[0124] Step S3.9: Combine the dynamically fused target ship features with the own ship feature Q to form the input of the collision avoidance model:

[0125] (12)

[0126] For Step S4, the collision avoidance model is a multi-layer fully connected neural network, which consists of 4 layers, and each layer has 256 nodes. The ReLU activation function is used after each layer. As a preferred solution of this embodiment, the specific process of using this model to calculate the collision avoidance decision is as follows:

[0127] Step S4.1: Input the calculated in Step S3 into the collision avoidance model.

[0128] Step S4.2: Use the policy network in SAC to calculate the collision avoidance actions of the unmanned ship. The policy network takes as the input, calculates the mean and standard deviation of the actions through a multi-layer neural network, and obtains the expected action 。

[0129] Step S4.3: The action output is mapped and converted into the thrust and torque required for the unmanned boat to avoid collision.

[0130] For Step S5, the control mechanism of the unmanned boat generally consists of a propeller and a steering gear. It is necessary to map the thrust T and torque calculated in Step S4 to the rudder angle and propeller speed commands. The specific steps are as follows:

[0131] Step S5.1: Convert the thrust into the propeller speed command , and its specific calculation formula is as follows:

[0132] (13)

[0133] where ρ is the seawater density, D is the propeller diameter, is the thrust coefficient obtained from the propeller characteristic curve, are the maximum safe rotational speed and the minimum rotational speed of the propeller respectively.

[0134] Step S5.2, convert the moment into a rudder angle command , and its specific calculation formula is as follows:

[0135] (14)

[0136] where k is the coefficient between the moment and the rudder angle, is the maximum rudder angle change.

[0137] For step S6, in the actual navigation environment, the state characteristics such as the number, position, speed, and course of the target (TS) change at any time. The vessel needs to collect the latest navigation state data in real time through sensors such as GNSS and millimeter-wave radar, and continuously calculate and execute collision avoidance decision-making actions to adapt to the dynamic changes of the actual navigation environment until it sails smoothly to the target point.

[0138] This intelligent collision avoidance method realizes the organic integration of the Actor and Critic networks in the SAC deep reinforcement learning framework by introducing a masked attention mechanism. The masked attention mechanism can dynamically adapt to the change in the number of target vessels, automatically filter invalid or redundant information, provide more accurate and efficient policy guidance for the Actor network, and construct a more accurate state value estimation for the Critic network, significantly improving the adaptability of the unmanned vessel in a dynamic and complex environment and effectively alleviating the problem of decision-making performance fluctuations caused by the fixed input dimension in traditional methods. In addition, the masked policy adaptively highlights key information during the decision-making process, thus ensuring the accuracy and robustness of the collision avoidance decision and further improving the stability and safety of the unmanned vessel's navigation.

[0139] When the unmanned vessel autonomously sails in a complex and dynamic marine environment, it needs to actively cope with the navigation risks caused by other vessels. Traditional collision avoidance algorithms are difficult to handle diverse dynamic navigation scenarios. SAC can continuously interact and learn with the environment to optimize the collision avoidance strategy in real time, and has strong dynamic decision-making and generalization capabilities. However, when solving the collision avoidance decision, it is necessary to consider the dynamic change in the number of the other vessel. By introducing a masked attention mechanism, feature extraction and adaptive fusion of the information of the dynamically changing target vessels are carried out, thus effectively solving the problems of non-fixed model input dimension, redundant interference of target vessel information, and insufficient feature effectiveness caused by the change in the number of target vessels, and improving the accuracy and robustness of the collision avoidance strategy decision of the unmanned vessel in a complex and dynamic marine environment.

[0140] The model architecture generated based on the training process is asFigure 3 As shown in Figure 3 , the core modules (masked attention, Actor / Critic networks) optimize their parameters through the following training steps and are finally deployed into the real-time collision avoidance decision-making process (Steps S1 - S6):

[0141] Step 1: Construct the training environment.

[0142] Step 2: Construct the state space.

[0143] Step 3: Construct the action space.

[0144] Step 4: Design the reward function.

[0145] Step 5: Construct the Actor network.

[0146] Step 6: Construct the Critic network.

[0147] Step 7: Train the collision avoidance model.

[0148] Step 8: Evaluate the model performance.

[0149] For Step 1, constructing the training environment is the basis for training the collision avoidance model of the unmanned boat, aiming to provide effective data samples for the reinforcement learning algorithm. By reasonably constructing the simulation environment, the stability of model training can be ensured, and the generalization performance of actual navigation can be improved. As the preferred solution of this embodiment, the specific construction method is as follows:

[0150] Step 1.1: Design the size range of the simulation scenario.

[0151] Step 1.2: Define the dynamic model of the unmanned boat in the training environment. A 7-meter ship model scaled down proportionally based on the KVLCC2 ship type is used as the simulation dynamic model of the unmanned boat. The three-degree-of-freedom MMG motion model is specifically used to accurately describe the motion state and dynamic characteristics of the unmanned boat. This model can relatively accurately reflect the actual ship motion state. The model is shown as follows:

[0152] (15)

[0153] Where respectively represent the pitch, roll, and yaw rates of the unmanned boat. respectively represent the longitudinal force, lateral force, and yaw moment, and the subscript respectively represent the relevant component forces and moments generated by the hull hydrodynamic force and the rudder. represents the moment of inertia of the ship.

[0154] Step 1.3: To improve the generalization performance of the model, the navigation state of the own ship, the position of the target point, the number of other ships, and the navigation state are randomly generated each time training is performed. Randomly generating the target ship state diversifies the training environment, enabling the model to handle actual complex environments.

[0155] Step 1.4: Design the maximum number of training steps to be 150 steps.

[0156] Step 1.5: Design the termination conditions. Training terminates when the following three conditions are met: 1. The own ship intrudes into the ship domain of other ships; 2. The maximum number of training steps is reached; 3. The target point is successfully reached.

[0157] For Step 2, the state space is an important part of the training process and directly determines the decision-making performance of the algorithm. Reasonably constructing the state space can not only improve the learning efficiency of the model but also significantly enhance the effectiveness and generalization ability of the algorithm. The following points need to be noted when constructing the state space: 1. Try to include environmental features that have a significant impact on decision-making, avoid information redundancy or omission, to improve the learning efficiency of the model; 2. Fully consider the dynamic changes in the number of target ships in the actual navigation environment, and use the masked attention mechanism to process state information, so as to ensure that the model has good adaptability to the changing number of target ships; 3. Normalization processing is required to eliminate the impact of different data dimensions on model training, thereby ensuring a stable training process, fast policy convergence, and stable and reliable performance. The state space constructed in the present invention is consistent with the masked attention input constructed in step S2.4 of the technical solution.

[0158] For Step 3, the action space determines the specific collision avoidance behavior forms that the unmanned boat can take. The designed action vector directly reflects the control variables when the unmanned boat executes the collision avoidance task, where represents the thrust, represents the torque. During the training process and are both values between [-1, 1], and need to be mapped through Equation (16) to convert them into specific thrust T and torque M before being input into Equation (14):

[0159] (16)

[0160] where respectively represent the minimum thrust and maximum thrust that the control mechanism of the unmanned boat can provide. respectively represent the minimum torque and maximum torque that the control mechanism of the unmanned boat can provide.

[0161] For step 4, the reward function directly affects whether the algorithm can learn an effective collision avoidance strategy. When designing the reward function, multiple factors such as the safety, navigation efficiency, and collision avoidance rules in the unmanned boat collision avoidance task are comprehensively considered. In the preferred solution of this embodiment, a comprehensive reward function is provided to effectively guide the training direction of the reinforcement learning model. The reward calculation function is shown as follows:

[0162] (17)

[0163] where represent the coefficients of distance reward, heading reward, collision risk penalty, collision penalty, and collision avoidance rule violation penalty respectively, and the weights are set to [1, 1, 1.5, 2.1, 1]. represent distance reward, heading reward, collision risk penalty, collision penalty, and collision avoidance rule violation penalty respectively. The specific meanings of each sub-item are as follows:

[0164] Distance reward term ( )

[0165] This term is mainly used to guide the unmanned boat to quickly and effectively move towards the target point, and its calculation is shown as follows:

[0166] (18)

[0167] where 、 represent the Euclidean distances from the unmanned boat to the target point at the previous moment and the current moment respectively. represents the current speed of the unmanned boat. represents the simulation time step.

[0168] Heading reward term ( )

[0169] This term is mainly used to guide the unmanned boat to head towards the target point as much as possible, and its calculation is shown as follows:

[0170] (19)

[0171] Collision risk penalty term ( )

[0172] This term is used to punish the unmanned boat for taking collision avoidance actions that increase navigation risks, and its calculation is shown as follows:

[0173] (20)

[0174] where represents the number of other boats encountered by the unmanned boat. represents the collision risk formed by the own boat and the i-th other boat.

[0175] Collision penalty term ( )

[0176] This term is used to guide the unmanned boat to reach the target point as soon as possible. When the unmanned boat successfully sails to the target point area, a positive reward is obtained; if it collides with the target ship or enters the ship domain of the target ship, a negative reward is given. Its calculation is shown in the following formula:

[0177] (21)

[0178] Where , , and respectively represent the distance between the unmanned boat and the target point, the distance between the unmanned boat and the i-th opposing ship, the distance threshold for judging whether the target point is reached, and the radius of the ship safety domain of the i-th opposing ship.

[0179] Collision avoidance rule violation penalty term ( )

[0180] This term is used to guide the unmanned boat to take collision avoidance measures in line with the International Regulations for Preventing Collisions at Sea. When the own ship encounters an opposing ship and forms a right crossing and head-on situation with it, the own ship needs to take a right turn collision avoidance measure and cannot change the speed. In addition, the own ship can take any collision avoidance measures to ensure navigation safety. Its calculation is shown in the following formula:

[0181] (22)

[0182] For step 5, the constructed Actor network adopts a fully connected structure of a multi-layer perceptron, including four hidden layers, each layer is set to 256 neurons, and the ReLU activation function is used to improve the network's learning ability for non-linear features. The input of the Actor network is the state features fused by the masked attention mechanism, and the output is the Gaussian distribution parameters of the collision avoidance action. Subsequently, the reparameterization sampling method is used to obtain the specific action to ensure the smoothness and robustness of the policy output action.

[0183] For step 6, the constructed Critic network adopts a double Q network structure to reduce the overestimation bias of value estimation, thereby improving the training stability. Each Critic network is a fully connected structure of a multi-layer perceptron, including four hidden layers, each layer is set to 256 neurons, and the ReLU activation function is used to improve the non-linear fitting ability of the network. The network input includes the state features fused by the masked attention mechanism and the action vector, and the output is the Q value of the current state-action pair, which is used to evaluate the long-term return expectation of the action policy in the current state and guide the Actor network to optimize the decision-making policy. In addition, a soft-updated target Critic network is also adopted in SAC to further enhance the convergence stability of the model.

[0184] For step 7, the SAC algorithm is used to train the intelligent collision avoidance model of the unmanned boat. This algorithm is an off-policy reinforcement learning algorithm based on the principle of maximum entropy, with advantages such as good policy stability, high training efficiency, and strong generalization performance. In this step, the model is trained through real-time interaction data in the simulation environment, and the parameters of the Actor network and the Critic network are continuously optimized to enable the unmanned boat to learn the optimal collision avoidance strategy. As an optimal solution of this embodiment, the specific training steps are as follows:

[0185] Step 7.1, interact with the simulation environment through the current Actor network to obtain sample data And store the above interaction data in the experience replay pool. respectively represent the current state, the currently executed action, the reward obtained by the current action, the new state after executing the action, and the termination situation.

[0186] Step 7.2, randomly extract a batch of data from the experience replay pool, and use the mean square error loss function to optimize the parameters of the Critic network. The objective function for updating the Critic network is as follows:

[0187] (23)

[0188] where the target Q value is defined as:

[0189] (24)

[0190] In the formula is the target Critic network parameter, is the discount factor, is the entropy temperature coefficient, is the current Actor network policy. Update the Critic network parameters through the Adam optimizer .

[0191] Step 7.3, update the Actor network parameters. The Actor network objective function is to maximize the entropy-regularized value of the policy:

[0192] (25)

[0193] In the formula is the Q value estimated by the target Critic network. Update the Actor network parameters through the Adam optimizer to ensure a balance between high expected rewards and policy entropy.

[0194] Step 7.4, update the entropy temperature coefficient To maintain the exploratory nature of the algorithm strategy, it is necessary to adaptively adjust the entropy temperature coefficient, and its objective function is as follows:

[0195] (26)

[0196] In the formula is the entropy value of the target policy, which is adaptively updated through gradient descent to ensure that the policy is not overly deterministic or overly random.

[0197] Step 7.5, softly update the target Critic network. To ensure the stability of training, the parameters of the target Critic network are smoothly updated through soft update:

[0198] (27)

[0199] In the formula is generally taken as a small value (such as 0.005) to ensure the smooth change of the target network parameters and further improve the stability and convergence of the algorithm.

[0200] Step 7.6, continuously repeat the training loop composed of steps 7.1 to 7.5 until the cumulative number of training rounds reaches the set maximum number of training rounds and stop training.

[0201] For step 8, based on the simulation environment constructed in step 1, a more complex simulation evaluation scenario is further designed, that is, increasing the number and diversity of target ships to evaluate and verify the generalization performance of the trained intelligent collision avoidance model in a complex dynamic marine environment. During the performance evaluation process, the following key performance indicators are focused on: collision avoidance success rate, time to reach the target, smoothness of actions, and the adaptive ability of the model to the complex dynamic scenarios of multiple target ships. By designing a navigation scenario in the simulation environment that is more challenging and contains a larger number of target ships than during training, the trained model is verified and tested, so as to confirm that the trained collision avoidance decision-making model can effectively cope with a more complex real marine environment and meet the safety requirements of the actual unmanned boat navigation task.

[0202] In summary, the unmanned boat intelligent collision avoidance decision-making method based on the masked attention mechanism and SAC proposed in the embodiment of the present invention. The masked attention mechanism is used to effectively handle the problem of the change in the input feature dimension caused by the dynamic change of the number of target ships, significantly improving the adaptability and decision-making efficiency of the unmanned boat to the complex dynamic navigation environment, and providing a new solution for the intelligent autonomous collision avoidance of the unmanned boat in the complex marine environment.

[0203] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.

[0204] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0205] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0206] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

[0207] The present invention is not limited to the above best mode. Anyone inspired by the present invention can derive various other forms of ship tracking control methods with predefined time-regulated performance. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. An intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC, characterized in that: include: Dynamic target ship information processing: Adaptive feature fusion is performed on the variable-length target ship information acquired in real time through the mask attention mechanism. The mask attention weight is dynamically generated based on the real-time collision risk score of the target ship. Invalid target ship data is shielded through padding or truncation strategies to generate fusion state features with unified dimensions. SAC framework collaborative optimization: Inputting the fusion state features into the Actor network to generate a continuous collision avoidance action strategy; The fused state features are input into the Critic network, and the state-action value function is estimated through the double Q network structure combined with the entropy regularization term to balance the strategy exploration and utilization; End-to-end decision execution: Optimize network parameters during the training phase and output collision avoidance control instructions in real time during the deployment phase; The generation rules of the mask include: Completely shielded processing: When the number of target ships exceeds the preset threshold, the N target ships with the largest collision risk scores are selected; Complete shielding is achieved by filling invalid target ships with 0 or truncating redundant targets; Gradient retention mechanism: During the training phase, a minimum value is imposed on the attention scores corresponding to invalid target ships to preserve gradient propagation.

2. According to claim 1, an unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The construction of the fusion state feature includes: Change own ship status to OS With the target ship status s TS Separate processing, generating query vector Q and key-value vector {K, V} through independent linear transformation; The query vector Q and each target ship key vector K i After splicing, input the additive attention network to generate dynamic attention weight α i ; According to the dynamic attention weight α i The value vector V i Weighted summation to generate the target ship fusion feature f TS and concatenated with the ship's feature Q to form the final input vector f in .

3. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 2 is characterized in that: The input of the Critic network is the fusion state feature f generated by the mask attention mechanism in , the value function is optimized by combining the double Q network with the entropy term. The target Q value calculation combines the state features generated by the masked attention mechanism and the negative logarithm of the policy entropy to dynamically balance exploration and utilization; the Actor network updates the policy parameters by maximizing the entropy regularization value after weighting the masked attention, and the Critic network parameters are soft-updated through the mixing coefficient τ.

4. According to claim 1, the unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The output of the collision avoidance action strategy includes: Convert the normalized action parameters output by the Actor network into thrust and torque through linear mapping; The thrust is converted into a speed command according to the propeller dynamics model, and the torque is converted into a rudder angle command according to the servo model.

5. According to claim 1, the unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The reward function in the training phase is a multi-objective weighted sum, including: Distance bonus: guide the unmanned boat to quickly approach the target point; Heading bonus: Keep heading towards the target point; Collision risk penalty item: The dynamic threat score is calculated based on the closest encounter distance and encounter time between the target ship and the unmanned boat; Compliance penalty: Enforcement of turning and speed constraints of the International Regulations for Preventing Collisions at Sea.

6. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The priority rule of the dynamic target ship information is: when the number of detected target ships exceeds a preset threshold, select N target ships with the greatest collision risk calculated based on the closest encounter distance and the closest encounter time.

7. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The Critic network input of the SAC framework is the fusion state feature generated by the mask attention mechanism.

8. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The unmanned boat dynamics model in the training stage is a three-degree-of-freedom MMG model, including surge, sway and bow pitch motion equations.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • A Collision Avoidance Method for Swarm Unmanned Surface Vessels Based on Deep Reinforcement Learning

    CN110658829B

  • Unmanned ship multi-ship collision avoidance decision-making method and system based on reinforcement learning

    CN116954232A

  • Deep reinforcement learning collaborative navigation method for multiple unmanned boats based on proximal strategy optimization

    CN117168468B

  • Unmanned ship collision avoidance model construction method and device and unmanned ship collision avoidance method and device

    CN117523925A

  • Ship collision risk assessment method based on attention mechanism

    CN113962153A