Unmanned ship intelligent collision avoidance method based on mask attention mechanism and SAC

Through the combination of mask attention mechanism and SAC framework, the problem of insufficient adaptability caused by dynamic changes in the number of target ships in the unmanned boat collision avoidance method is solved, and the stable collision avoidance decision and safe navigation of unmanned boats in complex environments is achieved.

CN120010526AActive Publication Date: 2025-05-16JIMEI UNIV

Patent Information

Application Number
CN202510484418.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing unmanned boat collision avoidance method based on SAC algorithm has problems of insufficient adaptability and generalization capabilities when facing the dynamically changing target ship count, resulting in unstable collision avoidance strategies and may cause safety hazards.

Method used

The mask attention mechanism is used to adaptively process the variable-length target ship information, combined with the SAC framework for reinforcement learning, and dynamically generate mask attention weight to block invalid target ship data, generate unified fusion state characteristics, and achieve continuous collision avoidance action output through Actor and Critic network optimization strategies.

Benefits of technology

It improves the adaptability and safety of unmanned boats in complex dynamic environments, ensures the stability and accuracy of decision-making, complies with international maritime collision avoidance rules, and improves navigation safety and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010526A_ABST
    Figure CN120010526A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned ship intelligent collision avoidance method based on a mask attention mechanism and SAC, and the method comprises the steps: dynamic target ship information processing: carrying out the adaptive feature fusion of variable-length target ship information obtained in real time through the mask attention mechanism, and dynamically generating a mask attention weight based on the real-time collision risk score of a target ship, invalid target ship data are shielded through a filling or truncation strategy, and fusion state features with unified dimensions are generated; sAC framework collaborative optimization: inputting the fusion state features into an Actor network to generate a continuous collision avoidance action strategy; inputting the fusion state features into a Critic network, estimating a state-action value function through a double-Q network structure in combination with an entropy regularization item, and balancing strategy exploration and utilization; and end-to-end decision execution: optimizing network parameters through a training stage, and outputting a collision avoidance control instruction in real time in a deployment stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of unmanned boat collision avoidance, reinforcement learning, etc., and specifically relates to an unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC. Background Art

[0002] As the core carrier of intelligent marine equipment, unmanned surface vehicles (USVs) play an irreplaceable role in the fields of marine environmental monitoring, maritime search and rescue, port inspection and military reconnaissance. However, with the increasing complexity of marine missions, traditional collision avoidance methods (such as artificial potential field method and dynamic window method) have gradually exposed their limitations such as poor environmental adaptability and insufficient computational real-time performance in scenarios such as dynamic obstacle interaction and multi-target collaborative collision avoidance. In recent years, collision avoidance methods based on deep reinforcement learning (DRL) have provided new solutions to the USV collision avoidance problem with their adaptability to complex dynamic environments and end-to-end decision-making advantages, and have become a research hotspot in this field.

[0003] At present, USV collision avoidance algorithms based on DRL are mainly divided into two categories: Value-based algorithms (such as DQN, Double DQN): This type of algorithm learns the state-action value function through the Q-learning framework to maximize the long-term reward. Its advantage is that the theoretical convergence is clear, but there is an overestimation bias problem, and it is limited by the discrete action space, which makes it difficult to meet the control requirements of USV continuous steering and speed change.

[0004] Policy-based algorithms (such as PPO, SAC): This type of algorithm directly optimizes the policy function and updates the parameters through policy gradients, avoiding the complexity of value function calculation. Policy-based algorithms support continuous action output and show greater flexibility and robustness in complex and continuous environments. They are especially suitable for high-dimensional and dynamic state spaces in unmanned boat collision avoidance tasks.

[0005] However, strategy-based algorithms still face challenges in practical applications. When the number of target ships changes dynamically, traditional methods need to fix the state space dimension, resulting in deficiencies in the algorithm's generalization and adaptability. Existing technologies usually use padding or truncation strategies to handle variable numbers of target ships, but padding may introduce redundant noise and reduce training efficiency; truncation may lead to the loss of key obstacle information and increase navigation risks.

[0006] Related prior art The prior art closest to the present invention includes: The patent with authorization announcement number CN117168468B proposes a collaborative navigation method for multiple unmanned boats based on proximal strategy optimization; The patent with authorization announcement number CN110658829B discloses a ship collision avoidance decision-making method based on generative adversarial imitation learning; Patent publication number CN117523925A designs a collision avoidance method for swarm unmanned boats that combines deep reinforcement learning with LSTM neural networks; Patent publication number CN116954232A proposes a multi-ship collision avoidance decision-making system based on reinforcement learning.

[0007] None of the above technologies effectively solves the contradiction between the dynamic change of the number of target ships and the fixed state space requirements of the SAC algorithm, resulting in limited adaptability and generalization of the collision avoidance strategy. In actual navigation, the number of target ships changes dynamically with the sea area, time and navigation conditions, while existing methods usually only model a fixed number of target ships and adapt the input dimensions by padding or truncating. This processing method has significant defects: filling redundant information may interfere with model training, and truncating key target ship information may cause safety hazards in high-density navigation scenarios. Summary of the invention

[0008] In order to solve the problem that the existing solutions do not take into account the dynamic changes in the number of target ships and the fixed state space dimension constraints of the SAC algorithm, resulting in insufficient adaptability and poor generalization of the collision avoidance strategy of the unmanned boat in complex sea conditions, and the inability to cope well with the dynamic collision avoidance requirements in different density navigation environments, the present invention proposes an unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC. By using the masked attention mechanism to adaptively process the variable-length target ship information, and combining the SAC algorithm for reinforcement learning training, it can effectively solve the problems that the existing methods cannot dynamically adapt to the changes in the number of target ships, the information loss caused by the filling or truncation strategy, and the fixed state dimension constraints affect the decision stability, thereby improving the intelligence level and navigation safety of the unmanned boat collision avoidance system.

[0009] Its core design is: An intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC, comprising: Dynamic target ship information processing: Adaptive feature fusion is performed on the variable-length target ship information acquired in real time through the mask attention mechanism. The mask attention weight is dynamically generated based on the real-time collision risk score of the target ship. Invalid target ship data is shielded through padding or truncation strategies to generate fusion state features with unified dimensions. SAC framework collaborative optimization: Inputting the fusion state features into the Actor network to generate a continuous collision avoidance action strategy; The fused state features are input into the Critic network, and the state-action value function is estimated through the double Q network structure combined with the entropy regularization term to balance the strategy exploration and utilization; End-to-end decision execution: Optimize network parameters during the training phase and output collision avoidance control instructions in real time during the deployment phase.

[0010] The core design of the present invention is to embed the mask attention mechanism into the SAC framework to solve the input dimension conflict caused by the change in the number of dynamic target ships. Figure 3 As shown, the fusion state feature (f in ) is processed by the mask attention generation module (step S3, see the embodiment section below, the same below), and then input to the Actor network (step S4.2) and the Critic network (step 6). In the training phase (steps 1-8), the model optimizes the policy network parameters through entropy regularization and combines the dynamic mask mechanism to achieve robust learning; in the deployment phase (steps S1-S6), the trained Actor network is loaded to output collision avoidance actions in real time. This design is verified to be effective in the "Specific Implementation Method" through dynamic priority rules (step S2.4) and end-to-end decision-making process (step S5).

[0011] Among them, the mask attention generation module: corresponds to steps S3.2-S3.9 ( Figure 2 Process), through the separation feature fusion (the state of the own ship and the state of the target ship are processed independently) to generate the fusion feature f in ; Actor-Critic Co-Optimization: Step S4.2 (Actor Network) and Steps 7.2-7.5 (Critic Network) explicitly input f in , balance exploration and utilization through entropy regularization (Equation 24-Equation 27).

[0012] End-to-end decision-making process: Step S5 (action mapping) and step S6 (real-time control) verify the effectiveness of the deployment phase.

[0013] Furthermore, the mask generation rule includes: Completely shielded processing: When the number of target ships exceeds the preset threshold, the N target ships with the largest collision risk scores are selected; Complete shielding is achieved by filling invalid target ships with 0 or truncating redundant targets; Gradient retention mechanism: During the training phase, a minimum value is imposed on the attention scores corresponding to invalid target ships to preserve gradient propagation.

[0014] Among them, the training stage (gradient preservation): Step S3.6: "The attention score for the invalid target ship is set to 10 -9 , avoid participating in Softmax calculation but retain gradient propagation” (corresponding to Equation 9-Equation 10).

[0015] Deployment phase (completely shielded): Step S2.4: “When the number of target ships exceeds the threshold, fill in 0 or truncate the excess targets” (corresponding to the state construction logic).

[0016] Dynamic priority rules: Step S2.4: “Select N target ships with the greatest collision risk”, with scoring based on Equation 20.

[0017] Furthermore, the construction of the fusion state feature includes: Change own ship status to OS With the target ship status s TS Separate processing, generating query vector Q and key-value vector {K, V} through independent linear transformation; The query vector Q and each target ship key vector K i After splicing, input the additive attention network to generate dynamic attention weight α i ; According to the dynamic attention weight α i The value vector V i Weighted summation to generate the target ship fusion feature f TS and concatenated with the ship's feature Q to form the final input vector f in .

[0018] Furthermore, the input of the Critic network is the fusion state feature f generated by the mask attention mechanism. in , the value function is optimized by combining the double Q network with the entropy term. The target Q value calculation combines the state features generated by the masked attention mechanism and the negative logarithm of the policy entropy to dynamically balance exploration and utilization; the Actor network updates the policy parameters by maximizing the entropy regularization value after weighting the masked attention, and the Critic network parameters are soft-updated through the mixing coefficient τ.

[0019] Among them, the dual Q network design: Step 7.2: "The Critic network adopts a dual Q structure and optimizes the value function by minimizing the mean square error of the outputs of the two Q networks."

[0020] Entropy regularization term: Equation 24 (target Q value calculation) contains an entropy term, and the adaptive temperature coefficient α is optimized by Equation 26.

[0021] Soft update mechanism: Equation 27 (target network parameter update), mixing coefficient τ = 0.005.

[0022] Furthermore, the output of the collision avoidance action strategy includes: Convert the normalized action parameters output by the Actor network into thrust and torque through linear mapping; The thrust is converted into a speed command according to the propeller dynamics model, and the torque is converted into a rudder angle command according to the servo model.

[0023] Furthermore, the reward function in the training phase is a multi-objective weighted sum, including: Distance bonus: guide the unmanned boat to quickly approach the target point; Heading bonus: Keep heading towards the target point; Collision risk penalty item: The dynamic threat score is calculated based on the closest encounter distance and encounter time between the target ship and the unmanned boat; Compliance penalty: Enforcement of turning and speed constraints of the International Regulations for Preventing Collisions at Sea.

[0024] Furthermore, the priority rule of the dynamic target ship information is: when the number of detected target ships exceeds a preset threshold, select N target ships with the greatest collision risk calculated based on the closest encounter distance and the closest encounter time.

[0025] Furthermore, the Critic network input of the SAC framework is the fusion state feature generated by the masked attention mechanism.

[0026] Furthermore, the unmanned boat dynamics model in the training phase is a three-degree-of-freedom MMG model, including surge, sway and bow pitch motion equations.

[0027] And, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0028] A non-transitory computer-readable storage medium stores a computer program, which implements the steps of the method described above when executed by a processor.

[0029] Although the embodiment uses SAC as the core framework, those skilled in the art may consider replacing it with TD3 or PPO based on the universality of the Actor-Critic architecture. However, the SAC framework solution of the embodiment of the present invention is still the best implementation solution.

[0030] Compared with the prior art, the intelligent collision avoidance method of the present invention and its preferred solution achieves the following outstanding technical effects in dynamic and complex navigation scenarios by deeply integrating the masked attention mechanism with the SAC framework: Improved adaptability to dynamic environments: In the training phase, the gradient of invalid target ships is retained by minimizing the value, and the weight distribution rules of valid targets are learned; During the deployment phase, redundant information interference is avoided through complete shielding processing, and key threat targets are dynamically focused.

[0031] Enhanced decision-making efficiency and safety: The Actor network generates continuous collision avoidance actions based on linear mapping to achieve precise control of heading and thrust; The critic network optimizes value estimation through the double Q structure and entropy regularization term, balancing exploration and utilization.

[0032] Rules compliance guarantee: The reward function design forces the collision avoidance strategy to comply with the International Regulations for Preventing Collisions at Sea, improving navigation safety and compliance. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: Figure 1 The present invention is an unmanned boat collision avoidance decision flow chart of an embodiment of the present invention.

[0034] Figure 2 This is a flowchart of the masked attention mechanism processing according to an embodiment of the present invention.

[0035] Figure 3 This is a structural diagram of an unmanned boat intelligent collision avoidance algorithm based on masked attention mechanism and SAC in an embodiment of the present invention. DETAILED DESCRIPTION

[0036] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are specifically described in detail as follows: It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0038] The embodiment of the present invention proposes an intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC. The method adaptively models the variable-length target ship information by introducing masked attention mechanism, effectively alleviating the high dependence of the existing algorithm on the fixed state space dimension. By dynamically adjusting the attention weight, the method can maintain stable collision avoidance decision-making ability under different target ship numbers, and improve the adaptability and safety of unmanned boats in complex navigation environments.

[0039] The collision avoidance decision process provided in this embodiment ( Figure 1 ) is the online deployment application process of the trained model. Its core modules (mask attention, actor / critic network) optimize parameters through the following training process (steps 1-8): Step S1: collecting sensor data; Step S2: data preprocessing; Step S3: Use the masked attention mechanism to process the data again and calculate the collision avoidance model input; Step S4: collision avoidance model calculation decision; Step S5: The unmanned boat control system executes a collision avoidance decision; Step S6: Repeat steps S1-S5 until the collision avoidance behavior is completed.

[0040] For step S1, the types of information required in the unmanned ship collision avoidance task include the navigation status information of the own ship (OS) and the target ship (TS). Among them, the navigation status of the own ship is mainly obtained through the global navigation satellite system (GNSS), including the real-time position, speed and heading of the ship; the navigation status of the target ship is mainly obtained through millimeter wave radar detection, including the current position, speed and heading of the target ship. This information provides important data support for the dynamic collision avoidance decision of the unmanned ship.

[0041] For step S2, data preprocessing is an important step in achieving collision avoidance, and is intended to process the originally acquired navigation data to meet the requirements of the algorithm input. As a preferred solution of this embodiment, the specific steps are as follows: In step S2.1, a sliding window filtering method is used to identify and remove outliers in the position, speed, and heading information provided by GNSS and millimeter-wave radar to ensure data authenticity and reliability.

[0042] Step S2.2, convert the relative coordinate data obtained by the millimeter wave radar into longitude and latitude coordinates according to the longitude and latitude information of the ship obtained by GNSS. The conversion formula is: (1) Where r is the average radius of the Earth, and are the longitude and latitude positions of own ship and opponent ship respectively, It is the relative position of the other ship detected by the millimeter wave radar.

[0043] The Mercator projection is then used to convert the longitude and latitude positions of the own ship and the other ship into a Cartesian coordinate system for easy input into the subsequent calculation model. The Mercator projection formula is shown in equation (2): (2) In the formula is the central meridian accuracy, are the latitude and longitude with transformed position respectively.

[0044] Step S2.3, linear interpolation is used to align the GNSS and millimeter wave radar data to ensure that the timestamps of all input data remain consistent. The specific calculation formula is shown in the following formula (3): (3) Where t is the time to be interpolated, are two adjacent moments and satisfy , They are Through this interpolation method, the consistency of all sensor data at the same timestamp can be guaranteed.

[0045] Step S2.4, construct the masked attention input based on the processed data. The input consists of three parts: , the calculation formula is shown in formula (4): (4) In the formula They represent the OS's speed, heading and turning speed respectively. Indicates the relative direction and distance from OS to the target point. It is composed of the information of each target ship, and the i-th target ship The calculation formula is as follows: (5) in Represents OS and distance. Indicates OS to relative position and relative heading. It's the speed. Indicates OS and The encounter between Represents OS and risk of collision between them. yes In order to adapt to the existing mainstream code interface, in the calculation In the process of , a sufficiently large number of TS n will be defined. When the number of TS detected within the range of the OS sensor is less than n, it will be filled with 0; when the number of TS detected within the range of the OS sensor is greater than n, the n TS with the greatest collision risk will be selected for calculation. .

[0046] For step S3, the attention mechanism used is additive attention. By using the masked attention mechanism to dynamically weight multiple TS information, the collision avoidance model can automatically identify key target information, capture the state characteristics that best reflect the impact of TS on OS navigation decisions, and improve the effectiveness of collision avoidance decisions. The processing flow of this step is as follows: Figure 2 As shown, as a preferred solution of this embodiment, the specific implementation steps are as follows: Step S3.1, dividing the data s preprocessed in step S2 into two parts, namely, and the characteristics of the ship's status .

[0047] Step S3.2, Perform a linear transformation to obtain the query vector Q. The calculation formula is as follows: (6) in , is a trainable parameter, is the hidden feature dimension.

[0048] Step S3.3, for each target ship state feature Do a linear transformation to get the key vector Sum value vector ,in : (7) in , is a trainable parameter.

[0049] Step S3.4, copy and expand the query vector Q to After that, with the key vector Concatenate to form a comprehensive feature vector : (8) In step S3.5, the attention score of each target ship is calculated through an additive attention network (composed of a linear layer and a Tanh activation function), and the calculation formula is as follows: (9) in , , and These are all trainable parameters.

[0050] Step S3.6, using the mask attention mechanism, the attention score corresponding to the invalid target ship is set to , to avoid participating in Softmax calculation.

[0051] Step S3.7, perform Softmax calculation on the attention score after mask processing to obtain the normalized attention weight of each target ship: (10) And meet .

[0052] Step S3.8, using the attention weights calculated in step S3.7 , for the target ship characteristics Perform dynamic weighted fusion to form the fusion characteristics of the target ship: (11) Step S3.9, dynamically fused target ship features Combined with own ship characteristic Q, it forms the collision avoidance model input : (12) For step S4, the collision avoidance model is a multi-layer fully connected neural network, which includes 4 layers, and each layer has 256 nodes. ReLU activation function is used after each layer. As a preferred solution of this embodiment, the specific process of using this model to calculate the collision avoidance decision is as follows: Step S4.1: convert the Input into the collision avoidance model.

[0053] Step S4.2, use the policy network in SAC to calculate the collision avoidance action of the unmanned ship. As input, the mean and standard deviation of the action are calculated through a multi-layer neural network to obtain the expected action .

[0054] In step S4.3, the action output is converted into the thrust and torque required for the unmanned boat to avoid collision through mapping.

[0055] For step S5, the control mechanism of the unmanned boat is generally composed of a propeller and a steering gear, and the thrust T and torque calculated in step S4 need to be Map to rudder angle and propeller speed command. The specific steps are as follows: Step S5.1, converting the thrust into a propeller speed command , the specific calculation formula is as follows: (13) Where ρ is the seawater density, D is the propeller diameter, is the thrust coefficient obtained from the propeller characteristic curve, They are the maximum safe speed and minimum speed of the propeller respectively.

[0056] Step S5.2, convert the torque into a rudder angle command , the specific calculation formula is as follows: (14) Where k is the coefficient between torque and rudder angle, is the maximum rudder angle change.

[0057] For step S6, in the actual navigation environment, the number, position, speed, heading and other status characteristics of the target (TS) may change at any time. The ship needs to collect the latest navigation status data in real time through sensors such as GNSS and millimeter-wave radar, and continuously calculate and execute collision avoidance decision actions to adapt to the dynamic changes of the actual navigation environment until it can successfully navigate to the target point.

[0058] This intelligent collision avoidance method achieves an organic integration of the two by introducing the masked attention mechanism into the Actor and Critic networks of the SAC deep reinforcement learning framework. The masked attention mechanism can dynamically adapt to changes in the number of target ships, automatically filter invalid or redundant information, provide more accurate and efficient strategy guidance for the Actor network, and build a more accurate state value estimate for the Critic network, significantly improving the adaptability of the unmanned boat in dynamic and complex environments, and effectively alleviating the decision-making performance fluctuation problem caused by the fixed input dimension of the traditional method. In addition, the mask strategy adaptively highlights key information in the decision-making process, thereby ensuring the accuracy and robustness of the collision avoidance decision, and further improving the stability and safety of the unmanned boat's navigation.

[0059] When unmanned boats navigate autonomously in complex and dynamic ocean environments, they need to actively respond to navigation risks caused by other ships. Traditional collision avoidance algorithms are difficult to cope with diverse dynamic navigation scenarios. SAC can optimize collision avoidance strategies in real time through continuous interactive learning with the environment, and has strong dynamic decision-making and generalization capabilities. However, when solving collision avoidance decisions, it is necessary to consider the dynamic changes in the number of opposing ships. The masked attention mechanism is introduced to extract and adaptively fuse the dynamically changing target ship information, thereby effectively solving the problems of non-fixed model input dimensions, redundant interference of target ship information, and insufficient feature validity caused by changes in the number of target ships, and improving the accuracy and robustness of collision avoidance strategy decisions of unmanned boats in complex and dynamic ocean environments.

[0060] The model architecture generated based on the training process is as follows Figure 3As shown in the figure, its core modules (mask attention, actor / critic network) are optimized through the following training steps and finally deployed in the real-time collision avoidance decision process (steps S1-S6): Step 1: Build the training environment.

[0061] Step 2: Construct the state space.

[0062] Step 3: Construct the action space.

[0063] Step 4: Design the reward function.

[0064] Step 5: Build the Actor network.

[0065] Step 6: Build the Critic network.

[0066] Step 7: Train the collision avoidance model.

[0067] Step 8: Evaluate model performance.

[0068] For step 1, the construction of the training environment is the basis for the training of the unmanned boat collision avoidance model, which aims to provide effective data samples for the reinforcement learning algorithm. By reasonably constructing the simulation environment, the stability of the model training can be ensured and the generalization performance of actual navigation can be improved. As a preferred solution of this embodiment, the specific construction method is as follows: Step 1.1, design the size of the simulation scene.

[0069] Step 1.2, define the unmanned boat dynamics model in the training environment. A 7-meter ship model proportionally reduced from the KVLCC2 ship type is used as the simulation dynamics model of the unmanned boat. Specifically, a three-degree-of-freedom MMG motion model is used to accurately describe the motion state and dynamic characteristics of the unmanned boat. This model can more accurately reflect the actual ship motion state. The model is shown in the following formula: (15) in They respectively represent the pitch, roll and bow speeds of the unmanned boat. They represent longitudinal force, lateral force and yaw moment respectively, and the subscripts They represent the relevant components and moments generated by the hull hydrodynamics and the rudder respectively. Represents the ship's moment of inertia.

[0070] Step 1.3, in order to improve the generalization performance of the model, the navigation status of the ship, the location of the target point, the number and navigation status of the other ship are randomly generated during each training. The random generation of the target ship status makes the training environment more diverse, so that the model has the ability to cope with actual complex environments.

[0071] In step 1.4, the maximum training step length is designed to be 150 steps.

[0072] Step 1.5, design the termination conditions. The training terminates when the following three conditions are met: 1. The ship invades the enemy ship's area; 2. The maximum training step length is reached; 3. The target point is reached successfully.

[0073] For step 2, the state space is an important part of the training process and directly determines the decision-making performance of the algorithm. Reasonable construction of the state space can not only improve the learning efficiency of the model, but also significantly enhance the effectiveness and generalization ability of the algorithm. The following points should be noted when constructing the state space: 1. Try to include environmental features that have a significant impact on decision-making, avoid information redundancy or missing, so as to improve the learning efficiency of the model; 2. Fully consider the dynamic changes in the number of target ships in the actual navigation environment, and use the masked attention mechanism to process the state information, so as to ensure that the model has good adaptability to the changing number of target ships; 3. Normalization processing is required to eliminate the influence of different dimensions on model training between different data, so as to ensure the stability of the training process, rapid strategy convergence and stable and reliable performance. The state space constructed by the present invention is consistent with the masked attention input constructed in step S2.4 of the technical solution.

[0074] For step 3, the action space determines the specific collision avoidance behavior that the unmanned boat can take. The designed action vector Directly reflects the control variables when the unmanned boat performs the collision avoidance task, among which Indicates thrust, Indicates torque. During training , are all values ​​between [-1,1] and need to be mapped through formula (16) to convert them into specific thrust T and torque M before they can be input into formula (14): (16) in They respectively represent the minimum thrust and maximum thrust that the unmanned boat control mechanism can provide. They respectively represent the minimum torque and maximum torque that the unmanned boat control mechanism can provide.

[0075] For step 4, the reward function directly affects whether the algorithm can learn an effective collision avoidance strategy. When designing the reward function, multiple factors such as safety, navigation efficiency, and collision avoidance rules in the unmanned boat collision avoidance task are comprehensively considered. In the preferred solution of this embodiment, a comprehensive reward function is provided to effectively guide the training direction of the reinforcement learning model. The reward calculation function is shown as follows: (17) in They represent the coefficients of distance reward, heading reward, collision risk penalty, collision penalty and penalty for violation of collision avoidance rules respectively, and the weights are set to [1, 1, 1.5, 2.1, 1]. They represent distance reward, heading reward, collision risk penalty, collision penalty and violation of collision avoidance rules penalty respectively. The specific meanings of each sub-item are as follows: Distance Bonus ( ) This item is mainly used to guide the unmanned boat to quickly and effectively drive to the target point. Its calculation is shown in the following formula: (18) in , They represent the Euclidean distance from the unmanned boat to the target point at the previous moment and the current moment respectively. Indicates the current speed of the unmanned boat. Represents the simulation step size.

[0076] Heading Bonus Items ( ) This item is mainly used to guide the unmanned boat to head towards the target point as much as possible. The calculation is shown in the following formula: (19) Collision risk penalty ( ) This item is used to punish the unmanned boat for taking collision avoidance actions that increase navigation risks. The calculation is shown in the following formula: (20) in Indicates the number of enemy ships encountered by the unmanned boat. Represents the collision risk between the own ship and the i-th opponent ship.

[0077] Collision penalty term ( ) This item is used to guide the unmanned boat to reach the target point as quickly as possible. When the unmanned boat successfully navigates to the target point area, it will receive a positive reward; if it collides with the target ship or enters the target ship's area, it will receive a negative reward. Its calculation is shown in the following formula: (twenty one) in , , and They respectively represent the distance between the unmanned boat and the target point, the distance between the unmanned boat and the i-th opponent ship, the distance threshold for judging whether the target point has been reached, and the safety area radius of the i-th opponent ship.

[0078] Penalty items for violation of collision avoidance rules ( ) This item is used to guide the unmanned boat to take avoidance measures in accordance with the International Regulations for Preventing Collisions at Sea. When the ship encounters an opposing ship and forms a right crossing and encounter situation with it, the ship must take a right turn to avoid the situation and cannot change the speed. In addition, the ship can take any avoidance measures to ensure navigation safety. The calculation is shown in the following formula: (twenty two) For step 5, the constructed Actor network adopts a fully connected structure of a multi-layer perceptron, including four hidden layers, each layer is set to 256 neurons, and uses the ReLU activation function to improve the network's learning ability for nonlinear features. The Actor network input is the state features fused by the masked attention mechanism, and outputs the Gaussian distribution parameters of the collision avoidance action. Then, the reparameterized sampling method is used to obtain the specific action to ensure the smoothness and robustness of the strategy output action.

[0079] For step 6, the constructed Critic network adopts a double Q network structure to reduce the excessive deviation of value estimation, thereby improving training stability. Each Critic network is a fully connected structure of a multi-layer perceptron, containing four hidden layers, each layer is set to 256 neurons, and uses a ReLU activation function to improve the nonlinear fitting ability of the network. The network input includes state features and action vectors fused by the masked attention mechanism, and the output is the Q value of the current state-action pair, which is used to evaluate the long-term return expectation of the action strategy in the current state and guide the Actor network to optimize the decision-making strategy. In addition, a soft-updated target Critic network is also used in SAC to further enhance the convergence stability of the model.

[0080] For step 7, the SAC algorithm is used to train the unmanned boat intelligent collision avoidance model. This algorithm is an off-policy reinforcement learning algorithm based on the maximum entropy principle, which has the advantages of good strategy stability, high training efficiency and strong generalization performance. This step uses the real-time interactive data of the simulation environment to train the model, and continuously optimizes the parameters of the Actor network and the Critic network, so that the unmanned boat learns the optimal collision avoidance strategy. As the preferred solution of this embodiment, the specific training steps are as follows: Step 7.1: Obtain sample data by interacting with the simulation environment through the current Actor network And store the above interaction data in the experience replay pool. They represent the current state, the currently executed action, the reward obtained for the current action, the new state after executing the action, and the termination situation respectively.

[0081] Step 7.2, randomly extract batches of data from the experience replay pool and use the mean square error loss function to optimize the critic network parameters. The objective function used to update the critic network is as follows: (twenty three) The target Q value Defined as: (twenty four) In the formula is the target Critic network parameter, is the discount factor, is the entropy temperature coefficient, The current Actor network strategy. Update the Critic network parameters through the Adam optimizer .

[0082] Step 7.3, update the Actor network parameters. The Actor network objective function is to maximize the entropy regularization value of the strategy: (25) In the formula The Q value estimated by the target Critic network is optimized by the Adam optimizer for the Actor network parameters. Updates are made to ensure that the policy strikes a balance between high expected reward and policy entropy.

[0083] Step 7.4, update the entropy temperature coefficient In order to maintain the exploratory nature of the algorithm strategy, the entropy temperature coefficient needs to be adaptively adjusted, and its objective function is as follows: (26) In the formula is the target policy entropy value, which is adaptively updated by gradient descent , ensuring that the strategy is not over-determined or over-random.

[0084] Step 7.5, soft update the target critic network. To ensure the stability of training, the parameters of the target critic network are smoothly updated through soft update: (27) In the formula Generally, a small value (such as 0.005) is taken to ensure smooth changes in the target network parameters and further improve the stability and convergence of the algorithm.

[0085] Step 7.6, continue to repeat the training cycle formed by steps 7.1 to 7.5 until the cumulative number of training rounds reaches the set maximum number of training rounds and then stop training.

[0086] For step 8, based on the simulation environment constructed in step 1, a more complex simulation evaluation scenario is further designed, that is, the number and diversity of target ships are increased to evaluate and verify the generalization performance of the trained intelligent collision avoidance model in a complex dynamic ocean environment. During the performance evaluation process, the following key performance indicators are focused on: collision avoidance success rate, time to reach the target, smoothness of the action, and the model's ability to adapt to complex dynamic scenarios with multiple target ships. By designing a navigation scenario in the simulation environment that is more challenging than during training and contains a larger number of target ships, the trained model is verified and tested to confirm that the trained collision avoidance decision model can effectively cope with a more complex real ocean environment and meet the safety requirements of actual unmanned boat navigation missions.

[0087] In summary, the intelligent collision avoidance decision method for unmanned boats based on masked attention mechanism and SAC proposed in the embodiment of the present invention effectively handles the problem of input feature dimension change caused by the dynamic change of the number of target ships by using masked attention mechanism, significantly improves the adaptability and decision-making efficiency of unmanned boats to complex dynamic navigation environments, and provides a new solution for intelligent autonomous collision avoidance of unmanned boats in complex marine environments.

[0088] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0089] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by the processor to execute the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0090] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0091] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

[0092] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive other various forms of ship tracking control methods with predefined time specification performance under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.

Claims

1. An intelligent collision avoidance method for unmanned boats based on masked attention mechanism and SAC, characterized in that: include: Dynamic target ship information processing: Adaptive feature fusion is performed on the variable-length target ship information acquired in real time through the mask attention mechanism. The mask attention weight is dynamically generated based on the real-time collision risk score of the target ship. Invalid target ship data is shielded through padding or truncation strategies to generate fusion state features with unified dimensions. SAC framework collaborative optimization: Inputting the fusion state features into the Actor network to generate a continuous collision avoidance action strategy; The fused state features are input into the Critic network, and the state-action value function is estimated through the double Q network structure combined with the entropy regularization term to balance the strategy exploration and utilization; End-to-end decision execution: Optimize network parameters during the training phase and output collision avoidance control instructions in real time during the deployment phase.

2. According to claim 1, an unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The generation rules of the mask include: Completely shielded processing: When the number of target ships exceeds the preset threshold, the N target ships with the largest collision risk scores are selected; Complete shielding is achieved by filling invalid target ships with 0 or truncating redundant targets; Gradient retention mechanism: During the training phase, a minimum value is imposed on the attention scores corresponding to invalid target ships to preserve gradient propagation.

3. The unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC according to claim 1 or 2, characterized in that: The construction of the fusion state feature includes: Change own ship status to OS With the target ship status s TS Separate processing, generating query vector Q and key-value vector {K, V} through independent linear transformation; The query vector Q and each target ship key vector K i After splicing, input the additive attention network to generate dynamic attention weight α i ; According to the dynamic attention weight α i The value vector V i Weighted summation to generate the target ship fusion feature f TS and concatenated with the ship's feature Q to form the final input vector f in .

4. According to claim 3, the unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The input of the Critic network is the fusion state feature f generated by the mask attention mechanism in , the value function is optimized by combining the double Q network with the entropy term. The target Q value calculation combines the state features generated by the masked attention mechanism and the negative logarithm of the policy entropy to dynamically balance exploration and utilization; the Actor network updates the policy parameters by maximizing the entropy regularization value after weighting the masked attention, and the Critic network parameters are soft-updated through the mixing coefficient τ.

5. According to claim 1, the unmanned boat intelligent collision avoidance method based on masked attention mechanism and SAC is characterized in that: The output of the collision avoidance action strategy includes: Convert the normalized action parameters output by the Actor network into thrust and torque through linear mapping; The thrust is converted into a speed command according to the propeller dynamics model, and the torque is converted into a rudder angle command according to the servo model.

6. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The reward function in the training phase is a multi-objective weighted sum, including: Distance bonus: guide the unmanned boat to quickly approach the target point; Heading bonus: Keep heading towards the target point; Collision risk penalty item: The dynamic threat score is calculated based on the closest encounter distance and encounter time between the target ship and the unmanned boat; Compliance penalty: Enforcement of turning and speed constraints of the International Regulations for Preventing Collisions at Sea.

7. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The priority rule of the dynamic target ship information is: when the number of detected target ships exceeds a preset threshold, select N target ships with the greatest collision risk calculated based on the closest encounter distance and the closest encounter time.

8. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1 is characterized in that: The Critic network input of the SAC framework is the fusion state feature generated by the mask attention mechanism.

9. The method for intelligent collision avoidance of unmanned boats based on masked attention mechanism and SAC according to claim 1, characterized in that: The unmanned boat dynamics model in the training stage is a three-degree-of-freedom MMG model, including surge, sway and bow pitch motion equations.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • A Collision Avoidance Method for Swarm Unmanned Surface Vessels Based on Deep Reinforcement Learning

    CN110658829B

  • Unmanned ship multi-ship collision avoidance decision-making method and system based on reinforcement learning

    CN116954232A

  • Deep reinforcement learning collaborative navigation method for multiple unmanned boats based on proximal strategy optimization

    CN117168468B

  • Unmanned ship collision avoidance model construction method and device and unmanned ship collision avoidance method and device

    CN117523925A

  • Open water area ship autonomous collision avoidance method, system and device, and storage medium

    CN113744569A

Cited By

  • Water area collision danger identification and management method based on propagation dynamics

    CN120748256A

  • Unmanned aerial vehicle visual obstacle avoidance and autonomous navigation method based on improved PPO

    CN121477964A

  • Improved unmanned aerial vehicle vision obstacle avoidance and autonomous navigation method

    CN121477964B

  • Ship autonomous collision avoidance decision-making method and system based on cognitive entropy near-end strategy optimization

    CN121764176A

  • Unmanned ship robust collision avoidance decision-making method considering ship load change

    CN122195014A