Unmanned ship collision avoidance decision-making method based on navigation intention perception
By introducing a deep fusion of recursive Bayesian intention modeling and reinforcement learning in the unmanned boat collision avoidance method, the problem of difficult to model the uncertainty of the target ship's behavior is solved, and more stable and robust collision avoidance decision-making capabilities are achieved.
Patent Information
- Application Number
- CN202510696136.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing unmanned boat collision avoidance methods fail to effectively model the uncertainty of the target ship's behavior, resulting in the strategy lacking the ability to predict the target's future actions, and it is prone to problems such as lagging response and increasing risk of misjudgment.
Using a deep fusion of recursive Bayesian intention modeling and reinforcement learning, the target ship's historical trajectory is dynamically analyzed through recursive Bayesian estimation, discrete navigation intentions are multiple behavioral patterns, the intent posterior probability distribution is updated in real time, and it is used as input to the reinforcement learning strategy network for forward-looking decisions.
It significantly improves the collision avoidance decision-making ability and strategy stability of unmanned boats in the face of sudden change or non-cooperation goals, and enhances the adaptability and robustness to complex dynamic environments.
Smart Images

Figure CN120217907A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields such as unmanned ship control (unmanned ship collision avoidance, reinforcement learning), and particularly relates to a collision avoidance decision-making method for an unmanned boat based on navigation intention perception. Background Art
[0002] In recent years, autonomous collision avoidance methods based on Deep Reinforcement Learning (DRL) have become a research hotspot in the field of unmanned boat autonomous navigation. Such methods can achieve end-to-end collision avoidance decisions by interacting with the environment and using reward signals to guide policy learning, and have good environmental adaptability and policy optimization capabilities. Compared with traditional rule-based or optimization-based methods, DRL can not only handle complex non-linear decision-making problems, but also achieve better collision avoidance effects in dynamic scenarios with multiple objectives and multiple constraints, and has broad application prospects.
[0003] In the DRL framework, the construction of the state space plays a decisive role in the policy performance. In order to enable the unmanned boat to accurately perceive the environment and make reasonable decisions, researchers have proposed various state modeling methods. Common practices include: 1) a method based on risk ranking, which only selects several target ships with the highest threat level at the current moment to construct a state vector; 2) a method based on space partitioning, which divides the perception area into multiple sectors, and combines the method based on risk ranking to select the most representative target in each area; 3) a grid map representation method, which converts the perceived obstacle information into grid map data, thereby improving the policy's modeling ability for multi-target interaction scenarios.
[0004] However, existing methods generally only use the geometric and motion states (such as position, speed, heading, etc.) of target ships at the current moment as input features, lack the modeling of the changing trends of target behaviors, and ignore the navigation intention information contained in their historical trajectories. In the actual marine environment, target ships often have obvious behavioral continuity, and their future actions can be largely inferred from the dynamics in the past period of time. If only relying on the single-moment state, it is difficult for the reinforcement learning policy to accurately judge the target behavior pattern. Especially when facing sudden course changes or non-cooperative ships, it is easy to cause response lags and policy failures, increasing the collision risk. Therefore, it is urgent to introduce a dynamic reasoning mechanism for navigation intention at the state modeling level to characterize target behavior features at a higher level, so as to improve the stability of the unmanned boat collision avoidance policy. To solve this problem, the present invention proposes a collision avoidance decision-making method for an unmanned boat based on navigation intention perception. Design a navigation intention perception module and explicitly integrate the intention perception information into the state space, thereby effectively improving the collision avoidance decision-making ability of the unmanned boat in an environment with uncertain target ship behaviors and ensuring its safe navigation.
[0005] The invention with the application number CN202510296976.6, "Method and device for collision avoidance decision-making of ships in restricted waters based on uncertainty modeling".
[0006] The invention with the application number CN202210962382.0, "Intelligent ship collision avoidance path planning method based on uncertain velocity obstacles".
[0007] The invention with the application number CN202411845856.9, "Full-process path planning method for autonomous collision avoidance of intelligent ships".
[0008] The invention with the application number CN202411439576.8, "Method for constructing a highly reliable ship autonomous collision avoidance model".
[0009] The above solutions are regarded as relevant prior arts, but none of them fully consider the impact of the uncertainty of the target ship's navigation intention on the stability and safety of the decision-making strategy during the collision avoidance process of unmanned boats. Generally, the state space is constructed only based on the geometric and motion states of the target ship at the current moment, ignoring the behavioral trend information contained in its historical trajectory. In actual navigation, target ships will exhibit non-linear and variable behavioral characteristics, such as sudden acceleration and deceleration, and temporary steering, which are particularly significant when facing non-cooperative or small ships. Existing methods fail to effectively model the uncertainty of such behaviors, resulting in the lack of predictive ability of the strategy for the future actions of the target, and prone to problems such as response lag and increased risk of misjudgment, reducing the adaptability and robustness of the collision avoidance strategy in complex dynamic environments. Summary of the Invention
[0010] Aiming at the defects and deficiencies existing in the prior art, the present invention provides an unmanned boat collision avoidance decision-making method based on the deep integration of recursive Bayesian intention modeling and reinforcement learning. Its innovative design points include: dynamically analyzing the historical trajectory of the target ship through recursive Bayesian estimation, discretizing the navigation intention into 9 combined patterns of heading (left turn / straight / right turn) and speed (acceleration / constant speed / deceleration), and updating the posterior probability distribution of the intention in real time; constructing a multi-dimensional state vector including the state of the own ship, the motion characteristics of the target ship and the intention distribution, and driving the Soft Actor-Critic (SAC) policy network to make forward-looking decisions; screening threat targets in combination with the five-factor collision risk assessment model (DCPA / TCPA / relative heading / distance / speed), designing a composite reward function that combines distance reward, heading reward, international rule penalty and speed obstacle penalty, and forcing the policy to comply with the COLREGs specification; simulating the random behaviors of the target ship of "maintaining speed and course - collision avoidance - intentional collision" (probability distribution [0.3, 0.3, 0.4]) through a dynamic simulation environment, strengthening the robustness of the policy in non-cooperative scenarios, and realizing the systematic integration of multi-source perception (GNSS / radar / AIS) data and the Bayesian inference-reinforcement learning framework, effectively solving the problems of response lag and misjudgment caused by traditional methods ignoring the behavior trends of target ships.
[0011] The solution adopted by the present invention to solve its technical problems specifically includes:
[0012] An unmanned boat collision avoidance decision-making method based on navigation intention perception:
[0013] Real-time collecting the motion state data of the own ship and the target ship through multi-source sensors, and constructing a collision risk assessment model to screen threat targets;
[0014] Based on the recursive Bayesian inference of the historical trajectory of the target ship, dynamically updating the posterior probability distribution of the navigation intention of the target ship, and discretizing the navigation intention into multiple behavior patterns;
[0015] Constructing the state of the own ship, the motion characteristics of the target ship and the intention distribution into a state vector, inputting it into the policy network based on reinforcement learning, and outputting collision avoidance control actions;
[0016] Guiding the optimization of the policy network through a composite reward function, and the reward function combines distance reward, rule penalty and collision risk penalty.
[0017] Furthermore, the multiple behavior patterns of the navigation intention include the combination of left turn, straight, right turn of the heading and acceleration, constant speed, deceleration of the speed, a total of 9 patterns.
[0018] Furthermore, the recursive Bayesian inference includes:
[0019] Predicting the state based on the dynamic model of the target ship's surge acceleration and yaw acceleration;
[0020] Model the observation noise through Gaussian distribution and calculate the likelihood probability between the predicted state and the actual observation;
[0021] Recursively update the posterior probability distribution of the intention.
[0022] Furthermore, when constructing the state vector, the data of the target ship with the highest collision risk index is preferentially selected, and if it is insufficient, it is supplemented with a zero vector.
[0023] Furthermore, the composite reward function includes an international collision avoidance rule penalty term and a speed obstacle penalty term, and a linear penalty is triggered when the action violates the rules or enters the speed obstacle area.
[0024] Furthermore, the reinforcement learning-based policy network is of SAC architecture, and the output action is sampled from a Gaussian distribution through the reparameterization technique and mapped to the physical control range through the tanh function.
[0025] Furthermore, the method is trained in a dynamic simulation environment, and the behavior patterns of the target ships are randomly selected according to preset probabilities, including three strategies: maintaining speed and course, collision avoidance, and intentional collision.
[0026] Furthermore, the multi-source sensors include the Global Navigation Satellite System, radar, and Automatic Identification System, and the collected data is input into the model after preprocessing such as anomaly rejection, time synchronization, and coordinate unification.
[0027] Furthermore, the collision risk assessment model calculates the risk index through five-factor linear weighting, including relative distance, relative speed, relative course, minimum collision distance, and time to the closest point of approach.
[0028] And, a collision avoidance decision-making system for an unmanned surface vehicle based on navigation intention perception, including:
[0029] A multi-source perception module, used to collect the motion state data of the own ship and the target ships in real time;
[0030] A threat target screening module, used to construct a collision risk assessment model to screen threat targets;
[0031] An intention reasoning module, used for recursive Bayesian reasoning based on the historical trajectories of the target ships to dynamically update the posterior probability distribution of the navigation intentions of the target ships;
[0032] A policy network module, used to generate collision avoidance control actions according to the state of the own ship, the motion characteristics of the target ships, and the intention distribution;
[0033] A control execution module, used to convert the action instruction into a ship course and speed control signal.
[0034] In addition, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-described method are implemented.
[0035] A non-transitory computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the above-described method are implemented.
[0036] Compared with the prior art, the present invention and its preferred solutions at least include the following beneficial effects:
[0037] Through the deep integration of recursive Bayesian and reinforcement learning, the perception limitation of traditional collision avoidance methods for the behavior trend of target ships is broken through. Based on historical trajectories, the intention distribution is dynamically modeled, significantly improving the pre-judgment ability and decision-making real-time performance for complex scenarios such as sudden course changes and non-cooperative targets.
[0038] The design of a composite reward function integrates international collision avoidance rules, speed obstacle constraints, and a multi-dimensional risk index system, driving the policy network to achieve a dynamic balance among safety, compliance, and economy, and avoiding sub-optimal decisions caused by rule violations or risk misjudgments.
[0039] The construction logic of a dynamic simulation environment strengthens the generalization ability of the policy in uncertain interaction scenarios by simulating the random behavior patterns of target ships such as "maintaining speed and course - collision avoidance - intentional collision", and solves the problem of policy vulnerability caused by the single traditional training scenario.
[0040] A multi-source perception and five-factor risk assessment model combines GNSS, radar, and AIS heterogeneous data fusion and anomaly rejection mechanisms to improve the reliability of environmental perception. At the same time, it linearly weights and quantifies the collision risk to achieve efficient screening and priority ranking of threat targets.
[0041] The optimization of the SAC algorithm framework, based on the reparameterized action sampling and physical constraint mapping mechanism, takes into account both the efficiency of policy exploration and the engineering feasibility of control instructions, ensuring the smoothness and execution stability of collision avoidance actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0043] Figure 1 It is a flowchart of the collision avoidance decision-making of the unmanned boat in the embodiment of the present invention.
[0044] Figure 2 It is a schematic diagram of the navigation intention of the target ship in the embodiment of the present invention.
[0045] Figure 3 It is a schematic diagram of the generation process of the target ship in the embodiment of the present invention.
[0046] Figure 4 This is a schematic diagram of the algorithm structure according to an embodiment of the present invention. Specific embodiments
[0047] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail as follows:
[0048] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0050] To address the problem that the existing solution fails to fully model the uncertainty of the navigation intention of the target ship, resulting in a lag in the collision avoidance strategy response and insufficient robustness when the unmanned boat faces non-cooperative targets or dynamic behavior changes, and it is difficult to effectively cope with the high-dynamic and variable navigation environment, the present invention proposes an intelligent collision avoidance method for unmanned boats based on the fusion of Bayesian navigation intention reasoning and deep reinforcement learning. By introducing a recursive Bayesian estimation method to dynamically analyze the historical trajectory of the target ship, probabilistic modeling of the navigation intention is achieved, and it is explicitly embedded in the state space as a navigation feature and jointly trained with the deep reinforcement learning strategy, which can effectively improve the perception and prediction ability of the collision avoidance strategy for the target behavior trend, thereby enhancing the decision-making stability, safety, and generalization ability of the unmanned boat in complex and uncertain sea conditions.
[0051] For this reason, the embodiments of the present invention propose an intelligent collision avoidance method for unmanned boats based on Bayesian intention reasoning and the SAC algorithm. By introducing a recursive Bayesian method to dynamically reason about the navigation intention of the target ship and embedding the intention distribution as a navigation feature of the target ship in the state space, the reinforcement learning strategy is guided to make forward-looking decisions, thereby improving the collision avoidance ability and strategy stability of the unmanned boat in complex, dynamic, and uncertain environments.
[0052] The collision avoidance decision flowchart proposed by the embodiments of the present invention is as Figure 1 shown and mainly includes the following steps:
[0053] Step 1: Environment perception to obtain navigation information;
[0054] Step 2: Determine whether there are obstacle ships around the unmanned boat according to the data in Step 1;
[0055] Step 3: Determine whether the obstacle ship poses a collision risk to the own ship;
[0056] Step 4: If there is no collision risk, maintain the current navigation state; if there is a collision risk, proceed to Step 5;
[0057] Step 5: Infer the navigation intention of the obstacle ship;
[0058] Step 6: Calculate the model input;
[0059] Step 7: Make a collision avoidance decision;
[0060] Step 8: Execute the collision avoidance decision and determine whether the target point has been reached;
[0061] Step 9: If the target point has not been reached, repeat Steps 1 - 7 until the target point is reached.
[0062] For Step 1: By integrating multi - source perception devices such as the Global Navigation Satellite System (GNSS), radar sensors, and Automatic Identification System (AIS), the navigation state information of the own ship and the target data of the surrounding environment are collected in real - time. The collected data includes the position, speed, heading, and turning angular velocity of the own ship, as well as the relative position, relative speed, and heading angle and other motion characteristics of the surrounding targets. After pre - processing such as abnormal data rejection, time synchronization, and coordinate system unification, various types of data constitute the basic input information for obstacle recognition and collision avoidance decision - making, ensuring that the system has stable and reliable environmental perception capabilities in a dynamic and complex environment.
[0063] For Step 2: Based on the multi - source perception information obtained in Step 1, the target bodies existing in the environment are effectively identified to achieve the preliminary screening and classification of obstacles. The specific steps are as follows:
[0064] Step 2.1: After pre - processing the original perception data obtained by multi - source sensor systems such as GNSS, radar, and AIS, extract the basic navigation characteristics of each target body;
[0065] Step 2.2: Eliminate noise data and only retain the suspicious obstacle data with physical consistency and navigation characteristics;
[0066] Step 2.3: Extract the data of the target ship.
[0067] For Step 3: After the extraction of the target ship data is completed, the system conducts a collision risk assessment for all target ships to determine whether there is a potential threat to the own ship. The specific assessment method uses a typical five-factor collision risk assessment model, comprehensively considering factors such as the relative distance, relative speed, relative course, minimum distance to closest point of approach (DCPA), and time to closest point of approach (TCPA) between two ships, and calculates the collision risk index of the target ship. The overall risk index is calculated through the following linear weighted model:
[0068] (1)
[0069] In the formula, J is the weight vector of each factor, reflecting its relative contribution degree to the collision risk, and M is the fuzzy membership function of each risk factor, with a value range of [0, 1].
[0070] For Step 4: According to the collision risk index calculated in Step 3, determine whether there are target ships that pose a substantial threat to the own ship. When the risk assessment results of all target ships are equal to 0, the system determines that the current navigation environment is safe, there is no need to calculate collision avoidance decisions, and the unmanned boat controller maintains the current navigation state; on the contrary, if there is at least one target ship with a collision risk index greater than 0, then proceed to Step 5.
[0071] For Step 5: After there are target ships that pose a potential collision threat to the own ship, use the recursive Bayesian method to infer the navigation intention of this type of target ship. This method dynamically models the deviation between the historical motion state of the target ship and the current observation, thereby estimating the posterior probability distribution of its behavioral intention, and using this as part of the features to combine with the current navigation data to construct the input of the collision avoidance decision model. The specific steps are as follows:
[0072] Step 5.1: Initialize the prior intention distribution.
[0073] The present invention discretizes the navigation intention of the target ship into nine typical patterns, which respectively represent the combined actions of course (left turn / straight / right turn) and speed (acceleration / constant speed / deceleration), such as Figure 2 shown. The system assigns a uniform prior distribution to each target ship at the initial moment:
[0074] (2)
[0075] Step 5.2: Calculate the predicted state.
[0076] The observed state of the target ship at time t - 1 is , where (x, y) represents the position of the target ship, ψ is the course angle of the target ship, and (u, v, r) are the surge, sway, and yaw angular velocities of the target ship respectively. Calculate the predicted state The prediction model is shown as follows:
[0077] (3)
[0078] where a u and a r are the surge acceleration and yaw acceleration of the target ship respectively.
[0079] Step 5.3: Calculate the likelihood estimate.
[0080] Compare the predicted state with the actual observed state O of the target ship at the current moment t to calculate the likelihood probability of the current observation under the intention I i . Assuming that the observation noise follows a zero-mean Gaussian distribution, the likelihood function is expressed as:
[0081] (4)
[0082] where ∑ is the observation error covariance matrix and n = 6 is the state dimension.
[0083] Step 5.4: Intention posterior update.
[0084] Use Bayes' formula to update the posterior probability of each candidate intention. The calculation formula is as follows:
[0085] (5)
[0086] This update process is recursively carried out within each control cycle to form the posterior distribution vector of the target ship's behavior intention at the current moment .
[0087] For step 6: After completing the inference of the target ship's navigation intention, the system constructs the input vector of the collision avoidance strategy network according to the state information of the own ship and the target environment at the current moment. This input vector consists of four parts, including: the navigation state of the own ship, the navigation target information, the motion state of the target ship, and the navigation intention distribution of the target ship.
[0088] (6)
[0089] The navigation state of the own ship mainly consists of the surge u os , sway v os and yaw angular velocity r os of the own ship, as shown below:
[0090] (7)
[0091] The navigation target information mainly consists of the relative azimuth angle between the own ship and the target point and the relative course angle Composed as follows:
[0092] (8)
[0093] The motion state of the target ship consists of the relative distance , relative azimuth angle , relative course angle , collision risk index C RI , encounter situation C OL and the turning angular velocity r of the target ship TS and the speed V, and is composed as follows:
[0094] (9)
[0095] The intention distribution of the target ship is obtained by calculation in step 5, as follows:
[0096] (10)
[0097] When constructing the input vector, the navigation information and intention distribution of the target ship with the largest collision risk index are preferentially selected to construct the input; if the quantity is insufficient, it is filled with a zero vector to ensure that the input dimensions are consistent. All feature variables are normalized before input to improve the training stability and policy convergence efficiency.
[0098] For step 7: After the state vector is constructed, the system takes it as input and sends it to the deep reinforcement learning policy network to calculate the collision avoidance action to be executed at the current moment. Preferably, the policy network in the Soft Actor-Critic (SAC) architecture is used in this embodiment to generate actions. The policy network takes the state s as input and outputs a Gaussian distribution with a mean of μ(s) and a standard deviation of σ(s), representing the probability distribution of the actions that can be taken in the current state. Subsequently, the system samples actions from this distribution through the reparameterization technique:
[0099] (11)
[0100] where the action a represents the collision avoidance control instruction at the current moment, usually the rudder angle force for controlling the bow direction of the ship. Before execution, the action will be mapped through the tanh function to ensure that its range conforms to the physical limitations of ship control.
[0101] For step 8: After the collision avoidance control action output by the policy network is parsed, it is converted into the corresponding control command and transmitted to the underlying control system of the unmanned boat. The control system adjusts the course and speed of the ship accordingly to implement the collision avoidance decision at the current moment.
[0102] For step 9: After each control cycle, the system evaluates the distance between the own ship and the target point in real time based on the latest own ship position data. If this distance is less than the preset threshold, it is determined that the own ship has successfully reached the target point, and the collision avoidance process terminates; otherwise, the system continues to collect the environmental state, reconstructs the state vector, and enters the next round of collision avoidance decision-making loop.
[0103] In the collision avoidance method for unmanned surface vehicles based on Bayesian intention reasoning proposed in the embodiments of the present invention, aiming at the deficiencies of the existing collision avoidance strategies in single state space modeling and difficulty in characterizing the uncertainty of the target ship's behavior, a modeling mechanism for dynamically integrating the target ship's navigation intention distribution into the state space is proposed. By performing recursive Bayesian inference on the historical motion trajectory of the target ship, its potential navigation intention is speculated in real time, and the inference result is embedded as a navigation feature into the state input, effectively enhancing the state space's ability to express the target behavior trend.
[0104] By introducing intention modeling, the collision avoidance decision-making system of the present invention enables the unmanned surface vehicle not only to judge the environmental situation based on the current geometric information, but also to make forward-looking decisions by combining the development trend of the target behavior, thereby improving the reaction speed and decision-making accuracy of the strategy. Compared with the traditional method that relies on the current instantaneous state modeling, this technology significantly enhances the adaptability of the deep reinforcement learning strategy in a dynamic and variable environment, improves the convergence stability during the training process, and effectively improves the collision avoidance success rate and overall navigation safety during actual navigation.
[0105] To improve the collision avoidance ability of unmanned surface vehicles under the uncertain behavior of the target ship, the above embodiments propose a collision avoidance decision-making method that integrates a recursive Bayesian intention reasoning mechanism based on the SAC algorithm. Aiming at the problem of insufficient adaptability of traditional collision avoidance strategies when facing the uncertainty of the target ship's behavior, the navigation intention of the target ship is dynamically estimated through recursive Bayesian inference, and the intention posterior distribution is incorporated into the state space as an explicit feature, thereby enhancing the collision avoidance strategy's perception ability of the target behavior pattern and achieving a more robust collision avoidance decision.
[0106] To implement the training and verification of the above algorithm, the collision avoidance decision-making strategy proposed in this embodiment is implemented by constructing a simulation platform based on the Python language and the OpenAI Gym framework during the training phase. The platform simulates a two-dimensional sea area environment with a regional range set to 200 meters × 200 meters, a simulation step size of 1 second, and supports dynamic intersection scenarios where multiple target ships sail simultaneously. The unmanned surface vehicle is modeled using a three-degree-of-freedom (3-DOF) nonlinear dynamics model, and the state variables include position, surge velocity, heading, sway velocity, and yaw rate of turn. The control inputs are fixed thrust and adjustable yaw moment. The target ship uses the same dynamics model to ensure interaction consistency.
[0107] In the simulation environment, the unmanned surface vehicle adopts a three-degree-of-freedom kinematic model, and the model form is as follows:
[0108] (12)
[0109] Wherein, T and τ are the thrust and torque inputs respectively; (m1, m2) is the added mass, is the damping coefficient, and L is the moment of inertia.
[0110] To ensure numerical stability and simulation accuracy, the platform uses the fourth-order Runge-Kutta method to perform numerical integration calculations on the above differential equations.
[0111] To generate a training scenario with collision risks, a set of dynamic target ship (Target Ship, TS) generation logics is designed, as Figure 3 shown. The specific generation process is as follows:
[0112] Step A1: First, randomly determine the number of target ships and the encounter type with the own ship (OwnShip, OS) according to the preset probability distribution;
[0113] Step A2: For each target ship, calculate its possible course range according to the selected encounter type and randomly generate a course ;
[0114] Step A3: Randomly generate the speed of the target ship and the collision time with the own ship ;
[0115] Step A4: Randomly generate the collision time , and calculate the initial coordinates of the target ship according to the following equations :
[0116] (13)
[0117] Step A5: If it is the first target ship, directly retain it; otherwise, it is necessary to judge whether the distance from the existing target ships meets the safety threshold to avoid generating target ships with overly dense positions;
[0118] Step A6: Repeat steps A2 to A5 until the set number of target ships is met.
[0119] This TS generation mechanism is dynamically executed at the beginning of each training round to generate scenarios containing multiple target ships, high-risk encounters, behavioral diversity, and uncertainty for training the collision avoidance model.
[0120] In addition, to simulate the navigation uncertainty of the TS in the actual marine environment, the simulation platform divides the navigation behavior of the TS during the driving process into two types: collision avoidance behavior and random behavior. When encountering an obstacle, the TS with collision avoidance behavior will use the Velocity Obstacle (VO) method to calculate the collision avoidance decision and execute it. The TS with random behavior will randomly select a navigation strategy in the three behaviors of "maintaining speed and course", "collision avoidance", and "intentional collision" with probabilities [0.3, 0.3, 0.4] in real time to form an uncertain strategy.
[0121] To realize the efficient collision avoidance strategy training of the unmanned boat in a complex dynamic environment, the present invention designs a collision avoidance decision-making method integrating navigation intention perception based on the Soft Actor-Critic (SAC) reinforcement learning framework, and its overall structure is as Figure 4 shown. The algorithm includes a policy network and a value network. The policy network consists of a neural network to generate control actions. The policy network consists of two value evaluation networks and two corresponding target networks to evaluate the action value. In the design of the policy network, the state space input includes the navigation state of the own ship, the target point information, the motion characteristics of the surrounding target ships and their intention posterior distribution. The above inputs are uniformly encoded into a fixed-length vector after preprocessing and input into the policy network and the value network.
[0122] The policy network adopts a 4-layer fully connected structure, with 256 neurons in each layer, and the activation function is ReLU. The output layer generates the mean and standard deviation of the Gaussian distribution, and the action vector is sampled through the differentiable random reparameterization technique. Considering the actual control requirements, the action is mapped to the corresponding steering thrust to realize the steering control of the unmanned boat. All the neural network structures in the value network are similar to the policy, and also consist of 4 fully connected layers. The input is the vector after concatenating the state and the action, and the Q-value estimate of the state-action pair is output. To alleviate the problem of overestimation of the Q value, a double Q structure is adopted and equipped with a target network, and the parameters are updated through exponential moving average to improve the stability of policy training.
[0123] In the training stage, the Adam optimizer is used, the capacity of the experience replay pool is 200,000, the batch size is 256, the discount factor is 0.99, and the soft update coefficient of the target network is 0.005. The remaining hyperparameters are kept consistent with the default values in the SAC algorithm.
[0124] To realize the effective training of the autonomous collision avoidance strategy, a composite reward function is constructed by comprehensively considering the target distance reward, the heading reward, the collision penalty, the collision risk penalty, the speed obstacle penalty and the international collision avoidance rule penalty. The calculation form of the total reward function is as follows:
[0125] (14)
[0126] Where They respectively represent the weighted coefficients of distance reward, heading reward, collision risk penalty, collision penalty, penalty for violating collision avoidance rules, and speed obstacle penalty. The specific numerical values of each part of the reward and the specific meanings of each sub-item are as follows:
[0127] 1. Distance reward term (r d )
[0128] It encourages the unmanned boat to continuously move towards the target point, and the specific calculation is:
[0129] (15)
[0130] Where and respectively represent the Euclidean distances from the unmanned boat to the target point at the previous moment and the current moment. represents the speed of the unmanned boat at the previous moment. j represents the simulation step size.
[0131] 2. Heading reward term (r h )
[0132] It encourages the unmanned boat to adjust its heading towards the target point, and the calculation formula is:
[0133] (16)
[0134] 3. Collision risk penalty term (r cri )
[0135] This term is used to penalize the unmanned boat for taking collision avoidance behaviors that increase navigation risks, and its calculation is shown as follows:
[0136] (17)
[0137] Where represents the collision risk formed by the own ship and the most dangerous target ship.
[0138] 4. Collision penalty term (r coll )
[0139] It is triggered when the agent collides or successfully reaches the target point, and its calculation is shown as follows:
[0140] (18)
[0141] Where 、 、 and respectively represent the distance between the unmanned boat and the target point, the distance between the unmanned boat and the i-th opposing ship, the distance threshold for judging whether the target point is reached, and the radius of the ship safety area of the i-th opposing ship.
[0142] 5. Penalty item for violating collision avoidance rules (r C ).
[0143] To improve the model's compliance ability with the International Regulations for Preventing Collisions at Sea, this penalty item is added in the single-ship intersection scenario. This item is used to guide the unmanned boat to take collision avoidance measures that comply with the collision avoidance rules. Its calculation is shown as follows:
[0144] (19)
[0145] 6. VO penalty item (r VO ).
[0146] In ship collision avoidance, VO is a collision avoidance modeling method based on the velocity space. It defines that if the velocity of the OS falls into the velocity obstacle area formed by the TS within a given time window, there may be a collision between the two ships. Therefore, the present invention introduces a VO penalty item into the reward function to punish those action selections that may potentially lead to entering the collision trajectory during the reinforcement learning training process, so as to drive the policy to automatically avoid the unsafe velocity area. Its calculation is shown as follows:
[0147] (20)
[0148] Through the provided technical solutions, the present invention organically integrates recursive Bayesian and deep reinforcement learning frameworks, and can improve the autonomous collision avoidance ability of unmanned boats under the condition of the uncertainty of the target ship's behavior. In the specific implementation process, the navigation intention of the target ship is dynamically estimated recursively, and its posterior probability distribution is explicitly introduced into the state space, effectively enhancing the modeling ability of the policy network and the value network for the behavior pattern of the target ship. At the same time, a multi-factor composite reward function system is designed to guide the intelligent agent to learn action strategies that are more in line with the maritime collision avoidance rules and have higher safety during the training process. To ensure the adaptability and generalization ability of the training strategy, the present invention constructs a simulation environment with high dynamics and uncertainty, and uses randomly generated parameters such as the number of target ships, headings, speeds, and positions, covering various typical behavior patterns such as "maintaining speed and course", "collision avoidance", and "intentional collision". Through the above technical means, the present invention realizes high-precision dynamic modeling of the target behavior intention and policy optimization, effectively improves the robustness and execution efficiency of the collision avoidance strategy of unmanned boats under multi-target intersection, behavior uncertainty, and rule constraints, has good engineering practicability, and is applicable to various autonomous navigation application scenarios such as intelligent shipping, maritime monitoring, and rescue search.
[0149] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0150] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0151] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0152] As described above, these are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0153] The present invention is not limited to the above best implementation mode. Anyone can obtain other various forms of a collision avoidance decision-making method for an unmanned boat based on navigation intention perception under the inspiration of the present invention. All equal changes and modifications made according to the scope of the present invention application shall fall within the coverage scope of the present invention.
Claims
1. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception, characterized in that: Motion state data of the own ship and target ships are collected in real time through multi-source sensors, and a collision risk assessment model is constructed to screen for threatening targets; Based on the recursive Bayesian inference of the historical trajectory of the target ship, the posterior probability distribution of the navigation intention of the target ship is dynamically updated, and the navigation intention is discretized into multiple behavior patterns; The state of the own ship, the motion characteristics of the target ship, and the intention distribution are constructed into a state vector and input into a policy network based on reinforcement learning to output a collision avoidance control action; The policy network is optimized by a composite reward function, and the reward function combines distance reward, rule penalty, and collision risk penalty.
2. The collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, wherein: The multiple behavior patterns of the navigation intention include combinations of left turn, straight ahead, right turn of the course and acceleration, constant speed, deceleration of the speed, for a total of 9 patterns.
3. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The recursive Bayesian inference includes: Predicting the state based on the dynamic model of the surge acceleration and yaw acceleration of the target ship; Modeling the observation noise by Gaussian distribution and calculating the likelihood probability of the predicted state and the actual observation; Recursively updating the posterior probability distribution of the intention.
4. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: When constructing the state vector, the data of the target ship with the highest collision risk index is preferentially selected, and when there is insufficient data, it is filled with a zero vector.
5. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The composite reward function includes an international collision avoidance rule penalty term and a speed obstacle penalty term, and a linear penalty is triggered when the action violates the rules or enters the speed obstacle area.
6. The collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The policy network based on reinforcement learning is of SAC architecture, and the output action is sampled from a Gaussian distribution through the reparameterization technique and mapped to the physical control range through the tanh function.
7. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The method is trained in a dynamic simulation environment, and the behavior patterns of the target ships are randomly selected according to preset probabilities, including three strategies: maintaining speed and course, collision avoidance, and intentional collision.
8. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The multi-source sensors include a global navigation satellite system, radar, and an automatic identification system, and the collected data is input into the model after preprocessing such as anomaly rejection, time synchronization, and coordinate unification.
9. A collision avoidance decision-making method for an unmanned boat based on navigation intention perception according to claim 1, characterized in that: The collision risk assessment model calculates the risk index through five-factor linear weighting, including relative distance, relative speed, relative course, minimum collision distance, and time to the closest point of approach.
10. An unmanned boat collision avoidance decision-making system based on navigation intention perception, characterized in that, It includes: A multi-source perception module for collecting motion state data of the own ship and target ships in real time; A threatening target screening module for constructing a collision risk assessment model to screen for threatening targets; An intention inference module for dynamically updating the posterior probability distribution of the navigation intention of the target ship based on the recursive Bayesian inference of the historical trajectory of the target ship; A policy network module for generating a collision avoidance control action according to the state of the own ship, the motion characteristics of the target ship, and the intention distribution; A control execution module for converting the action instruction into a ship's course and speed control signal.
Citation Information
Patent Citations
Unmanned ship collision avoidance model construction method and device and unmanned ship collision avoidance method and device
CN117523925A
Ship collision avoidance optimization method under condition of uncertain obstacle ship motion information
CN119207166A
Multi-sensor adaptive anti-collision method and device based on deep learning and medium
CN119989218A
Learning device, learning method, and learning program
WO2022230038A1
Cited By
Intelligent obstacle avoidance method and system applied to ship
CN121411453A
Unmanned ship smooth collision avoidance method considering marine environment disturbance
CN121596884A
Ship autonomous collision avoidance decision-making method and system based on cognitive entropy near-end strategy optimization
CN121764176A