Unmanned ship collision avoidance method based on memory mechanism deep reinforcement learning
By adopting a collision avoidance method based on memory mechanism and deep reinforcement learning in unmanned ships, the problem of insufficient autonomous collision avoidance performance of unmanned ships against dynamic obstacles under perceptually restricted conditions is solved, and more efficient and safe collision avoidance decisions and path planning are achieved.
Patent Information
- Application Number
- CN202510484094.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art has insufficient autonomous collision avoidance performance of unmanned ships to dynamic obstacles under perceptually restricted conditions, especially in complex marine environments. The incomplete information caused by limited sensor detection range and environmental interference have led to a significant reduction in the reliability of collision avoidance decisions.
The unmanned ship collision avoidance method based on memory mechanism and deep reinforcement learning is adopted. By constructing a fixed-length memory space, the unmanned ship's historical navigation state sequence is dynamically stored, and the collision avoidance action instructions are generated in combination with the gated cycle unit (GRU) and the multi-layer perceptron (MLP) to achieve autonomous collision avoidance under perception limitations.
It significantly improves the collision avoidance performance of unmanned ships under perceptually restricted conditions, improves the collision avoidance success rate by 22%, shortens the decision response time to 85ms, optimizes the path planning efficiency by 15%, and ensures safe and compliant collision avoidance decisions in complex marine environments.
Smart Images

Figure CN120010498A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of unmanned ship collision avoidance and autonomous decision-making, and specifically relates to an unmanned ship collision avoidance method based on deep reinforcement learning of a memory mechanism. Background Art
[0002] With the rapid development of unmanned marine system technology, unmanned ships are increasingly used in the fields of marine resource exploration, maritime search and rescue, etc. In actual navigation missions, unmanned ships need to deal with the dual threats of static obstacles such as buoys and islands, and dynamic obstacles such as ships with uncertain motion states, which puts strict requirements on autonomous collision avoidance algorithms. Especially in complex marine environments, due to the limitations of sensor detection range (such as millimeter-wave radar field of view) and environmental interference (such as waves and tidal currents), unmanned ships are often unable to obtain the complete status information of obstacles in real time (including key parameters such as heading, speed, and size), resulting in a significant reduction in the decision-making reliability of existing collision avoidance algorithms in perception-limited scenarios.
[0003] At present, the solutions to the problem of limited perception can be divided into two categories: the first is to improve the environmental perception capability by adding perception devices such as lidar and multispectral sensors, but this method significantly increases the hardware cost and system complexity, and there are technical bottlenecks in the fusion of multi-source sensor data; the second is to improve the design of decision-making algorithms, but traditional algorithms have inherent defects: although the dynamic improved A* algorithm combines the COLREGs rule, it is highly dependent on the accuracy of the environmental model and has low efficiency in path replanning in dynamic obstacle-intensive scenes; the artificial potential field method (APF) performs well in static obstacle avoidance, but the superposition of multiple obstacle potential fields in a dynamic environment can easily lead to local optimal traps; the collision avoidance strategy based on geometric prediction of the relative velocity obstacle method (RVO) is difficult to cope with complex environmental disturbances, and has a high risk of failure when obstacle information is missing. Although deep reinforcement learning (DRL) has shown potential in decision-making tasks, the traditional DRL model has problems such as strong dependence on environmental information and low efficiency of sparse reward learning, and the collision avoidance performance drops sharply when perception is limited.
[0004] In the related prior art, although the authorized patent CN202411556502.2 (Unmanned boat obstacle avoidance path planning method and system based on global optimization) improves the accuracy of path planning, it does not solve the problems of dynamic obstacle interaction and limited perception; the patent CN202010717418.X (An unmanned boat path planning method based on deep reinforcement learning and taking into account marine environmental elements) focuses on marine environment modeling, but does not involve decision optimization in information-missing scenarios; the patent CN201911043840.5 (A group of unmanned boats intelligent collision avoidance method based on deep reinforcement learning) is designed for multi-agent collaborative scenarios, but its assumption based on complete environmental information is limited in applicability when the perception of a single ship is limited. It can be seen that the prior art has not effectively solved the problem of autonomous collision avoidance of unmanned boats against dynamic obstacles under sensor-constrained conditions. Summary of the invention
[0005] The present invention discloses an unmanned ship collision avoidance decision-making scheme based on memory mechanism and reinforcement learning, which is used to solve the problem of failure of unmanned ship collision avoidance decision-making due to incomplete environmental information under perception-constrained conditions. Its core design is to combine historical navigation data with the collaborative optimization of the reinforcement learning framework to achieve safe and compliant collision avoidance of dynamic obstacles.
[0006] Its core design is: An unmanned ship collision avoidance method based on memory mechanism and reinforcement learning comprises the following steps: (1) Dynamically store the historical navigation state sequence of the unmanned ship and construct a fixed-length memory space. The state sequence includes the relative position of the target point, the relative position of the obstacle, and the historical action data; (2) inputting the state sequence in the memory space into the reinforcement learning decision network, extracting the temporal features through the gated recurrent unit (GRU), and generating collision avoidance action instructions in combination with the multi-layer perceptron (MLP); (3) Based on the collision avoidance action instructions, the unmanned boat is controlled to navigate, and the network parameters are optimized by calculating the instantaneous reward value according to the composite reward function to achieve autonomous collision avoidance under limited perception.
[0007] Furthermore, the reinforcement learning decision network adopts a soft actor-critic algorithm to optimize network parameters by minimizing the value network loss function and maximizing the policy entropy.
[0008] Furthermore, the memory space is implemented by a first-in-first-out queue, and the oldest state data is removed and the current state data is added each time it is updated.
[0009] The memory space uses a first-in-first-out (FIFO) queue to achieve rolling updates. Each time the current state is added, the oldest historical data is automatically eliminated to ensure that the stored n-step state sequence is strictly arranged in chronological order. This design solves the situation assessment bias problem caused by the interference of invalid historical data in traditional methods, and the unmanned ship can still plan a safe path when the radar field of view is limited.
[0010] Furthermore, the hidden state dimension of the gated recurrent unit GRU is 128, the MLP layer dimension is 256, and the decision network outputs Gaussian distribution parameters of propulsion force and torque. This lightweight network structure reduces computing resource usage by 30% while ensuring decision accuracy, and is suitable for embedded device deployment.
[0011] Furthermore, the compound reward function includes: Target proximity reward: dynamically adjusted based on the Euclidean distance between the unmanned ship and the target point; Collision avoidance safety reward: calculated in segments based on the ratio of the relative distance to the obstacle and the safety threshold, with a penalty imposed when the distance is below the minimum safety threshold; COLREGs compliance reward: dynamically adjusts the penalty weight by counting the number of violations in historical action sequences.
[0012] Among them, the target proximity reward is dynamically adjusted based on the Euclidean distance between the unmanned ship and the target point (step 6.1) to guide efficient navigation; the collision avoidance safety reward is calculated in segments according to the ratio of the relative distance to the obstacle (d_OT) to the safety threshold (d_safe) (step 6.2). When the distance is lower than the minimum safety threshold (d_min), a penalty of -6000 is imposed to force the maintenance of a safe distance; the COLREGs compliance reward dynamically adjusts the penalty weight by counting the number of violations in the historical action sequence (step 6.5). This method complies with the requirements of the rules in scenarios such as overtaking and encountering.
[0013] Furthermore, the network parameter update adopts continuous n-step historical data sampling and is updated through a KL divergence constraint strategy to avoid strategy oscillation caused by memory space data distribution offset.
[0014] This method constructs a fixed-length memory space to dynamically store the historical navigation state sequence of the unmanned ship in the last n steps, including the relative position of the target point, the relative position of the obstacle, and the historical action data. When the sensor detection range is limited (such as the millimeter-wave radar only covers -60° to 60°), the gated recurrent unit (GRU) is used to extract time series features, infer the movement trend of obstacles (such as speed and heading changes), and combine the multi-layer perceptron (MLP) to generate collision avoidance action instructions for propulsion and torque, so as to achieve autonomous collision avoidance in dynamic obstacle scenarios. The network parameters are optimized by the soft actor-critic (SAC) algorithm. The effectiveness of this method has been verified in simulation and real ship experiments, with the collision avoidance success rate increased by 22% and the decision response time shortened to 85ms.
[0015] And, corresponding to the method, an unmanned ship collision avoidance decision system, comprising: A memory module is used to store the historical navigation state sequence of the most recent n steps; the state sequence includes the relative position of the target point, the relative position of the obstacle and the historical action data; A decision module, integrating a gated recurrent unit and a reinforcement learning network of a multi-layer perceptron, outputs collision avoidance action instructions based on the historical state sequence, and optimizes network parameters through a composite reward function; A control module controls the propulsion and steering of the unmanned boat based on the action instructions; The training module optimizes network parameters through a compound reward function to achieve autonomous collision avoidance under limited perception.
[0016] The system successfully completed the collision avoidance task in a dynamic obstacle scenario.
[0017] Furthermore, the memory module implements rolling storage of the state sequence through a first-in-first-out queue, and removes the oldest state data and adds the current state data each time it is updated.
[0018] Furthermore, the reinforcement learning network adopts a soft actor-critic algorithm to optimize network parameters by minimizing the value network loss function and maximizing the policy entropy; the GRU hidden layer dimension of the decision module is 128, the MLP layer dimension is 256, and the Gaussian distribution parameters of the output propulsion force and torque are output; the composite reward function of the training module includes a target proximity reward, a collision avoidance safety reward, and a COLREGs compliance reward, wherein the collision avoidance safety reward is calculated in segments according to the ratio of the relative distance of the obstacle to the safety threshold; the training module adopts continuous n-step historical data sampling and is updated through a KL divergence constraint strategy; it also includes a millimeter-wave radar and a GNSS module, the detection range of the millimeter-wave radar is -60° to 60°, which is used to obtain the relative position of obstacles; the GNSS module is used to locate the coordinates of the unmanned ship in real time.
[0019] And, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0020] A non-transitory computer-readable storage medium stores a computer program, which implements the steps of the method described above when executed by a processor.
[0021] Compared with the prior art, the present invention and its preferred solution significantly improve the collision avoidance performance of the unmanned ship under perception-restricted conditions through the coordinated optimization of the following core invention points: 1. Situation completion with memory mechanism and GRU timing coordination By building a fixed-length memory space (FIFO queue), storing the most recent n-step historical states (target point position, obstacle position, historical actions), and combining with the GRU module to extract time series features. When the millimeter-wave radar field of view is limited (-60° to 60°), the obstacle movement trend is inferred through historical data (speed change rate ≥ 0.3m / s², heading angle deviation ≤ 5°), and the situation completion accuracy is improved by 35%. The collision avoidance success rate is increased from 73% of traditional DRL to 95%, and the path planning efficiency (navigation time) is optimized by 15%.
[0022] 2. Multi-objective collaborative optimization of compound reward functions The target proximity reward, collision avoidance safety reward and COLREGs compliance reward are integrated, and the weights are adjusted dynamically. The collision avoidance safety reward is calculated in segments to force the vehicle to maintain a safe distance, with a close distance penalty value of -6000, which reduces the risk of collision by 90%; the COLREGs compliance reward uses a sliding window (m=5) to count the number of historical violations, dynamically adjust the penalty weight, and reduce violations by 90%. The target proximity reward guides the path efficiency, shortens the flight distance by 15%, and improves the learning efficiency by 50% in reward-sparse scenarios.
[0023] 3. Lightweight network and stable training mechanism The GRU (128-dimensional) + MLP (256-dimensional) network structure is adopted, and the KL divergence constraint strategy is updated. The number of network parameters is reduced by 30%. The continuous n-step (n=8) historical data sampling combined with the KL divergence constraint improves the training stability by 40% and the convergence speed is accelerated by 1.8 times.
[0024] 4. COLREGs rule-embedded reinforcement learning decision The COLREGs rule is converted into a reward function constraint term, combined with the policy entropy maximization feature of the SAC algorithm. The compliance rate of key scenarios of collision avoidance behavior in complex scenarios such as encounter and overtaking is 100%; the path tracking error in dynamic obstacle interaction scenarios is reduced by 42%.
[0025] Through the deep collaboration of memory mechanism and reinforcement learning, the present invention overcomes the problems of real-time, safety and rule compatibility of unmanned ship collision avoidance decision-making under limited perception, and provides an efficient and reliable technical solution for operations in complex marine environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: Figure 1 This is a flow chart of collision avoidance of an unmanned ship according to an embodiment of the present invention; Figure 2 A collision avoidance path planning diagram for a simulation scenario according to an embodiment of the present invention; Figure 3 This is a diagram of the deep reinforcement learning network structure of the memory mechanism of an embodiment of the present invention; Figure 4 A diagram showing the construction of an unmanned ship model according to an embodiment of the present invention; Figure 5 A collision avoidance scene division diagram according to an embodiment of the present invention; Figure 6 This is a collision avoidance rule diagram of an embodiment of the present invention; Figure 7 This is a graph of the memory mechanism deep reinforcement learning network update of an embodiment of the present invention; Figure 8 A comparison chart of the success rates of various algorithms in the embodiments of the present invention; Fig. 9 This is a comparison chart of rewards for various algorithms in an embodiment of the present invention; Fig.10 It is an obstacle-free navigation trajectory diagram of an embodiment of the present invention; Fig.11 It is a static obstacle avoidance trajectory diagram of an embodiment of the present invention; Fig.12 This is a collision avoidance trajectory diagram of the port side of an unmanned ship according to an embodiment of the present invention; Fig.13 This is a collision avoidance trajectory diagram of the starboard side of an unmanned ship according to an embodiment of the present invention; Fig.14 This is a trajectory diagram of an unmanned ship encounter according to an embodiment of the present invention; Fig.15 This is a trajectory diagram of an unmanned ship overtaking according to an embodiment of the present invention; Fig.16 This is a diagram of the equipment of an unmanned ship according to an embodiment of the present invention; Fig.17 This is a diagram of a real ship experiment of the unmanned ship collision avoidance method with a memory mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to make the features and advantages of the present invention more clearly understood, the following embodiments are specifically described in detail as follows: It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0029] In order to solve the problem of autonomous collision avoidance of unmanned ships under limited perception, the present invention proposes an autonomous collision avoidance decision method based on a memory mechanism, which can use historical navigation data to construct the surrounding situation under limited perception, thereby improving the accuracy and safety of collision avoidance decisions. Its deep reinforcement learning algorithm (MDRL) based on the memory mechanism is used to optimize the collision avoidance decision of unmanned ships. By constructing a memory space, the algorithm stores and processes historical navigation data, thereby effectively evaluating the surrounding situation in the absence of real-time environmental information. The core of the solution is to use historical navigation data to make up for the lack of immediate perception, and combine it with an improved decision-making network to ensure that unmanned ships make safe and compliant collision avoidance decisions in dynamic and complex marine environments. In addition, by combining historical information with short-term decisions, the present invention improves the adaptability of unmanned ships to changing sea conditions and achieves more efficient autonomous collision avoidance, thereby providing safety guarantees for unmanned ships when performing various maritime tasks and promoting the further development of unmanned ship technology.
[0030] The embodiment of the present invention proposes a collision avoidance decision method for unmanned ships based on a memory mechanism. Based on the soft actor-critic (SAC) algorithm in deep reinforcement learning, the memory mechanism of the gated recurrent unit (GRU) is integrated to construct a decision model that can use historical navigation data to solve the problem of autonomous collision avoidance of unmanned ships under limited perception. By optimizing the reward function and improving the design of the experience pool, the collision avoidance decision-making ability and learning efficiency of the unmanned ship are improved.
[0031] The process steps of the unmanned ship collision avoidance decision training method are as follows: Figure 1 As shown: Step 1, initialize the training environment; Step 2, obtaining environmental information; Step 3, update the memory space; Step 4: According to step 3, the decision network makes a decision; Step 5: According to step 4, execute corresponding decision-making behavior; Step 6, according to step 5, calculate the reward value; Step 7, determine whether the navigation mission is completed or a collision occurs, update the decision network or end the navigation.
[0032] For the implementation of step 1, this embodiment establishes a simulated ocean environment and an intelligent agent including dynamic and static obstacles on a simulation platform. Figure 2 shown.
[0033] As a preferred implementation, initializing the training environment is specifically divided into the following steps: Step 1.1, set the environmental parameters, including water area, waves, ocean currents and other environmental interference factors, as well as the number, location, size, speed and heading of obstacles; Step 1.2, initialize the position information, speed, heading and other state parameters of the unmanned ship; Step 1.3, setting the target point and navigation path of the unmanned ship, and planning the navigation mission target for the unmanned ship; Step 1.4, build the network structure, such as Figure 3 As shown in the figure, the SAC algorithm decision network and value network are constructed. Both the policy network and the value network contain GRU and multi-layer perceptron (MLP) modules. The dimension of the GRU layer is 128, and the dimension of the MLP layer is 256. The policy network contains the mean layer and the standard deviation layer, which are used to output propulsion force and torque. The value network directly outputs the behavior value for network update; For step 2, this embodiment can obtain limited environmental information by simulating the real scene of the unmanned ship, collecting the current navigation status ,like Figure 4 As shown, it contains the position coordinates of the target point relative to the unmanned ship , target azimuth , the obstacle position coordinates relative to the unmanned ship The action performed at the previous moment a .
[0034] As a demonstration of one of the core design points of the present invention, for step 3, this embodiment simulates the human short-term memory and forgetting mechanism by updating the memory space, retaining key historical information to assist decision-making. t Store in memory space c t , the memory space stores the navigation status of the most recent n steps, forming a state sequence .
[0035] For step 4, the state sequence C in the memory space obtained in step 3 is t , is input into the decision network. As a preferred implementation, the specific decision process of the network is divided into the following steps: Step 4.1, the GRU module in the policy network processes the state sequence data to extract the dynamic characteristics and time correlation of the environment; Step 4.2, based on the environmental dynamic features and time correlation extracted in step 4.1, such as Figure 5 As shown in the figure, it is determined whether the unmanned ship will perform collision avoidance actions at present. If collision avoidance actions are performed, then according to Figure 6 The International Regulations for Preventing Collisions at Sea (COLREGs) are used for decision making; Step 4.3: After the extracted features are processed by a multi-layer perceptron (MLP), the action to be performed is output. ,in Indicates thrust, Indicates torque; For step 5, in this embodiment, the unmanned ship takes actions according to the output of the strategy network. , perform corresponding propulsion and steering operations in the simulation environment. The environment simulates the state changes after the unmanned ship performs the action, and updates the position information, speed, heading, etc. of the unmanned ship.
[0036] For step 6, calculate the immediate reward based on the action performed by the unmanned ship and the new state , the reward function consists of multiple indicators, For each reward weight, as a preferred implementation, the specific calculation steps are as follows: Step 6.1, calculate the reward for approaching the target , used to encourage the unmanned ship to sail towards the target point. The calculation formula is as follows:
[0037] Among them, K1 is the target proximity reward weight; (x o ,y o ) represents the current coordinates of the unmanned ship, which can be obtained in real time through GNSS.
[0038] Step 6.2, calculate the collision avoidance reward , according to the distance between the unmanned ship and the obstacle , give rewards or penalties to encourage safe navigation, the calculation formula is as follows:
[0039] Among them, K2 is the collision avoidance safety reward weight, d OT Represents the real-time distance between the unmanned ship and the obstacle, d safe is the safety threshold, d min It is the size of an unmanned ship.
[0040] Step 6.3, calculate the angle reward , control the heading change of the unmanned ship, avoid excessive turning, and ensure smooth navigation. The calculation formula is as follows:
[0041]
[0042]
[0043] in, is the penalty item for heading deviation from the target; is the penalty term for heading angle change; is the change in the heading angle of the unmanned ship, which can be measured by a gyroscope; K3 is the angle reward weight.
[0044] Step 6.4, calculate speed bonus , encourage the unmanned ship to sail at an appropriate speed and keep the speed stable. The calculation formula is as follows:
[0045]
[0046]
[0047] in, It is the speed error penalty term, which can be adjusted dynamically according to the task type (such as overtaking / other); is the speed change penalty term; is the speed change of the unmanned ship, which can be measured by an accelerometer; K4 is the speed reward weight; Step 6.5, calculate COLREGs reward r COLREGs , to ensure that the collision avoidance behavior of the unmanned ship complies with COLREGs. The calculation formula is as follows:
[0048] Among them, K5 is the violation penalty weight; Step 6.6, calculate the total reward :
[0049] For step 7, the policy network is optimized by continuously interacting with the environment. Determine whether the navigation task is completed or a collision occurs. If the task is completed or a collision occurs, calculate the total reward value and result. Otherwise, update the network model parameters, such as Figure 7 The specific update steps are as follows: Step 7.1, sample data from the experience pool, randomly sample a batch of n consecutive steps of navigation data from the experience pool to form training samples , where C is the current state sequence, a is the action performed in state C, r is the total reward calculated in step 6, C' is the state sequence at the next moment after the action is performed, and d is the termination signal, indicating whether the task is completed or a collision occurs.
[0050] Step 7.2, calculate the state value target , the calculation formula is
[0051] Among them, γ is the discount factor, α is the entropy weight used to adjust the randomness of the strategy, is the target value network, For the strategy π in state C' θ Sampling action.
[0052] Step 7.3, Update the value network , minimize the loss function, value network Used to evaluate the quality of the policy network, the update formula is:
[0053] in .
[0054] Step 7.4, update the policy network to maximize the expected reward and entropy. The update formula is:
[0055] in is obtained from the policy π by reparameterization technique. θ The action sampled in .
[0056] Step 7.5, soft update target value network , the update formula is:
[0057] Where ρ is the soft update coefficient.
[0058] The collision avoidance algorithm is based on the memory-based deep reinforcement learning (MDRL) algorithm, forming an improved strategy that combines the memory mechanism and GRU. When the unmanned ship performs autonomous collision avoidance, the memory space stores historical navigation data, the GRU module processes time series information, and comprehensively utilizes current and historical environmental information to assess the risk of collision. When an obstacle is found within the detectable range, the unmanned ship uses the memory mechanism and limited environmental perception, and uses historical navigation data to determine the movement trend of the obstacle and possible collision threats.
[0059] The algorithm performs collision avoidance behavior according to the size of the collision risk. The collision avoidance behavior rules comply with COLREGs and complete the safe navigation task with the optimal path. By introducing a memory mechanism and an improved reward function, the algorithm effectively solves the problem of sparse rewards in reinforcement learning, improves the network update speed and learning efficiency, and improves the shortcomings of high randomness and low learning rate. The algorithm's collision avoidance ability is enhanced, and it has good generalization ability, which improves the safety navigation efficiency of unmanned ships in complex marine environments.
[0060] The specific simulation implementation and actual ship verification process corresponding to the solution of this embodiment are further elaborated in detail below.
[0061] 1. Simulation experiment part: 1.1 Initialize the training environment A simulated ocean environment with dynamic and static obstacles is established on the simulation platform. The environmental parameters are set, including the water area (e.g., an area of 300×300m), wave flow (size, direction), and the number of obstacles (set to fixed values, such as one static obstacle and three dynamic obstacles), location, size, speed, and heading. These parameters are dynamically simulated through a random number generator to simulate the complexity of the actual ocean environment.
[0062] 1.2 Initialize the unmanned ship status Set the initial position of the unmanned ship ,speed , heading angle As well as thrust and torque. The initial state of the unmanned ship starts from a preset starting point and the target point position is ,The dynamic behavior model of the unmanned ship is described by combining the dynamic equation and the kinematic equation.
[0063] 1.3 Obtaining environmental information In the simulation environment, the only information about the surrounding environment known to the unmanned ship is the location information of obstacles.
[0064] 1.4 Update memory space The environmental information and historical states obtained by the unmanned ship are stored in the memory space for input into the subsequent decision network. Specifically, the memory space stores the state sequence of the most recent n steps. This short-term memory mechanism simulates human memory behavior and helps unmanned ships make reasonable decisions under limited perception conditions.
[0065] 1.5 Decision-making process of decision network The state sequence in the memory space Input to the policy network. The GRU module first processes the input time series data, extracts the dynamic features and time correlation of the environment, and then converts the hidden state vector The input is sent to the MLP, and the propulsion force and torque are output. The output of the propulsion force and torque is calculated by the mean layer and the standard deviation layer to generate a Gaussian distribution, and the policy network selects actions based on this Gaussian distribution.
[0066] 1.6 Execution of decision-making behavior The unmanned ship performs propulsion and steering operations in the simulation environment according to the propulsion and torque output by the policy network. The magnitude of the thrust controls the forward speed of the unmanned ship, and the torque controls the change of its heading angle. After executing the behavior, the simulation platform updates the state of the unmanned ship and calculates its distance from obstacles and target points to facilitate the subsequent calculation of rewards.
[0067] 1.7 Calculating Reward Value After each action is performed, the reward value of the current state of the unmanned ship is calculated , used to optimize the decision network.
[0068] 1.8 Update the decision network Through continuous interaction with the environment, the policy network is continuously optimized to improve the collision avoidance capability of the unmanned ship.
[0069] 1.9 Circuit Training The above process is repeated until the unmanned ship can effectively perform collision avoidance behavior and reach the target point. Through multiple rounds of training, the model is continuously optimized, and finally a stable decision network is obtained for the actual collision avoidance decision of the unmanned ship.
[0070] 2. Actual ship verification part: 2.1 Experimental Equipment and Setup 2.1.1 Unmanned boat platform The unmanned boat used in the experiment is 2 meters long, and its main body is made of carbon fiber material, which has excellent navigation performance. The unmanned boat is equipped with a variety of sensors (such as cameras, millimeter-wave radars, GNSS, etc.) to collect environmental information. Since the ship is equipped with a millimeter-wave radar and the monitoring range is between -60° and 60°, and it is impossible to obtain accurate obstacle speed and size information, there is a situation of limited perception. The actual ship control system includes a shipboard computer and a shore-based computer. The unmanned boat runs the policy network in real time and executes decision-making behaviors through the shipboard computer. The shore-based computer is mainly used for data monitoring, collection and experimental recording.
[0071] 2.1.2 Experimental Environment The experimental site was selected in a water area of about 500 meters by 500 meters. Dynamic obstacles (small artificial boats) were set up in the water area. When the experiment was carried out, the wind speed was about 16 kilometers per hour and there were slight waves on the water surface.
[0072] 2.2 Experimental procedures 2.2.1 Unmanned boat initialization Set the initial position of the unmanned boat in the experimental waters and target point The navigation task is set to start from the starting point and reach the target point. When the unmanned boat starts, it first obtains the starting position information and the target position, and the initial speed is set to 0.5m / s.
[0073] 2.2.2 Obtaining environmental information During navigation, the unmanned ship senses the location information of obstacles through the millimeter-wave radar installed on the ship, obtains its own location data through the GNSS module, and forms the current state vector for subsequent decision-making network input. The sensor information is updated at a frequency of 10 Hz to ensure that the unmanned ship can obtain environmental data in a dynamic environment in a timely manner.
[0074] 2.2.3 Real-time decision-making and collision avoidance control The unmanned ship inputs the collected state information into the decision network. The strategy network outputs thrust and torque based on the time series features extracted by the GRU module and the feature calculations of the MLP module.
[0075] 2.3 Experimental scenario testing 2.3.1 Encounter During the experiment, the unmanned boat encountered a manned boat head-on. According to COLREGs, the unmanned boat should turn to the right to avoid it. The experimental results show that the unmanned boat turned right, avoided the oncoming boat and successfully resumed its original route.
[0076] 2.3.2 Starboard Crossing The target ship came from the right side of the unmanned ship, and the decision network judged that it was a starboard crossing situation, and the unmanned ship should turn right to avoid it. In the experiment, the unmanned ship showed a stable collision avoidance path and successfully avoided the target ship.
[0077] 2.3.3 Port crossing The target ship came from the left side of the unmanned boat, and the unmanned boat kept its course and speed unchanged according to the rules. In the experiment, because the target ship did not take evasive action, it posed a threat to our unmanned boat. The unmanned boat took evasive action to the right and kept a safe distance from the target ship.
[0078] 2.3.4 Overtaking The unmanned ship approached the target ship from behind and attempted to overtake it. The decision network determined the speed and course of the target ship based on its location information and selected the best overtaking route, successfully completing the overtaking while maintaining a safe distance from the target ship.
[0079] like Figure 8-Figure 17As shown, this embodiment verifies the effectiveness and feasibility of the unmanned ship collision avoidance decision method based on memory mechanism proposed by the present invention through experiments. Under the condition of limited perception, the unmanned ship can reasonably judge the situation and make collision avoidance decisions in accordance with COLREGs, which has better safety, stability and generalization ability than traditional reinforcement learning algorithms.
[0080] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0081] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by the processor to execute the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0082] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0083] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
[0084] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive other various forms of ship tracking control methods with predefined time specification performance under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention should fall within the scope of the present invention.
Claims
1. A method for unmanned ship collision avoidance based on deep reinforcement learning of memory mechanism, characterized by: Dynamically store the historical navigation state sequence of the unmanned ship and construct a fixed-length memory space, wherein the state sequence includes the relative position of the target point, the relative position of the obstacle and the historical action data; The state sequence in the memory space is input into a reinforcement learning decision network, the timing features are extracted through a gated recurrent unit, and a collision avoidance action instruction is generated in combination with a multi-layer perceptron; The navigation of the unmanned boat is controlled based on the collision avoidance action instruction, and the network parameters are optimized by calculating the instant reward value according to the composite reward function to achieve autonomous collision avoidance under limited perception.
2. The unmanned ship collision avoidance method based on deep reinforcement learning of memory mechanism according to claim 1 is characterized in that: The reinforcement learning decision network adopts a soft actor-critic algorithm to optimize network parameters by minimizing the value network loss function and maximizing the policy entropy.
3. The unmanned ship collision avoidance method based on deep reinforcement learning of memory mechanism according to claim 1 is characterized in that: The memory space is implemented by a first-in-first-out queue, and the oldest state data is removed and the current state data is added each time it is updated.
4. The unmanned ship collision avoidance method based on deep reinforcement learning of memory mechanism according to claim 1 is characterized in that: The hidden state dimension of the gated recurrent unit is 128, the MLP layer dimension is 256, and the decision network outputs Gaussian distribution parameters of propulsion force and torque.
5. The unmanned ship collision avoidance method based on deep reinforcement learning of memory mechanism according to claim 1 is characterized in that: The compound reward function includes: Target proximity reward: dynamically adjusted based on the Euclidean distance between the unmanned ship and the target point; Collision avoidance safety reward: calculated in segments based on the ratio of the relative distance to the obstacle and the safety threshold, with a penalty imposed when the distance is below the minimum safety threshold; COLREGs compliance reward: dynamically adjusts the penalty weight by counting the number of violations in historical action sequences.
6. The unmanned ship collision avoidance method based on deep reinforcement learning of memory mechanism according to claim 1 is characterized in that: The network parameter update adopts continuous n-step historical data sampling and is updated through the KL divergence constraint strategy to avoid strategy oscillation caused by memory space data distribution offset.
7. An unmanned ship collision avoidance decision system, characterized in that: include: A memory module is used to store the historical navigation state sequence of the most recent n steps; the state sequence includes the relative position of the target point, the relative position of the obstacle and the historical action data; A decision module, integrating a gated recurrent unit and a reinforcement learning network of a multi-layer perceptron, outputs collision avoidance action instructions based on the historical state sequence, and optimizes network parameters through a composite reward function; A control module controls the propulsion and steering of the unmanned boat based on the action instructions; The training module optimizes network parameters through a compound reward function to achieve autonomous collision avoidance under limited perception.
8. The unmanned ship collision avoidance decision system according to claim 7, characterized in that: The memory module implements rolling storage of the state sequence through a first-in-first-out queue, removing the oldest state data and adding the current state data each time an update occurs.
9. The unmanned ship collision avoidance decision system according to claim 7, characterized in that: The reinforcement learning network adopts the soft actor-critic algorithm to optimize the network parameters by minimizing the value network loss function and maximizing the policy entropy; the GRU hidden layer dimension of the decision module is 128, the MLP layer dimension is 256, and the Gaussian distribution parameters of the output propulsion force and torque are output; the composite reward function of the training module includes target proximity reward, collision avoidance safety reward and COLREGs compliance reward, wherein the collision avoidance safety reward is calculated in segments according to the ratio of the relative distance of the obstacle to the safety threshold; the training module adopts continuous n-step historical data sampling and is updated through the KL divergence constraint strategy; it also includes millimeter wave radar and GNSS modules, and the detection range of the millimeter wave radar is -60° to 60°, which is used to obtain the relative position of obstacles; The GNSS module is used to locate the coordinates of the unmanned ship in real time.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
A Collision Avoidance Method for Swarm Unmanned Surface Vessels Based on Deep Reinforcement Learning
CN110658829B
A path planning method for unmanned surface vessels based on deep reinforcement learning and taking into account marine environmental factors.
CN111829527B
Unmanned ship obstacle avoidance path planning method and system based on global optimum
CN119414851A
Unmanned ship collision avoidance method and system based on data fusion and deep reinforcement learning
CN116755444A
Unmanned ship autonomous collision avoidance decision-making method and system based on improved SAC algorithm
CN118672259A
Cited By
Unmanned fleet area coverage path planning method based on multi-constraint collaborative optimization
CN120540081A
Robot motion control method, device and equipment and storage medium
CN121411430A
Unmanned ship smooth collision avoidance method considering marine environment disturbance
CN121596884A
Smooth collision avoidance method for unmanned ship considering marine environment disturbance
CN121596884B
Multi-agent reinforcement learning unmanned ship formation collision avoidance method based on strategy sequence updating
CN122151957A