Unmanned Ship Dynamic Environment Path Planning System and Method Based on Deep Reinforcement Learning
By adopting the technology of combining deep reinforcement learning and multimodal sensors in the unmanned ship path planning system, the problems of insufficient environmental perception capabilities and low decision-making efficiency in complex dynamic water environments are solved, and the effects of efficient perception, intelligent decision-making and safe navigation are achieved.
Patent Information
- Application Number
- CN202510446319.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-10
AI Technical Summary
When facing complex and dynamic water environments, existing unmanned ship path planning technology has limited environmental perception capabilities, low decision-making efficiency, and difficult to ensure navigation safety.
The dynamic environmental path planning system of unmanned ships based on deep reinforcement learning is adopted, combined with the environment perception module, data processing module, deep reinforcement learning module, inference decision-making module and safety guarantee module, through multi-modal sensor data processing, dual-self attention mechanism feature extraction, lightweight deep reinforcement learning model training and transfer learning technology, the efficient perception, intelligent decision-making and safe navigation of unmanned ships in dynamic environments is achieved.
It improves the environmental perception accuracy and robustness of unmanned ships in dynamic environments, significantly improves decision-making efficiency and real-time, ensures navigation safety, and achieves a seamless transition from offline training to online applications.
Smart Images

Figure CN119961579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an unmanned ship dynamic environment path planning system and method based on deep reinforcement learning, which is applicable to the autonomous navigation and path planning of unmanned ships in complex dynamic water area environments. Background Art
[0002] With the rapid development of artificial intelligence and autonomous control technologies, unmanned ships, as an important type of marine intelligent equipment, have broad application prospects in fields such as marine resource investigation, marine environment monitoring, maritime search and rescue, and maritime transportation. In practical applications, unmanned ships often face complex and changeable water area environments, such as dynamic factors like suddenly emerging water area obstacles, sudden changes in water flow, and sensor interference, which pose great challenges to the autonomous navigation and path planning of unmanned ships.
[0003] Traditional path planning methods mainly include rule-based methods and optimization algorithm-based methods. Rule-based methods usually rely on predefined behavior patterns and decision rules and are difficult to handle uncertain factors in complex dynamic environments; optimization algorithm-based methods such as the A* algorithm and artificial potential field method, although they can achieve path optimization to a certain extent, often cannot meet the requirements of real-time performance and adaptability when facing highly dynamic environments.
[0004] In recent years, deep reinforcement learning, as a cutting-edge technology in the field of artificial intelligence, has demonstrated great potential in solving complex decision-making problems by combining the advantages of deep learning and reinforcement learning. However, existing path planning methods based on deep reinforcement learning still have the following problems when applied to the dynamic environment of unmanned ships: First, the environmental perception ability is limited, and it is difficult to effectively process multi-modal environmental information; second, the training efficiency of the deep reinforcement learning model is low, and there is a problem of mismatch between training and actual application scenarios; in addition, existing methods often lack effective safety guarantee mechanisms and are difficult to ensure the navigation safety of unmanned ships in complex environments.
[0005] In view of the above problems, there is an urgent need for an unmanned ship path planning method and system that can effectively cope with dynamic environments, have efficient decision-making capabilities, and ensure navigation safety. Summary of the Invention
[0006] The present invention aims to solve the problems existing in the existing unmanned ship path planning technology, such as limited environmental perception ability, low decision-making efficiency, and difficult-to-guarantee safety when facing dynamic and complex water area environments, and provides an unmanned ship dynamic environment path planning system and method based on deep reinforcement learning to achieve efficient perception, intelligent decision-making, and safe navigation of unmanned ships in dynamic environments.
[0007] The present invention provides an unmanned ship dynamic environment path planning system based on deep reinforcement learning, including:
[0008] An environmental perception module for obtaining dynamic environmental data around the unmanned ship and the status information of the unmanned ship;
[0009] A data processing module, connected to the environmental perception module, for converting the dynamic environmental data into a depth image and extracting features based on an attention mechanism;
[0010] A deep reinforcement learning module, connected to the data processing module, for making heading decisions based on the extracted features;
[0011] An inference and decision-making module, connected to the deep reinforcement learning module, for generating the optimal bow forward heading;
[0012] A safety guarantee module, connected to the inference and decision-making module, for comparing the decision-making heading with the actual heading and providing a safe heading when they are inconsistent;
[0013] The inference and decision-making module includes:
[0014] An online inference sub-module for receiving environmental observation data and performing real-time inference;
[0015] A dynamic scene recognition sub-module for classifying and recognizing potential dynamic scenes;
[0016] A heading generation sub-module for generating the optimal bow forward heading direction according to the inference result.
[0017] Preferably, the environmental perception module includes:
[0018] A multi-modal sensor sub-module for collecting environmental data including prominent water area obstacles, sudden changes in water flow, and sensor interference;
[0019] A status information collection sub-module for obtaining status information including the position, speed, and attitude of the unmanned ship;
[0020] A data fusion sub-module for performing spatio-temporal alignment and fusion of the multi-modal sensor data and the status information.
[0021] Preferably, the data processing module includes:
[0022] An image conversion sub-module for converting the dynamic environmental data into a heading depth image;
[0023] A feature extraction sub-module designed based on a dual self-attention mechanism for extracting depth features from the heading depth image;
[0024] A temporal encoding sub-module using an LSTM network structure for encoding and processing temporal features.
[0025] Preferably, the dual self-attention mechanism of the feature extraction sub-module is implemented through the following steps:
[0026] Perform global average pooling on the input features;
[0027] Divide the input features into a first data subset and a second data subset;
[0028] Calculate the first weight, the second weight, and the third weight respectively;
[0029] Perform weighted processing on the data subsets based on the weights;
[0030] Output the processing result as the sixth intermediate data through the feature extractor.
[0031] Preferably, the deep reinforcement learning module includes:
[0032] A value evaluation network, using the LightGBM model, for evaluating the value of each possible action in the current state;
[0033] An action decision network, for selecting the optimal heading decision based on the value evaluation result;
[0034] A training optimization sub-module, for optimizing and training the network through a composite reward function, and the composite reward function takes into account the lane reward, the environmental perception reward, the obstacle avoidance reward, and the heading stability reward.
[0035] Preferably, the value evaluation network is trained using the TD algorithm, and the value update formula is:
[0036] ,
[0037] where represents the state value at time t, represents the state value at time t+1.
[0038] Preferably, the calculation time of the inference and decision-making process of the inference and decision-making module is less than 100 ms.
[0039] Preferably, the safety guarantee module includes:
[0040] A heading comparison sub-module, for comparing the inference and decision-making heading with the actual control heading of the actual ship;
[0041] A safe heading generation sub-module, for generating a preset safe heading when the headings are inconsistent;
[0042] A rudder angle compensation sub-module, for performing steering adjustment compensation in the movement direction of the unmanned ship.
[0043] Preferably, it further includes:
[0044] A transfer learning module, connected between the deep reinforcement learning module and the inference decision-making module, is used to transfer the neural network model parameters trained offline to the online inference decision-making model;
[0045] A parameter adaptive sub-module is used to fine-tune the transferred parameters according to the differences between the actual environment and the simulation environment;
[0046] An execution feedback sub-module is used to collect execution results and adjust the parameters of the inference decision-making module based on error feedback.
[0047] An unmanned ship dynamic environment path planning method based on deep reinforcement learning, applied to the said system, includes:
[0048] An environment perception step, which acquires the dynamic environment data around the unmanned ship and the state information of the unmanned ship;
[0049] A data processing step, which converts the dynamic environment data into a depth image and extracts features based on the attention mechanism;
[0050] A deep reinforcement learning step, which trains a heading decision neural network model based on the extracted features;
[0051] A model transfer step, which transfers the trained neural network model parameters to the online inference decision-making model;
[0052] An online inference step, which receives environment observation data and conducts real-time inference to generate the optimal bow forward heading;
[0053] A safety guarantee step, which compares the decision-making heading with the actual heading and provides a safe heading when they are inconsistent.
[0054] Through the organic combination of deep reinforcement learning and the attention mechanism, combined with transfer learning technology and a multi-layer safety guarantee mechanism, the present invention realizes the adaptive path planning of an unmanned ship in a dynamic environment, and has the following beneficial effects:
[0055] 1. The data processing module based on the dual self-attention mechanism can effectively extract key features in the dynamic environment, improving the accuracy and robustness of environment perception;
[0056] 2. Adopting a lightweight deep reinforcement learning model design, combined with an optimized composite reward function, significantly improves the decision-making efficiency and realizes intelligent decision-making under real-time requirements;
[0057] 3. Through transfer learning technology, the problem of differences between the training environment and the actual application environment is solved, realizing a seamless transition from offline training to online application;
[0058] 4. The design of the multi - layer security mechanism ensures the navigation safety of the unmanned ship in various complex situations and greatly improves the reliability of the system;
[0059] 5. The modular system architecture design enables the system to have good scalability and maintainability, facilitating subsequent function upgrades and optimizations. Brief Description of the Drawings
[0060] Figure 1 It is the structural block diagram of the unmanned ship dynamic environment path planning system based on deep reinforcement learning of the present invention;
[0061] Figure 2 It is the structural block diagram of the environment perception module of the present invention;
[0062] Figure 3 It is the structural block diagram of the data processing module of the present invention;
[0063] Figure 4 It is the implementation flowchart of the double self - attention mechanism in the feature extraction sub - module of the present invention;
[0064] Figure 5 It is the structural block diagram of the deep reinforcement learning module of the present invention;
[0065] Figure 6 It is the structural block diagram of the inference and decision - making module of the present invention;
[0066] Figure 7 It is the structural block diagram of the security guarantee module of the present invention;
[0067] Figure 8 It is the structural block diagram of the transfer learning module of the present invention;
[0068] Figure 9 It is the flowchart of the method for path planning of the unmanned ship dynamic environment based on deep reinforcement learning of the present invention. Detailed Embodiments
[0069] Next, in combination with specific embodiments, the technical solutions of the present invention will be described in detail.
[0070] Embodiment 1: Overall System Architecture
[0071] As Figure 1 shown, the unmanned ship dynamic environment path planning system based on deep reinforcement learning provided by the present invention includes an environment perception module 1, a data processing module 2, a deep reinforcement learning module 3, an inference and decision - making module 4, a security guarantee module 5, and a transfer learning module 6.
[0072] The environmental perception module 1 is used to obtain the dynamic environmental data around the unmanned ship and the status information of the unmanned ship. Specifically, the environmental perception module 1 receives data from a variety of sensors, including but not limited to environmental data and ship status data collected by devices such as cameras, lidars, sonars, and GPSs.
[0073] The data processing module 2 is connected to the environmental perception module 1 and is used to convert the dynamic environmental data into a depth image and extract features based on the attention mechanism. Preferably, the data processing module 2 adopts a dual self-attention mechanism, which can adaptively focus on key regions and features in the environment and improve the effectiveness of feature extraction.
[0074] The deep reinforcement learning module 3 is connected to the data processing module 2 and is used to make a heading decision based on the extracted features. This module selects the optimal navigation strategy by evaluating the value functions of different headings. The present invention adopts a lightweight network design and an optimized training algorithm, which significantly improves the decision-making efficiency.
[0075] The inference and decision-making module 4 is connected to the deep reinforcement learning module 3 and is used to generate the best forward heading of the bow. This module infers the dynamic scene in the real-time environment and makes a decision on the best heading direction accordingly, providing guidance for the navigation of the unmanned ship.
[0076] The safety guarantee module 5 is connected to the inference and decision-making module 4 and is used to compare the decision-making heading with the actual heading and provide a safe heading when they are inconsistent. This module ensures that the unmanned ship can still sail safely even in a complex environment or when there is a deviation in model inference.
[0077] The transfer learning module 6 is connected between the deep reinforcement learning module 3 and the inference and decision-making module 4 and is used to transfer the neural network model parameters trained offline to the online inference and decision-making model to solve the problem of the difference between the training environment and the actual application environment.
[0078] Embodiment 2: Specific implementation of the environmental perception module
[0079] As Figure 2 shown, the environmental perception module 1 includes a multi-modal sensor sub-module 11, a status information acquisition sub-module 12, and a data fusion sub-module 13.
[0080] The multimodal sensor sub-module 11 is used to collect environmental data including prominent water area obstacles, sudden changes in water flow, and sensor interference. This sub-module integrates a variety of sensing devices, including optical cameras, depth cameras, lidar, millimeter-wave radars, sonars, etc. Preferably, the resolution of the optical camera is 1920×1080 pixels, and the acquisition frequency is 30 frames per second; the scanning range of the lidar is 360°, the scanning frequency is 10Hz, and the ranging accuracy is ±2cm; the detection range of the sonar is 500 meters, and the resolution is 0.1 meter. Through the collaborative work of multiple sensors, the environmental conditions around the unmanned boat can be perceived in all directions.
[0081] The status information acquisition sub-module 12 is used to obtain status information including the position, speed, and attitude of the unmanned boat. This sub-module mainly obtains accurate position, heading, speed, attitude, etc. information of the unmanned boat through devices such as GPS, inertial measurement unit (IMU), and electronic compass. Preferably, the GPS positioning accuracy is better than 0.5 meters, and the angle measurement accuracy of the IMU is better than 0.1°, and the measurement frequency is 100Hz. These high-precision status information provides reliable basic data for subsequent path planning.
[0082] The data fusion sub-module 13 is used to perform spatio-temporal alignment and fusion of multimodal sensor data and status information. This sub-module uses the extended Kalman filter algorithm for the fusion processing of multi-source heterogeneous data, effectively solving the problems of time synchronization and spatial registration of different sensor data. Preferably, the processing frequency of data fusion is 50Hz, which can meet the real-time requirements. The data fusion processing can be expressed as:
[0083] ,
[0084] ,
[0085] where, represents the state vector at time k, represents the state transition matrix, represents the control input matrix, represents the control vector, represents the process noise, represents the observation vector, represents the observation matrix, represents the observation noise. In this way, the environmental perception module 1 can provide high-quality fusion data, providing reliable input for subsequent processing.
[0086] Example 3: Specific implementation of the data processing module
[0087] As Figure 3 shown, the data processing module 2 includes an image conversion sub-module 21, a feature extraction sub-module 22, and a temporal coding sub-module 23.
[0088] The image conversion sub-module 21 is used to convert dynamic environment data into a heading depth image. Specifically, this sub-module converts the multi-modal data collected by the environment perception module 1 into a depth image in a unified format, where the depth value represents the heading angle. Preferably, the resolution of the converted depth image is 60×40 pixels, the depth value range is 0-255, and the corresponding heading angle range is 0-360 degrees. This conversion method gives the environmental data a spatial expression form, facilitating subsequent feature extraction.
[0089] The feature extraction sub-module 22 is designed based on a dual self-attention mechanism and is used to extract depth features from the heading depth image. As Figure 4 shown, the dual self-attention mechanism is implemented through the following steps: First, perform global average pooling on the input feature X to obtain the global feature G; then, divide the input feature into a first data subset and a second data subset Next, calculate the first weight , the second weight and the third weight respectively through a neural network. Then, perform weighted processing on the data subsets based on these weights to obtain intermediate results , , , and ; Finally, pass through a feature extractor, and the output is the sixth intermediate data 0. This process can be expressed as:
[0090] ,
[0091] ,
[0092] ,
[0093] ,
[0094] ,
[0095] ,
[0096] where AvgPool represents the global average pooling operation, and represent the feature subset division function, , , and Denote different neural networks, and Extractor denotes the feature extractor. Through this dual self-attention mechanism, the system can adaptively focus on key regions and features in the environment, greatly improving the effectiveness of feature extraction.
[0097] The time series encoding sub-module 23 adopts an LSTM network structure for encoding time series features. The introduction of the LSTM network enables the system to consider historical information, capture the time series characteristics of environmental changes, and provide a more comprehensive information basis for decision-making. Preferably, the hidden layer dimension of the LSTM network is 128, and the time series step is 10, which can effectively process continuous environmental changes. The calculation process of the LSTM cell can be expressed as:
[0098] ,
[0099] ,
[0100] ,
[0101] ,
[0102] ,
[0103] ,
[0104] Among them, 、 and represent the forget gate, input gate, and output gate respectively, represents the cell state, represents the hidden state, 、 、 and represent weight matrices, 、 、 and represent bias terms, represents the sigmoid activation function, represents the hyperbolic tangent activation function.
[0105] Through this design, the data processing module 2 can effectively extract key features in the environment and consider time series information, providing high-quality feature representations for subsequent decision-making.
[0106] Example 4: Specific implementation of the deep reinforcement learning module
[0107] As Figure 5 shown, the deep reinforcement learning module 3 includes a value evaluation network 31, an action decision network 32, and a training optimization sub-module 33.
[0108] The value evaluation network 31 adopts the LightGBM model to evaluate the value of each possible action in the current state. LightGBM is an efficient gradient boosting decision tree framework, which features fast training speed and low memory occupancy, and is very suitable for scenarios that require real-time response. Preferably, the maximum depth of the decision tree is set to 8, the number of leaf nodes is set to 31, and the learning rate is set to 0.05. These parameters have been experimentally verified to maintain high computational efficiency while ensuring accuracy.
[0109] The action decision network 32 is used to select the optimal course decision based on the value evaluation result. This network adopts the Actor-Critic architecture, where the Actor network generates the policy and the Critic network evaluates the policy. Preferably, the Actor network consists of 3 fully connected layers, with the number of neurons in each layer being 256, 128, and 64 respectively; the Critic network also consists of 3 fully connected layers, with the number of neurons in each layer being 256, 128, and 1 respectively. Both networks use ReLU as the activation function, except for the last layer.
[0110] The training optimization sub-module 33 is used to optimize and train the network through a composite reward function. The composite reward function takes into account the lane reward, environmental perception reward, obstacle avoidance reward, and course stability reward, and can be expressed as:
[0111] ,
[0112] where represents the lane reward, which considers whether the unmanned ship is sailing within the predetermined lane; represents the environmental perception reward, which is adjusted according to the perception accuracy, and the calculation formula is:
[0113] ,
[0114] where and are the variances of the observed target and the obstacle respectively, represents the Euclidean distance between them; represents the obstacle avoidance reward, which is related to the collision risk; represents the course stability reward, which is used to avoid frequent course changes.
[0115] The training optimization sub-module 33 is used to optimize and train the network through a composite reward function. The composite reward function takes into account the lane reward, environmental perception reward, obstacle avoidance reward, and course stability reward, and can be expressed as:
[0116] The value evaluation network 31 is trained using the TD algorithm, which is a temporal difference learning method that can update the value estimate of the current state using the estimate of subsequent states. The value update formula is:
[0117] ,
[0118] where, represents the state value at time t, represents the state value at time t+1. Through the TD algorithm, the system can effectively learn the state value function and thus guide decision-making behaviors.
[0119] ,
[0120] where, represents the expectation sampled from the trajectory dataset D, represents the discount factor, represents the immediate reward obtained by executing the action in the state , λ represents the entropy regularization coefficient, represents the policy entropy. By introducing the entropy regularization term, the exploration behavior of the policy can be encouraged to avoid premature convergence to the local optimal solution.
[0121] With this design, the deep reinforcement learning module 3 can efficiently learn the optimal decision-making policy, meeting the real-time requirements while ensuring the decision-making quality.
[0122] Example 5: Specific implementation of the inference and decision-making module
[0123] As Figure 6 shown, the inference and decision-making module 4 includes an online inference sub-module 41, a dynamic scenario recognition sub-module 42, and a heading generation sub-module 43.
[0124] The online inference sub-module 41 is used to receive environmental observation data and perform real-time inference. This sub-module receives real-time environmental observation data from the environmental perception module 1, and after being processed by the data processing module 2, it is input into the inference model migrated by the transfer learning module 6 for inference. Preferably, the calculation time of the inference process is controlled within 5 ms to ensure that the system can quickly respond to environmental changes.
[0125] The dynamic scene recognition sub-module 42 is used to classify and recognize potential dynamic scenes. Based on the extracted features, this sub-module identifies the types of dynamic scenes in the current environment, such as static obstacles, moving obstacles, water flow changes, etc. Preferably, the scene classification uses a SoftMax classifier, and the number of categories is set to 10, including obstacle-free scenes, static obstacle scenes, single moving obstacle scenes, multiple moving obstacle scenes, sudden water flow change scenes, etc. The calculation of scene classification can be expressed as:
[0126] ,
[0127] where, represents the probability that the sample x belongs to the category , represents the score of the category . Through scene recognition, the system can adopt corresponding decision-making strategies for different types of scenes.
[0128] The heading generation sub-module 43 is used to generate the best bow forward heading direction according to the inference result. Based on the results of online inference and scene recognition, this sub-module comprehensively considers factors such as safety and efficiency to generate the best heading instruction. Preferably, the accuracy of the heading angle is 0.1°, which is sufficient to meet the requirements of fine navigation.
[0129] It should be noted that the calculation time of the entire inference and decision-making process is less than 100 ms. This performance index ensures that the system can make timely responses in a dynamic environment and meet the requirements of real-time control. Through this design, the inference and decision-making module 4 can quickly generate the best heading decision based on real-time environmental observations, providing reliable guidance for the autonomous navigation of the unmanned ship.
[0130] Example 6: Specific implementation of the safety guarantee module
[0131] As Figure 7 shown, the safety guarantee module 5 includes a heading comparison sub-module 51, a safe heading generation sub-module 52, and a rudder angle compensation sub-module 53.
[0132] The heading comparison sub-module 51 is used to compare the inference and decision-making heading with the actual control heading of the real ship. This sub-module receives the heading instruction from the inference and decision-making module 4 and the actual heading feedback from the unmanned ship's own control system for real-time comparison. Preferably, the frequency of heading comparison is 10 Hz, that is, a comparison is made every 100 ms, which can timely detect heading deviations.
[0133] The safe heading generation sub-module 52 is used to generate a preset safe heading when the headings are inconsistent. When the heading comparison result shows that the deviation between the decision-making heading and the actual heading exceeds the preset threshold, this sub-module will generate a safe heading instruction to overwrite the original decision. Preferably, the heading deviation threshold is set to 5°, which is based on experimental verification and can avoid excessive intervention while ensuring sensitivity. The generation strategy of the safe heading takes into account factors such as the current environmental conditions, the motion state of the unmanned ship, and the navigation rules, etc., to ensure that the generated heading is safe and feasible.
[0134] The rudder angle compensation sub-module 53 is used to perform steering adjustment compensation in the motion direction of the unmanned ship. This sub-module calculates the required rudder angle compensation amount according to the safe heading instruction, and adjusts the rudder angle to make the unmanned ship gradually turn to the safe heading. Preferably, the rudder angle compensation adopts the proportional-integral-derivative (PID) control algorithm, and the control parameters , this set of parameters has been tested on the actual ship and can provide smooth and effective steering performance. The PID control algorithm can be expressed as:
[0135] ,
[0136] where, represents the control output, that is, the rudder angle adjustment amount; represents the error, that is, the difference between the current heading and the target heading; , and respectively represent the proportional, integral and derivative coefficients.
[0137] Through this multi-level safety guarantee mechanism, even in complex environments or when there are deviations in model inference, the system can still ensure the navigation safety of the unmanned ship, significantly improving the reliability and robustness of the system.
[0138] Embodiment 7: Specific implementation of the transfer learning module
[0139] As Figure 8 shown, the transfer learning module 6 includes a parameter transfer sub-module 61, a parameter adaptation sub-module 62, and an execution feedback sub-module 63.
[0140] The parameter transfer sub-module 61 is used to transfer the neural network model parameters trained offline to the online inference decision model. This sub-module realizes the transfer of the model parameters trained by the deep reinforcement learning module 3 to the model parameters used by the inference decision module 4. Preferably, the parameter transfer adopts a combination of direct copying and fine-tuning, that is, directly copy the parameters for the shared network structure part, and perform parameter adaptation and fine-tuning for the parts with different structures.
[0141] The parameter adaptive sub-module 62 is used to finely tune the migration parameters according to the differences between the actual environment and the simulation environment. This sub-module monitors the performance of the model in the actual environment, identifies the differences between the simulation environment and the actual environment, and accordingly makes targeted adjustments to the model parameters. Preferably, the adaptive adjustment adopts an online learning method, with the learning rate set to 0.01, and the parameters are updated once every 100 samples are processed. This setting ensures the adjustment effect while avoiding the instability caused by overly frequent updates.
[0142] The execution feedback sub-module 63 is used to collect the execution results and adjust the parameters of the inference decision module based on error feedback. This sub-module monitors the actual effect after the unmanned ship executes the decision, calculates the error between the execution result and the expected result, and adjusts the model parameters based on this error. Preferably, the error feedback adopts the gradient descent method, with the learning rate set to 0.001 and the update frequency of once every 10 minutes. This setting can achieve the continuous optimization of the model while ensuring stability.
[0143] Through the design of the transfer learning module 6, the system solves the problem of differences between the training environment and the actual application environment, realizes the seamless transition from offline training to online application, and greatly improves the adaptability and performance of the model in the actual environment.
[0144] Embodiment 8: Unmanned Ship Dynamic Environment Path Planning Method Based on Deep Reinforcement Learning
[0145] As Figure 9 shown, the present invention also provides an unmanned ship dynamic environment path planning method based on deep reinforcement learning, which is applied to the above system. This method includes an environment perception step, a data processing step, a deep reinforcement learning step, a model transfer step, an online inference step, and a safety guarantee step.
[0146] The environment perception step is to obtain the dynamic environment data around the unmanned ship and the state information of the unmanned ship. Specifically, the environment data and state information are obtained through a variety of sensors and subjected to data fusion processing.
[0147] The data processing step is to convert the dynamic environment data into a depth image and extract features based on the attention mechanism. Specifically, the environment data is converted into a depth image in a unified format, and then the key features are extracted through a dual self-attention mechanism, and the temporal information is processed through an LSTM network.
[0148] The deep reinforcement learning step is to train a heading decision neural network model based on the extracted features. Specifically, a value evaluation network and an action decision network are designed and optimized and trained through a composite reward function and a TD algorithm.
[0149] Model migration step: Transfer the parameters of the trained neural network model to the online inference and decision-making model. Specifically, through parameter transfer and adaptive fine-tuning, the model is migrated from the training environment to the actual environment.
[0150] Online inference step: Receive environmental observation data and perform real-time inference to generate the optimal bow forward course. Specifically, based on the migrated model, online inference is performed to identify dynamic scenarios and generate the optimal course instruction.
[0151] Safety guarantee step: Compare the decision-making course with the actual course and provide a safe course when they are inconsistent. Specifically, through course comparison, safe course generation, and rudder angle compensation, the navigation safety of the unmanned ship is ensured.
[0152] Example 9: Application of the system in the actual water area environment
[0153] The application of the system of the present invention in the actual water area environment demonstrates its superior performance and adaptability. In this embodiment, a certain bay area is selected as the test scenario, which has various typical complex water area characteristics, including reefs, channels, water flow change areas, and other vessel activity areas, etc.
[0154] During the test, the unmanned ship was equipped with the complete system of the present invention and passed through multiple preset test points in sequence. During the test, various dynamic challenge scenarios were artificially set, including suddenly appearing floating objects, small vessels passing through quickly, and artificially created water flow disturbances, etc.
[0155] The test results show that the system of the present invention can successfully cope with various dynamic challenge scenarios while meeting the real-time requirement (inference and decision-making time is less than 100 ms). Compared with traditional methods, the system shows significant advantages in terms of navigation safety, path planning efficiency, and environmental adaptability: in terms of safety, the system successfully avoided 100% of the dynamic obstacles; in terms of efficiency, compared with traditional algorithms, the path length was shortened by an average of 23.5%, and the fuel consumption was reduced by 18.7%; in terms of adaptability, the system can automatically adjust the decision-making strategy to adapt to different water flow conditions and traffic densities, and the adaptability is significantly enhanced.
[0156] Preferably, in actual applications, the parameter configuration of the system can be adjusted according to specific application scenarios and requirements. For example, in open waters, the ship speed can be appropriately increased and the obstacle avoidance sensitivity can be reduced, while in narrow channels or port areas, the ship speed should be reduced and the obstacle avoidance sensitivity should be increased. Through this flexible parameter configuration, the system can better adapt to various complex actual application scenarios.
[0157] Example 10: Performance evaluation and optimization of the system
[0158] The performance evaluation of the system of the present invention is mainly carried out from four aspects: environmental perception ability, decision-making efficiency, safety, and robustness.
[0159] In terms of environmental perception ability, the system realizes efficient perception of complex dynamic environments through multi-modal sensor fusion and dual self-attention mechanisms. Tests show that under various lighting conditions, the detection accuracy of the system for obstacles reaches 96.8%, far higher than that of traditional single-sensor solutions (usually 82% - 85%). Especially under harsh conditions such as low light and foggy days, the system can still maintain a high perception accuracy, which benefits from the complementarity of multi-modal sensing data and the effectiveness of the attention mechanism.
[0160] In terms of decision-making efficiency, the system adopts a lightweight deep reinforcement learning model design and combines an optimized TD algorithm, significantly improving the decision-making speed. Test data shows that the average decision-making time of the system is 42ms, and the maximum decision-making time does not exceed 85ms, fully meeting the requirements of real-time control (usually required to be less than 100ms). Compared with traditional deep reinforcement learning methods, the computational efficiency is improved by about 65%, while maintaining comparable decision-making quality.
[0161] In terms of safety, the multi-layer safety guarantee mechanism of the system ensures the safe navigation of the unmanned ship in various situations. In test scenarios containing sudden risks, the system can identify risks and execute safe course adjustments within 0.5 seconds, avoiding potential collision risks. The intervention rate of the safety guarantee mechanism is about 8.3%, indicating that the system can ensure safety relying on the normal decision-making process in most cases, and only requires the intervention of the safety mechanism in a few high-risk situations.
[0162] In terms of robustness, the system demonstrates good adaptability to environmental changes through transfer learning and parameter self-adaptive techniques. In tests under different water conditions, the performance degradation of the system does not exceed 7%, while traditional methods usually experience a 20% - 30% performance decline when the environment changes. This result proves the advantage of the present system in dealing with environmental variability.
[0163] Based on the performance evaluation results, the system has been further optimized. The main optimization measures include: adjusting the reward function for specific scenarios, fine-tuning the parameters of the attention mechanism, and optimizing the setting of safety thresholds. These optimization measures further improve the performance of the system while maintaining its original advantages, especially the stability and reliability in extreme situations.
[0164] The unmanned ship dynamic environment path planning system and method based on deep reinforcement learning provided by the present invention have broad industrial application prospects. First, the system can be applied to fields such as marine resource investigation, marine environment monitoring, and maritime search and rescue, greatly improving operation efficiency and safety. Second, in terms of maritime traffic and port management, the system can be used for intelligent ship guidance and traffic coordination. In addition, the system can also be extended to other autonomous navigation scenarios, such as the fields of drones and unmanned vehicles.
[0165] The core advantages of the system of the present invention lie in its adaptability to complex dynamic environments and its safe and reliable decision-making ability, which enable it to play an important role in actual industrial applications. Compared with traditional methods, the present system has significant advantages in aspects such as environmental perception, decision-making efficiency, safety, and robustness, and can better meet the actual application requirements.
[0166] In summary, the unmanned ship dynamic environment path planning system and method based on deep reinforcement learning provided by the present invention, through the combination of deep reinforcement learning and the attention mechanism, as well as the design of transfer learning and multi-layer security guarantee mechanisms, achieve efficient perception, intelligent decision-making, and safe navigation of unmanned ships in dynamic environments, providing a new solution for the development and application of unmanned ship technology.
[0167] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An unmanned ship dynamic environment path planning system based on deep reinforcement learning, characterized by: include: Environmental perception module, used to obtain dynamic environmental data around the unmanned ship and the status information of the unmanned ship; A data processing module, connected to the environment perception module, for converting the dynamic environment data into a depth image and extracting features based on an attention mechanism; A deep reinforcement learning module, connected to the data processing module, for making heading decisions based on the extracted features; A reasoning and decision-making module, connected to the deep reinforcement learning module, for generating an optimal bow forward heading; A safety assurance module, connected to the reasoning decision module, for comparing the decision heading with the actual heading and providing a safe heading when there is a discrepancy; The data processing module comprises: An image conversion submodule, used for converting the dynamic environment data into a heading depth image; A feature extraction submodule, designed based on a dual self-attention mechanism, for extracting depth features from the heading depth image; The time series encoding submodule adopts the LSTM network structure to encode the time series features; The dual self-attention mechanism is implemented through the following steps: First, the input feature X is globally averaged pooled to obtain the global feature G; then, the input feature is divided into the first data subset and the second data subset Next, the first weights are calculated through the neural network. , second weight and the third weight Then, based on these weights, the data subsets are weighted to obtain the intermediate results , , , and ; Finally, Through the feature extractor, the output is the sixth intermediate data 0. This process can be expressed as: , , , , , , Among them, AvgPool represents the global average pooling operation, and represents the feature subset partitioning function, , , and Represents different neural networks, and Extractor represents feature extractor; The calculation process of the LSTM unit of the LSTM network structure can be expressed as: , , , , , , in, , and They represent the forget gate, input gate and output gate respectively. Indicates the unit status, Indicates the hidden state, , , and represents the weight matrix, , , and represents the bias term, represents the sigmoid activation function, represents the hyperbolic tangent activation function.
2. The system according to claim 1, characterized in that The environment perception module comprises: Multimodal sensor submodule for collecting environmental data including highlighting water obstacles, sudden changes in water flow, and sensor interference; The status information acquisition submodule is used to obtain status information including the position, speed and attitude of the unmanned ship; The data fusion submodule is used to align and fuse the multimodal sensor data with the state information in time and space.
3. The system according to claim 1, characterized in that The deep reinforcement learning module includes: The value evaluation network uses the LightGBM model to evaluate the value of each possible action in the current state; an action decision network, for selecting an optimal heading decision based on the value assessment result; The training optimization submodule is used to optimize the training of the network through a composite reward function, wherein the composite reward function considers a channel reward, an environment perception reward, an obstacle avoidance reward, and a heading stability reward.
4. The system according to claim 3, characterized in that The value evaluation network is trained using the TD algorithm, and the value update formula is: , in, represents the state value at time t, Indicates the state value at time t+1.
5. The system according to claim 1, characterized in that The reasoning decision module includes: The online reasoning submodule is used to receive environmental observation data and perform real-time reasoning; Dynamic scene recognition submodule, used to classify and recognize potential dynamic scenes; A heading generation submodule is used to generate the best bow forward heading direction according to the reasoning result; The calculation time of the inference decision process is less than 100ms.
6. The system according to claim 1, characterized in that The security module comprises: The heading comparison submodule is used to compare the inference decision heading with the actual control heading of the actual ship; A safety heading generation submodule is used to generate a preset safety heading when the headings are inconsistent; The rudder angle compensation submodule is used to perform steering adjustment compensation in the direction of movement of the unmanned ship.
7. The system according to claim 1, characterized in that Also includes: A transfer learning module, connected between the deep reinforcement learning module and the reasoning and decision-making module, for migrating the parameters of the offline trained neural network model to the online reasoning and decision-making model; The parameter adaptation submodule is used to fine-tune the migration parameters according to the difference between the actual environment and the simulation environment; The execution feedback submodule is used to collect execution results and adjust the parameters of the reasoning decision module based on error feedback.
8. A method for unmanned ship dynamic environment path planning based on deep reinforcement learning, using the system described in any one of claims 1 to 7, characterized in that: include: Environmental perception step, obtaining dynamic environmental data around the unmanned ship and the status information of the unmanned ship; A data processing step, converting the dynamic environment data into a depth image and extracting features based on an attention mechanism; Deep reinforcement learning step, training the heading decision neural network model based on the extracted features; Model migration step, migrating the trained neural network model parameters to the online inference decision model; The online reasoning step receives environmental observation data and performs real-time reasoning to generate the optimal bow heading; The safety assurance step compares the decision heading with the actual heading and provides a safe heading when there is a discrepancy.
Citation Information
Patent Citations
Unmanned ship collision avoidance method and device based on DBSCAN algorithm
CN119763371A
Apparatus, method, and recording medium for autonomous ship navigation
US20210114698A1