Unmanned ship dynamic environment path planning system and method based on deep reinforcement learning

By introducing deep reinforcement learning and multi-layer safety assurance mechanisms into the unmanned ship path planning system, the problems of insufficient environmental perception ability, low decision efficiency and poor navigation safety of unmanned ship path planning in complex dynamic water environments are solved, and the effects of efficient perception, intelligent decision-making and safe navigation are achieved.

CN119961579AActive Publication Date: 2025-05-09JIMEI UNIV

Patent Information

Application Number
CN202510446319.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

When facing complex and dynamic water environments, existing unmanned ship path planning technology has limited environmental perception capabilities, low decision-making efficiency, and difficult to ensure navigation safety.

Method used

The dynamic environmental path planning system of unmanned ships based on deep reinforcement learning is adopted, combined with the environment perception module, data processing module, deep reinforcement learning module, reasoning decision-making module and safety assurance module, and features are extracted through multi-modal sensor data processing and dual self-attention mechanism, and intelligent decision-making is made using deep reinforcement learning, and navigation safety is ensured through multi-layer safety assurance mechanisms.

Benefits of technology

It realizes efficient perception, intelligent decision-making and safe navigation of unmanned ships in dynamic environments, improves the accuracy and robustness of environmental perception, and enhances decision-making efficiency and navigation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961579A_ABST
    Figure CN119961579A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an unmanned ship dynamic environment path planning system and method based on deep reinforcement learning, and the system comprises an environment sensing module which is used for obtaining the dynamic environment data around an unmanned ship and the state information of the unmanned ship; the data processing module is connected with the environment sensing module and used for converting the dynamic environment data into a depth image and extracting features based on an attention mechanism; the deep reinforcement learning module is connected with the data processing module and is used for performing course decision based on the extracted features; the reasoning decision module is connected with the deep reinforcement learning module and is used for generating an optimal ship bow forward course; the safety guarantee module is connected with the reasoning decision module and used for comparing the decision course with the actual course and providing the safe course when the decision course is inconsistent with the actual course, the design of a multi-layer safety guarantee mechanism ensures the sailing safety of the unmanned ship under various complex conditions, and the reliability of the system is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a system and method for unmanned ship dynamic environment path planning based on deep reinforcement learning, which is suitable for autonomous navigation and path planning of unmanned ships in complex dynamic water environments. Background Art

[0002] With the rapid development of artificial intelligence and autonomous control technology, unmanned ships, as an important marine intelligent equipment, have broad application prospects in the fields of marine resource survey, marine environment monitoring, maritime search and rescue, and maritime transportation. In practical applications, unmanned ships often face complex and changeable water environments, such as prominent water obstacles, sudden changes in water flow, sensor interference and other dynamic factors, which bring great challenges to the autonomous navigation and path planning of unmanned ships.

[0003] Traditional path planning methods mainly include rule-based methods and optimization algorithm-based methods. Rule-based methods usually rely on pre-defined behavior patterns and decision rules, and are difficult to deal with uncertainties in complex dynamic environments; optimization algorithm-based methods such as the A* algorithm and artificial potential field method can achieve path optimization to a certain extent, but in the face of highly dynamic environments, their real-time and adaptability are often difficult to meet requirements.

[0004] In recent years, deep reinforcement learning, as a cutting-edge technology in the field of artificial intelligence, has shown great potential in solving complex decision-making problems by combining the advantages of deep learning and reinforcement learning. However, the existing path planning methods based on deep reinforcement learning still have the following problems when applied to the dynamic environment of unmanned ships: first, the environmental perception ability is limited, and it is difficult to effectively process multimodal environmental information; second, the training efficiency of deep reinforcement learning models is low, and there is a problem of mismatch between training and actual application scenarios; in addition, existing methods often lack effective safety assurance mechanisms, making it difficult to ensure the navigation safety of unmanned ships in complex environments.

[0005] In view of the above problems, there is an urgent need for an unmanned ship path planning method and system that can effectively cope with dynamic environments, have efficient decision-making capabilities and ensure navigation safety. Summary of the invention

[0006] The present invention aims to solve the problems of limited environmental perception ability, low decision-making efficiency, and difficulty in ensuring safety in the existing unmanned ship path planning technology when facing dynamic and complex water environments, and to provide an unmanned ship dynamic environment path planning system and method based on deep reinforcement learning to achieve efficient perception, intelligent decision-making and safe navigation of unmanned ships in dynamic environments.

[0007] The present invention provides an unmanned ship dynamic environment path planning system based on deep reinforcement learning, comprising:

[0008] Environmental perception module, used to obtain dynamic environmental data around the unmanned ship and the status information of the unmanned ship;

[0009] A data processing module, connected to the environment perception module, for converting the dynamic environment data into a depth image and extracting features based on an attention mechanism;

[0010] A deep reinforcement learning module, connected to the data processing module, for making heading decisions based on the extracted features;

[0011] A reasoning and decision-making module, connected to the deep reinforcement learning module, for generating an optimal bow forward heading;

[0012] A safety assurance module, connected to the reasoning decision module, for comparing the decision heading with the actual heading and providing a safe heading when there is a discrepancy; The reasoning decision module includes: The online reasoning submodule is used to receive environmental observation data and perform real-time reasoning; Dynamic scene recognition submodule, used to classify and recognize potential dynamic scenes; The heading generation submodule is used to generate the optimal bow forward heading direction according to the reasoning result.

[0013] Preferably, the environment perception module comprises:

[0014] Multimodal sensor submodule for collecting environmental data including highlighting water obstacles, sudden changes in water flow, and sensor interference;

[0015] The status information acquisition submodule is used to obtain status information including the position, speed and attitude of the unmanned ship;

[0016] The data fusion submodule is used to align and fuse the multimodal sensor data with the state information in time and space.

[0017] Preferably, the data processing module comprises:

[0018] An image conversion submodule, used for converting the dynamic environment data into a heading depth image;

[0019] A feature extraction submodule, designed based on a dual self-attention mechanism, for extracting depth features from the heading depth image;

[0020] The time series encoding submodule adopts the LSTM network structure to encode the time series features.

[0021] Preferably, the dual self-attention mechanism of the feature extraction submodule is implemented by the following steps:

[0022] Perform global average pooling on the input features;

[0023] Dividing the input features into a first data subset and a second data subset;

[0024] Calculate the first weight, the second weight and the third weight respectively;

[0025] Performing weighted processing on the data subset based on the weight;

[0026] The processing result is output as sixth intermediate data through the feature extractor.

[0027] Preferably, the deep reinforcement learning module comprises:

[0028] The value evaluation network uses the LightGBM model to evaluate the value of each possible action in the current state;

[0029] an action decision network, for selecting an optimal heading decision based on the value assessment result;

[0030] The training optimization submodule is used to optimize the training of the network through a composite reward function, wherein the composite reward function considers a channel reward, an environment perception reward, an obstacle avoidance reward, and a heading stability reward.

[0031] Preferably, the value evaluation network is trained using the TD algorithm, and the value update formula is:

[0032] ,

[0033] in, represents the state value at time t, Indicates the state value at time t+1.

[0034] Preferably, the calculation time of the reasoning and decision-making process of the reasoning and decision-making module is less than 100 ms.

[0035] Preferably, the security assurance module comprises:

[0036] The heading comparison submodule is used to compare the inference decision heading with the actual control heading of the actual ship;

[0037] A safety heading generation submodule is used to generate a preset safety heading when the headings are inconsistent;

[0038] The rudder angle compensation submodule is used to perform steering adjustment compensation in the direction of movement of the unmanned ship.

[0039] As a preference, it also includes:

[0040] A transfer learning module, connected between the deep reinforcement learning module and the reasoning and decision-making module, for migrating the parameters of the offline trained neural network model to the online reasoning and decision-making model;

[0041] The parameter adaptation submodule is used to fine-tune the migration parameters according to the difference between the actual environment and the simulation environment;

[0042] The execution feedback submodule is used to collect execution results and adjust the parameters of the reasoning decision module based on error feedback.

[0043] The unmanned ship dynamic environment path planning method based on deep reinforcement learning is applied to the system, including:

[0044] Environmental perception step, obtaining dynamic environmental data around the unmanned ship and the status information of the unmanned ship;

[0045] A data processing step, converting the dynamic environment data into a depth image and extracting features based on an attention mechanism;

[0046] Deep reinforcement learning step, training the heading decision neural network model based on the extracted features;

[0047] Model migration step, migrating the trained neural network model parameters to the online inference decision model;

[0048] The online reasoning step receives environmental observation data and performs real-time reasoning to generate the optimal bow heading;

[0049] The safety assurance step compares the decision heading with the actual heading and provides a safe heading when there is a discrepancy.

[0050] The present invention realizes adaptive path planning of unmanned ships in dynamic environments by organically combining deep reinforcement learning with attention mechanism, combined with transfer learning technology and multi-layer security guarantee mechanism, and has the following beneficial effects:

[0051] 1. The data processing module based on the dual self-attention mechanism can effectively extract key features in dynamic environments and improve the accuracy and robustness of environmental perception;

[0052] 2. The lightweight deep reinforcement learning model design combined with the optimized compound reward function significantly improves the decision-making efficiency and realizes intelligent decision-making under real-time requirements;

[0053] 3. Through transfer learning technology, the difference between the training environment and the actual application environment is solved, and a seamless transition from offline training to online application is achieved;

[0054] 4. The design of multi-layer safety assurance mechanism ensures the navigation safety of unmanned ships in various complex situations, greatly improving the reliability of the system;

[0055] 5. The modular system architecture design enables the system to have good scalability and maintainability, facilitating subsequent functional upgrades and optimizations. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a structural block diagram of the unmanned ship dynamic environment path planning system based on deep reinforcement learning of the present invention;

[0057] Figure 2 This is a structural block diagram of the environment perception module of the present invention;

[0058] Figure 3 This is a structural block diagram of the data processing module of the present invention;

[0059] Figure 4 This is a flowchart for implementing the dual self-attention mechanism in the feature extraction submodule of the present invention;

[0060] Figure 5 This is a structural block diagram of the deep reinforcement learning module of the present invention;

[0061] Figure 6 This is a structural block diagram of the reasoning decision module of the present invention;

[0062] Figure 7 This is a structural block diagram of the security module of the present invention;

[0063] Figure 8 This is a structural block diagram of the transfer learning module of the present invention;

[0064] Fig. 9 This is a flow chart of the unmanned ship dynamic environment path planning method based on deep reinforcement learning of the present invention. DETAILED DESCRIPTION

[0065] The technical solution of the present invention is described in detail below in conjunction with specific embodiments.

[0066] Example 1: Overall system architecture

[0067] like Figure 1 As shown, the unmanned ship dynamic environment path planning system based on deep reinforcement learning provided by the present invention includes an environment perception module 1, a data processing module 2, a deep reinforcement learning module 3, a reasoning and decision-making module 4, a safety assurance module 5 and a transfer learning module 6.

[0068] The environment perception module 1 is used to obtain dynamic environment data around the unmanned ship and the status information of the unmanned ship. Specifically, the environment perception module 1 receives data from a variety of sensors, including but not limited to environmental data and ship status data collected by devices such as cameras, lidars, sonars, and GPS.

[0069] The data processing module 2 is connected to the environment perception module 1, and is used to convert the dynamic environment data into a depth image and extract features based on the attention mechanism. Preferably, the data processing module 2 adopts a dual self-attention mechanism, which can adaptively focus on key areas and features in the environment and improve the effectiveness of feature extraction.

[0070] The deep reinforcement learning module 3 is connected to the data processing module 2 and is used to make heading decisions based on the extracted features. The module selects the optimal navigation strategy by evaluating the value functions of different headings. The present invention adopts a lightweight network design and an optimized training algorithm, which significantly improves the decision-making efficiency.

[0071] The reasoning decision module 4 is connected to the deep reinforcement learning module 3 to generate the best bow forward heading. This module infers the dynamic scene in a real-time environment and decides the best heading direction accordingly to provide guidance for the navigation of the unmanned ship.

[0072] The safety assurance module 5 is connected to the reasoning decision module 4 to compare the decision heading with the actual heading and provide a safe heading when there is a discrepancy. This module ensures that the unmanned ship can maintain safe navigation even in complex environments or when model reasoning deviates.

[0073] The transfer learning module 6 is connected between the deep reinforcement learning module 3 and the reasoning and decision-making module 4, and is used to migrate the parameters of the offline trained neural network model to the online reasoning and decision-making model to solve the difference between the training environment and the actual application environment.

[0074] Example 2: Specific implementation of the environment perception module

[0075] like Figure 2 As shown, the environment perception module 1 includes a multimodal sensor submodule 11 , a state information acquisition submodule 12 and a data fusion submodule 13 .

[0076] The multimodal sensor submodule 11 is used to collect environmental data including highlighting water obstacles, sudden changes in water flow, and sensor interference. This submodule integrates a variety of sensing devices, including optical cameras, depth cameras, laser radars, millimeter-wave radars, sonars, etc. Preferably, the resolution of the optical camera is 1920×1080 pixels, and the acquisition frequency is 30 frames per second; the scanning range of the laser radar is 360°, the scanning frequency is 10Hz, and the ranging accuracy is ±2cm; the detection range of the sonar is 500 meters and the resolution is 0.1 meters. Through the collaborative work of multiple sensors, the environmental conditions around the unmanned ship can be fully perceived.

[0077] The status information acquisition submodule 12 is used to obtain status information including the position, speed and attitude of the unmanned ship. This submodule mainly obtains the precise position, heading, speed, attitude and other information of the unmanned ship through GPS, inertial measurement unit (IMU), electronic compass and other equipment. Preferably, the GPS positioning accuracy is better than 0.5 meters, the angle measurement accuracy of IMU is better than 0.1°, and the measurement frequency is 100Hz. These high-precision status information provides reliable basic data for subsequent path planning.

[0078] The data fusion submodule 13 is used to align and fuse the multimodal sensor data with the state information in time and space. This submodule uses the extended Kalman filter algorithm to perform fusion processing of multi-source heterogeneous data, effectively solving the time synchronization and spatial registration problems of different sensor data. Preferably, the processing frequency of data fusion is 50Hz, which can meet the real-time requirements. The data fusion process can be expressed as:

[0079] ,

[0080] ,

[0081] in, represents the state vector at time k, represents the state transfer matrix, represents the control input matrix, represents the control vector, represents the process noise, represents the observation vector, represents the observation matrix, In this way, the environment perception module 1 can provide high-quality fusion data and provide reliable input for subsequent processing.

[0082] Example 3: Specific implementation of the data processing module

[0083] like Figure 3 As shown, the data processing module 2 includes an image conversion submodule 21 , a feature extraction submodule 22 and a time series encoding submodule 23 .

[0084] The image conversion submodule 21 is used to convert dynamic environmental data into a heading depth image. Specifically, this submodule converts the multimodal data collected by the environmental perception module 1 into a depth image in a unified format, in which the depth value represents the heading angle. Preferably, the converted depth image has a resolution of 60×40 pixels, a depth value range of 0-255, and a corresponding heading angle range of 0-360 degrees. This conversion method gives the environmental data a spatial expression, which is convenient for subsequent feature extraction.

[0085] The feature extraction submodule 22 is designed based on a dual self-attention mechanism and is used to extract depth features from the heading depth image. Figure 4 As shown in Figure 1, the dual self-attention mechanism is implemented by the following steps: first, the input feature X is globally averaged pooled to obtain the global feature G; then, the input feature is divided into the first data subset and the second data subset Next, the first weights are calculated through the neural network. , second weight and the third weight Then, based on these weights, the data subsets are weighted to obtain the intermediate results , , , and ; Finally, After the feature extractor, the output is the sixth intermediate data 0. This process can be expressed as:

[0086] ,

[0087] ,

[0088] ,

[0089] ,

[0090] ,

[0091] ,

[0092] Among them, AvgPool represents the global average pooling operation, and represents the feature subset partitioning function, , , and Represents different neural networks, and Extractor represents the feature extractor. Through this dual self-attention mechanism, the system can adaptively focus on key areas and features in the environment, greatly improving the effectiveness of feature extraction.

[0093] The time series encoding submodule 23 adopts an LSTM network structure to encode time series features. The introduction of the LSTM network enables the system to consider historical information, capture the time series characteristics of environmental changes, and provide a more comprehensive information basis for decision-making. Preferably, the hidden layer dimension of the LSTM network is 128, and the time series step size is 10, which can effectively handle continuous environmental changes. The calculation process of the LSTM unit can be expressed as:

[0094] ,

[0095] ,

[0096] ,

[0097] ,

[0098] ,

[0099] ,

[0100] in, , and They represent the forget gate, input gate and output gate respectively. Indicates the unit status, Indicates the hidden state, , , and represents the weight matrix, , , and represents the bias term, represents the sigmoid activation function, represents the hyperbolic tangent activation function.

[0101] Through this design, the data processing module 2 can effectively extract key features in the environment and consider the timing information to provide high-quality feature representation for subsequent decision-making.

[0102] Example 4: Specific implementation of deep reinforcement learning module

[0103] like Figure 5 As shown, the deep reinforcement learning module 3 includes a value assessment network 31, an action decision network 32 and a training optimization submodule 33.

[0104] The value evaluation network 31 uses the LightGBM model to evaluate the value of each possible action in the current state. LightGBM is an efficient gradient boosting decision tree framework with the characteristics of fast training speed and low memory usage, which is very suitable for scenarios that require real-time response. Preferably, the maximum depth of the decision tree is set to 8, the number of leaf nodes is set to 31, and the learning rate is set to 0.05. These parameters have been experimentally verified to be able to maintain high computational efficiency while ensuring accuracy.

[0105] The action decision network 32 is used to select the optimal heading decision based on the value evaluation result. The network adopts the Actor-Critic architecture, in which the Actor network generates the strategy and the Critic network evaluates the strategy. Preferably, the Actor network consists of 3 fully connected layers, and the number of neurons in each layer is 256, 128 and 64 respectively; the Critic network also consists of 3 fully connected layers, and the number of neurons in each layer is 256, 128 and 1 respectively. Both networks use ReLU as the activation function, except for the last layer.

[0106] The training optimization submodule 33 is used to optimize the network training through a composite reward function. The composite reward function takes into account the channel reward, the environment perception reward, the obstacle avoidance reward and the heading stability reward, which can be expressed as:

[0107] ,

[0108] in, represents the channel reward, considering whether the unmanned ship is sailing within the predetermined channel; Represents the environmental perception reward, which is adjusted according to the perception accuracy. The calculation formula is:

[0109] ,

[0110] in, and are the variances of the observed target and obstacle, respectively. represents the Euclidean distance between them; represents the obstacle avoidance reward, which is related to the collision risk; Represents a heading stability reward, used to avoid frequent direction changes.

[0111] The training optimization submodule 33 is used to optimize the network training through a composite reward function. The composite reward function takes into account the channel reward, the environment perception reward, the obstacle avoidance reward and the heading stability reward, which can be expressed as:

[0112] The value evaluation network 31 is trained using the TD algorithm, which is a temporal difference learning method that can use the estimation of the subsequent state to update the value estimation of the current state. The value update formula is:

[0113] ,

[0114] in, represents the state value at time t, Represents the state value at time t+1. Through the TD algorithm, the system can effectively learn the state value function and guide decision-making behavior.

[0115] ,

[0116] in, represents the expectation sampled from the trajectory dataset D, represents the discount factor, Indicates in status Next action The instant reward obtained, in represents the entropy regularization coefficient, Representation strategy By introducing the entropy regularization term, the exploration behavior of the strategy can be encouraged to avoid converging to the local optimal solution too early.

[0117] Through this design, the deep reinforcement learning module 3 can efficiently learn the optimal decision-making strategy and meet the real-time requirements while ensuring the quality of decision-making.

[0118] Example 5: Specific implementation of the reasoning decision module

[0119] like Figure 6 As shown, the reasoning decision module 4 includes an online reasoning submodule 41 , a dynamic scene recognition submodule 42 and a heading generation submodule 43 .

[0120] The online reasoning submodule 41 is used to receive environmental observation data and perform real-time reasoning. This submodule receives real-time environmental observation data from the environmental perception module 1, and after being processed by the data processing module 2, it is input into the reasoning model obtained by the transfer learning module 6 for reasoning. Preferably, the calculation time of the reasoning process is controlled within 5ms to ensure that the system can respond quickly to environmental changes.

[0121] The dynamic scene recognition submodule 42 is used to classify and identify potential dynamic scenes. Based on the extracted features, this submodule identifies the types of dynamic scenes in the current environment, such as static obstacles, moving obstacles, water flow changes, etc. Preferably, the scene classification uses a SoftMax classifier, and the number of categories is set to 10, including obstacle-free scenes, static obstacle scenes, single moving obstacle scenes, multiple moving obstacle scenes, water flow mutation scenes, etc. The calculation of scene classification can be expressed as:

[0122] ,

[0123] in, Indicates that the sample x belongs to the category The probability of Indicates category Through scene recognition, the system can adopt corresponding decision-making strategies for different types of scenes.

[0124] The heading generation submodule 43 is used to generate the best bow heading direction according to the inference result. Based on the result of online inference and the result of scene recognition, this submodule generates the best heading instruction by comprehensively considering factors such as safety and efficiency. Preferably, the accuracy of the heading angle is 0.1°, which is sufficient to meet the requirements of fine navigation.

[0125] It is worth noting that the calculation time of the entire reasoning and decision-making process is less than 100ms. This performance indicator ensures that the system can respond promptly in a dynamic environment and meet the requirements of real-time control. Through this design, the reasoning and decision-making module 4 can quickly generate the best heading decision based on real-time environmental observations, providing reliable guidance for the autonomous navigation of the unmanned ship.

[0126] Example 6: Specific implementation of the security module

[0127] like Figure 7 As shown, the safety assurance module 5 includes a heading comparison submodule 51 , a safety heading generation submodule 52 and a rudder angle compensation submodule 53 .

[0128] The heading comparison submodule 51 is used to compare the inference decision heading with the actual control heading of the actual ship. This submodule receives the heading instruction from the inference decision module 4 and the actual heading fed back by the unmanned ship's own control system, and performs real-time comparison. Preferably, the heading comparison frequency is 10Hz, that is, a comparison is performed every 100ms, so that the heading deviation can be detected in time.

[0129] The safe heading generation submodule 52 is used to generate a preset safe heading when the heading is inconsistent. When the heading comparison result shows that the deviation between the decision heading and the actual heading exceeds the preset threshold, the submodule will generate a safe heading instruction to overwrite the original decision. Preferably, the heading deviation threshold is set to 5°. This value is based on experimental verification and can avoid excessive intervention while ensuring sensitivity. The safe heading generation strategy takes into account factors such as the current environmental conditions, the motion status of the unmanned ship, and the navigation rules to ensure that the generated heading is safe and feasible.

[0130] The rudder angle compensation submodule 53 is used to perform steering adjustment compensation in the direction of the unmanned ship's movement. This submodule calculates the required rudder angle compensation amount according to the safe heading instruction, and gradually turns the unmanned ship to the safe heading by adjusting the rudder angle. Preferably, the rudder angle compensation adopts a proportional-integral-differential (PID) control algorithm, and the control parameters This set of parameters has been tested on a real ship and can provide smooth and effective steering performance. The PID control algorithm can be expressed as:

[0131] ,

[0132] in, Indicates the control output, i.e. the rudder angle adjustment amount; Represents the error, that is, the difference between the current heading and the target heading; , and represent the proportional, integral, and differential coefficients respectively.

[0133] Through this multi-level safety assurance mechanism, the system can ensure the navigation safety of the unmanned ship even in complex environments or when model reasoning deviates, significantly improving the reliability and robustness of the system.

[0134] Example 7: Specific implementation of the transfer learning module

[0135] like Figure 8 As shown, the transfer learning module 6 includes a parameter transfer submodule 61 , a parameter adaptation submodule 62 and an execution feedback submodule 63 .

[0136] The parameter migration submodule 61 is used to migrate the parameters of the offline trained neural network model to the online reasoning decision model. This submodule realizes the migration of the model parameters trained by the deep reinforcement learning module 3 to the model parameters used by the reasoning decision module 4. Preferably, the parameter migration adopts a combination of direct copying and fine-tuning, that is, the parameters are directly copied for the shared network structure part, and the parameters are adapted and fine-tuned for the parts with different structures.

[0137] The parameter adaptive submodule 62 is used to fine-tune the migration parameters according to the difference between the actual environment and the simulation environment. This submodule monitors the performance of the model in the actual environment, identifies the difference between the simulation environment and the actual environment, and makes targeted adjustments to the model parameters accordingly. Preferably, the adaptive adjustment adopts an online learning method, the learning rate is set to 0.01, and the parameters are updated once every 100 samples are processed. This setting ensures the adjustment effect while avoiding the instability caused by too frequent updates.

[0138] The execution feedback submodule 63 is used to collect the execution results and adjust the parameters of the reasoning decision module based on the error feedback. This submodule monitors the actual effect of the unmanned ship after executing the decision, calculates the error between the execution result and the expected result, and adjusts the model parameters based on this error. Preferably, the error feedback adopts the gradient descent method, the learning rate is set to 0.001, and the update frequency is once every 10 minutes. This setting can achieve continuous optimization of the model while ensuring stability.

[0139] Through the design of transfer learning module 6, the system solves the problem of the difference between the training environment and the actual application environment, realizes a seamless transition from offline training to online application, and greatly improves the adaptability and performance of the model in the actual environment.

[0140] Example 8: Unmanned ship dynamic environment path planning method based on deep reinforcement learning

[0141] like Fig. 9 As shown, the present invention also provides an unmanned ship dynamic environment path planning method based on deep reinforcement learning, which is applied to the above-mentioned system. The method includes an environment perception step, a data processing step, a deep reinforcement learning step, a model migration step, an online reasoning step and a safety assurance step.

[0142] The environmental perception step is to obtain the dynamic environmental data around the unmanned ship and the status information of the unmanned ship. Specifically, the environmental data and status information are obtained through multiple sensors and data fusion processing is performed.

[0143] In the data processing step, the dynamic environment data is converted into a depth image and features are extracted based on the attention mechanism. Specifically, the environment data is converted into a depth image in a unified format, and then the key features are extracted through the dual self-attention mechanism, and the time series information is processed through the LSTM network.

[0144] In the deep reinforcement learning step, the heading decision neural network model is trained based on the extracted features. Specifically, the value evaluation network and action decision network are designed, and optimized and trained through the composite reward function and TD algorithm.

[0145] The model migration step migrates the trained neural network model parameters to the online reasoning decision model. Specifically, the model is migrated from the training environment to the actual environment through parameter migration and adaptive fine-tuning.

[0146] In the online reasoning step, the environmental observation data is received and real-time reasoning is performed to generate the optimal bow heading. Specifically, online reasoning is performed based on the migrated model to identify dynamic scenes and generate the optimal heading instruction.

[0147] Safety assurance step: compare the decision heading with the actual heading and provide a safe heading when they are inconsistent. Specifically, the navigation safety of the unmanned ship is ensured through heading comparison, safe heading generation and rudder angle compensation.

[0148] Example 9: Application of the system in actual water environment

[0149] The application of the system of the present invention in the actual water environment demonstrates its superior performance and adaptability. This embodiment selects a bay area as the test scene, which has a variety of typical complex water features, including islands and reefs, waterways, water flow change areas and other ship activity areas.

[0150] During the test, the unmanned boat was equipped with the complete system of the present invention and passed through multiple preset test points in sequence. During the test, a variety of dynamic challenge scenarios were artificially set up, including the sudden appearance of floating objects, fast-moving small boats, and artificially created water disturbances.

[0151] The test results show that the system of the present invention can successfully cope with various dynamic challenge scenarios while meeting the real-time requirements (inference decision time is less than 100ms). Compared with traditional methods, this system has significant advantages in navigation safety, path planning efficiency and environmental adaptability: in terms of safety, the system successfully avoids 100% of dynamic obstacles; in terms of efficiency, compared with traditional algorithms, the path length is shortened by an average of 23.5%, and fuel consumption is reduced by 18.7%; in terms of adaptability, the system can automatically adjust the decision-making strategy to adapt to different water flow conditions and traffic density, and its adaptability is significantly enhanced.

[0152] Preferably, in actual applications, the system parameter configuration can be adjusted according to specific application scenarios and requirements. For example, in open waters, the speed can be appropriately increased and the obstacle avoidance sensitivity can be reduced, while in narrow waterways or port areas, the speed should be reduced and the obstacle avoidance sensitivity should be increased. Through this flexible parameter configuration, the system can better adapt to various complex actual application scenarios.

[0153] Example 10: System performance evaluation and optimization

[0154] The performance evaluation of the system of the present invention is mainly carried out from four aspects: environmental perception ability, decision-making efficiency, security and robustness.

[0155] In terms of environmental perception, the system achieves efficient perception of complex dynamic environments through multimodal sensor fusion and dual self-attention mechanism. Tests show that under various lighting conditions, the system's obstacle detection accuracy reaches 96.8%, which is much higher than the traditional single sensor solution (usually 82% to 85%). In particular, under harsh conditions such as low light and foggy days, the system can still maintain a high perception accuracy, thanks to the complementarity of multimodal sensor data and the effectiveness of the attention mechanism.

[0156] In terms of decision-making efficiency, the system adopts a lightweight deep reinforcement learning model design, combined with an optimized TD algorithm, which significantly improves the decision-making speed. Test data shows that the average decision time of the system is 42ms, and the maximum decision time does not exceed 85ms, which fully meets the requirements of real-time control (usually required to be less than 100ms). Compared with traditional deep reinforcement learning methods, the computing efficiency is improved by about 65%, while maintaining a considerable decision quality.

[0157] In terms of safety, the system's multi-layered safety assurance mechanism ensures the safe navigation of the unmanned ship in various situations. In the test scenario involving sudden risks, the system was able to identify risks and perform safe course adjustments within 0.5 seconds, avoiding potential collision risks. The intervention rate of the safety assurance mechanism was about 8.3%, indicating that the system can rely on normal decision-making processes to ensure safety in most cases, and the safety mechanism intervention is only required in a few high-risk situations.

[0158] In terms of robustness, the system has demonstrated good adaptability to environmental changes through transfer learning and parameter adaptation technology. In tests under different water conditions, the system performance degradation did not exceed 7%, while traditional methods usually experience a 20% to 30% performance degradation when the environment changes. This result proves the advantages of this system in dealing with environmental variability.

[0159] Based on the performance evaluation results, the system was further optimized. The main optimization measures include: reward function adjustment for specific scenarios, fine-tuning of attention mechanism parameters, and optimization of safety threshold settings. These optimization measures have enabled the system to further improve performance while maintaining its original advantages, especially stability and reliability in extreme situations.

[0160] The unmanned ship dynamic environment path planning system and method based on deep reinforcement learning provided by the present invention has broad industrial application prospects. First, the system can be applied to the fields of marine resource survey, marine environment monitoring, maritime search and rescue, etc., greatly improving the operation efficiency and safety; secondly, in terms of maritime traffic and port management, the system can be used for intelligent ship guidance and traffic coordination; in addition, the system can also be extended to other autonomous navigation scenarios, such as drones, unmanned vehicles and other fields.

[0161] The core advantage of the system of the present invention lies in its adaptability to complex dynamic environments and its safe and reliable decision-making capabilities, which enable it to play an important role in practical industrial applications. Compared with traditional methods, this system has significant advantages in environmental perception, decision-making efficiency, security and robustness, and can better meet the needs of practical applications.

[0162] In summary, the unmanned ship dynamic environment path planning system and method based on deep reinforcement learning provided by the present invention, through the combination of deep reinforcement learning and attention mechanism, as well as the design of transfer learning and multi-layer safety assurance mechanism, realizes the efficient perception, intelligent decision-making and safe navigation of unmanned ships in dynamic environments, and provides a new solution for the development and application of unmanned ship technology.

[0163] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, replacement, and improvement made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An unmanned ship dynamic environment path planning system based on deep reinforcement learning, characterized by: include: Environmental perception module, used to obtain dynamic environmental data around the unmanned ship and the status information of the unmanned ship; A data processing module, connected to the environment perception module, for converting the dynamic environment data into a depth image and extracting features based on an attention mechanism; A deep reinforcement learning module, connected to the data processing module, for making heading decisions based on the extracted features; A reasoning and decision-making module, connected to the deep reinforcement learning module, for generating an optimal bow forward heading; A safety assurance module, connected to the reasoning decision module, for comparing the decision heading with the actual heading and providing a safe heading when there is a discrepancy; The reasoning decision module includes: The online reasoning submodule is used to receive environmental observation data and perform real-time reasoning; Dynamic scene recognition submodule, used to classify and recognize potential dynamic scenes; The heading generation submodule is used to generate the optimal bow forward heading direction according to the reasoning result.

2. The system according to claim 1, characterized in that The environment perception module comprises: Multimodal sensor submodule for collecting environmental data including highlighting water obstacles, sudden changes in water flow, and sensor interference; The status information acquisition submodule is used to obtain status information including the position, speed and attitude of the unmanned ship; The data fusion submodule is used to align and fuse the environmental data with the state information in time and space.

3. The system according to claim 1, characterized in that The data processing module comprises: An image conversion submodule, used for converting the dynamic environment data into a heading depth image; A feature extraction submodule, designed based on a dual self-attention mechanism, for extracting depth features from the heading depth image; The time series encoding submodule adopts the LSTM network structure to encode the time series features.

4. The system according to claim 3, characterized in that The dual self-attention mechanism of the feature extraction submodule is implemented by the following steps: Perform global average pooling on the input features; Dividing the input features into a first data subset and a second data subset; Calculate the first weight, the second weight and the third weight respectively; Performing weighted processing on the data subset based on the weight; The processing result is output as sixth intermediate data through the feature extractor.

5. The system according to claim 1, characterized in that The deep reinforcement learning module includes: The value evaluation network uses the LightGBM model to evaluate the value of each possible action in the current state; an action decision network, for selecting an optimal heading decision based on the value assessment result; The training optimization submodule is used to optimize the training of the network through a composite reward function, wherein the composite reward function considers a channel reward, an environment perception reward, an obstacle avoidance reward, and a heading stability reward.

6. The system according to claim 5, characterized in that The value evaluation network is trained using the TD algorithm, and the value update formula is: , in, represents the state value at time t, Indicates the state value at time t+1.

7. The system according to claim 1, characterized in that The calculation time of the reasoning and decision-making process of the reasoning and decision-making module is less than 100ms.

8. The system according to claim 1, characterized in that The security module comprises: The heading comparison submodule is used to compare the inference decision heading with the actual control heading of the actual ship; A safety heading generation submodule is used to generate a preset safety heading when the headings are inconsistent; The rudder angle compensation submodule is used to perform steering adjustment compensation in the direction of movement of the unmanned ship.

9. The system according to claim 1, characterized in that Also includes: A transfer learning module, connected between the deep reinforcement learning module and the reasoning and decision-making module, for migrating the parameters of the offline trained neural network model to the online reasoning and decision-making model; The parameter adaptation submodule is used to fine-tune the migration parameters according to the difference between the actual environment and the simulation environment; The execution feedback submodule is used to collect execution results and adjust the parameters of the reasoning decision module based on error feedback.

10. A method for unmanned ship dynamic environment path planning based on deep reinforcement learning, using the system according to any one of claims 1 to 9, characterized in that: include: Environmental perception step, obtaining dynamic environmental data around the unmanned ship and the status information of the unmanned ship; A data processing step, converting the dynamic environment data into a depth image and extracting features based on an attention mechanism; Deep reinforcement learning step, training the heading decision neural network model based on the extracted features; Model migration step, migrating the trained neural network model parameters to the online inference decision model; The online reasoning step receives environmental observation data and performs real-time reasoning to generate the optimal bow heading; The safety assurance step compares the decision heading with the actual heading and provides a safe heading when there is a discrepancy.

Citation Information

Patent Citations

  • Ship berthing system and method based on unmanned aerial vehicle

    CN113534811A

  • Target submarine path prediction and tracking method for unmanned ship

    CN119399248A

  • Unmanned ship collision avoidance method and device based on DBSCAN algorithm

    CN119763371A

  • Apparatus, method, and recording medium for autonomous ship navigation

    US20210114698A1

  • Integrated Automated Driving System for Maritime Autonomous Surface Ship (MASS)

    US20210116922A1

Cited By

  • Intelligent ship navigation control and adjustment system and method based on multi-sensor data

    CN120135406A

  • Ship autonomous navigation and obstacle avoidance method

    CN120686808A

  • Automatic driving path planning method and system and medium

    CN121048652A

  • Unmanned vehicle path planning method, system and equipment based on deep reinforcement learning

    CN121297884A

  • Unmanned vehicle path planning method, system and device based on deep reinforcement learning

    CN121297884B