Unmanned ship path planning method in complex water area

The integration of deep reinforcement learning and A* algorithm with embedded sensors and multi-agent frameworks enables adaptive path planning for unmanned ships, addressing safety and efficiency challenges in complex water environments.

CN120315440APending Publication Date: 2025-07-15SUZHOU UNIV OF SCI & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510456524.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional path planning methods are difficult to meet safety, response speed and energy consumption optimization in complex waters at the same time, and lack real-time perception and adaptive adjustment of the hydrological environment state, resulting in insufficient safety and high energy consumption and time costs when performing tasks.

Method used

Combining global deep reinforcement learning and local improvement A* algorithm, we collect water area data through embedded sensors, build hydrological state vectors, adopt multi-agent reinforcement learning framework for path planning, and dynamic adjustment at the global and local levels, introduce dynamic adjustment mechanisms for risk compensation terms and reward function to achieve adaptive path optimization.

Benefits of technology

It realizes high safety, high efficiency and low energy consumption navigation of unmanned ships in complex waters, can respond to environmental changes in real time, adapt to multi-task needs, and improves the scientific nature of path planning and the accuracy of execution processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315440A_ABST
    Figure CN120315440A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned ship path planning method in a complex water area, and aims to solve the problem that an existing path planning method cannot be adaptively adjusted in a dynamic hydrological environment. A sensor (RTK-GPS, laser radar) on an unmanned ship is used for environment information collection, a hydrological state vector is constructed, and a global path is preliminarily planned based on a multi-agent reinforcement learning framework. Then, local path fine adjustment is carried out in combination with an improved A * algorithm, a risk compensation item is dynamically adjusted, and it is ensured that the unmanned ship can realize efficient and safe path planning in different hydrological environments. According to the method, the weight of the reward function is adjusted in real time, so that the unmanned ship can flexibly deal with different hydrological environments and adapt to complex water area conditions, and the accuracy and efficiency of path planning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned ship navigation and path planning, and particularly relates to a path planning method for an unmanned ship in complex waters, which is applicable to multi-task application scenarios such as environmental monitoring, pollution prevention and control, and emergency avoidance. Background Art

[0002] With the increase in the complexity and dynamic changes of the water environment, traditional path planning methods based on static rules are difficult to simultaneously meet the requirements of safety, response speed, and energy consumption optimization. Most of the existing technologies only focus on global planning or local obstacle avoidance, lacking real-time perception and adaptation to the state of the hydrological environment (such as high water period, low water period, rapid flow area, slow flow area). As a result, when an unmanned ship executes tasks in complex waters, the following problems may be faced: it is difficult to avoid high-risk waters in a timely manner, resulting in insufficient safety; lack of comprehensive consideration of dynamic factors such as flow velocity and obstacle density, and the path planning is too rough; it is difficult to adaptively adjust the planning scheme according to environmental feedback during the task execution process, resulting in an increase in energy consumption and time costs.

[0003] In summary, there is an urgent need to develop a path planning method that can combine global deep reinforcement learning and local improved A* algorithm to adaptively perceive and dynamically adjust the path for diverse hydrological environments, so as to achieve high safety, high efficiency, and low energy consumption of unmanned ships in complex waters. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, to solve the technical problems in the prior art that it is impossible to perform safe and efficient path planning for complex water environments and adapt to dynamic changes;

[0005] The purpose of the present invention is also to provide a path planning method for an unmanned ship in complex waters, including the following steps:

[0006] S1: Use embedded sensors (such as RTK-GPS, lidar, flow velocity sensors, etc.) to collect preliminary environmental information of the target water area, including water level (W level ), flow velocity (V flow ), obstacle density (D obs ), etc. data, and construct an initial hydrological state vector E t = [W level , V flow , D obs ; Complete system initialization in the cloud or local computing module, load historical hydrological environment data and basic parameters of the unmanned ship (such as hull size, speed limit, energy consumption model, etc.), and provide basic data support for subsequent global and local planning.

[0007] S2: Set up a multi-agent reinforcement learning framework in the cloud or a centralized control center according to the mission type of the unmanned ship (such as environmental monitoring, pollution prevention and control, emergency avoidance, etc.).

[0008] Define the global state vector

[0009] S t = [p ship , V ship , F static , F dynamic , W env , D task

[0010] Among them, P ship represents the ship position, V ship represents the ship speed, F static and F dynamic respectively represent the static and dynamic environmental risks, W env represents the hydrological environment weight (such as water level, flow velocity), D task represents the mission urgency.

[0011] Design the action space (such as heading angle adjustment, speed control, mission switching instruction) and the multi-objective reward function:

[0012] R = αR target + βR safety + γR time + δR energy + ∈R multitask

[0013] The initial coefficients are set as

[0014] α = 1.0, β = 1.0, γ = 0.8, δ = 1.0, ∈ = 0.5

[0015] , and are dynamically adjusted according to the hydrological environment.

[0016] Perform a preliminary global path planning through the reinforcement learning algorithm (TD3) to obtain a better path planning sequence.

[0017] S3: During the movement of the unmanned ship, collect the surrounding obstacle information in real time and use the improved A* algorithm for local path fine-tuning.

[0018] Define the cost function of the improved A* algorithm:

[0019] f(n) = g(n) + h(n) + L(n)

[0020] Among them, g(n) is the actual cost (distance, energy consumption, local resistance, etc.), and h(n) is the heuristic estimate value;

[0021] Add a risk compensation term L(n):​

[0022]

[0023] The λ value is dynamically adjusted according to the real-time hydrological state; Risk(n) represents the local risk level at node n, and Risk max is the historical maximum risk value, and ε is a constant to prevent the denominator from being zero;

[0024] Through the local search of the improved A* algorithm above and combining with the global path obtained by S2, the dynamic fine-tuning and precise obstacle avoidance of the unmanned ship's navigation route are realized.

[0025] S4: On the basis of global and local path planning, by real-time identifying the hydrological environment, the reward function and the risk compensation coefficient are dynamically adjusted:

[0026] When in the flood season, the weight β of the safety term is increased to 1.5 - 2.0, and the time term γ and the energy consumption term δ are reduced;

[0027] When in the dry season, the time term γ is increased to 1.0, and the safety term β is moderately reduced;

[0028] When entering the rapid flow area, the energy consumption term δ and the safety term β are increased to 1.5, and the multi-task term ∈ is reduced;

[0029] When in the slow flow area, the multi-task term ∈ is increased to 1.0, and other terms are kept in balance;

[0030] Meanwhile, the λ parameter in the local planning is adjusted accordingly, so that the unmanned ship is more inclined to avoid danger in a high-risk environment and pays more attention to the balance of efficiency and energy consumption in a low-risk environment.

[0031] S5: When the unmanned ship completes the local planning and executes the navigation, the system transmits the path execution result and the environmental feedback information back to the multi-agent reinforcement learning module:

[0032] Based on the trial-reward feedback mechanism, the reinforcement learning module continuously adjusts the parameters of the policy network, so that the unmanned ship can better adapt to the environment in subsequent tasks;

[0033] The soft update and experience replay mechanism are adopted to avoid parameter oscillation and ensure the stable convergence of the learning process;

[0034] Through long-term data accumulation, the LSTM model can be further used to predict the trend of the hydrological environment and provide data support for more accurate path planning.

[0035] Among them, the dynamic adjustment mechanism of the reward function in S2 includes:

[0036] According to the water level W level 、flow velocity V flow, obstacle density D obs and other information constitute the hydrological state vector

[0037] E t = level , V flow , D obs ,

[0038] Combined with the time series prediction model LSTM to identify the current hydrological state (flood season, dry season, rapid flow, slow flow), and adjust the five coefficients of α, β, γ, δ, ∈ according to the identification results, so that the reward function adapts to the environment.

[0039] Preferably, the dynamic adjustment method of the risk compensation term λ in local planning is as follows:

[0040] Flood season: λ takes 1.5 - 2.0;

[0041] Dry season: λ takes 0.5 - 1.0;

[0042] Rapid flow area: moderately increase the weight of g(n) to 1.2 - 1.5 and keep λ at a medium level;

[0043] Slow flow area: λ is about 1.0 to maintain cost balance.

[0044] Preferably, the unmanned ship can adopt a multi-agent reinforcement learning framework:

[0045] Regard different tasks (such as monitoring, salvage, emergency, etc.) as different agents and make collaborative decisions;

[0046] By sharing global environmental information and their respective local observations, improve the overall path planning efficiency.

[0047] Preferably, to improve the generalization and stability of the model:

[0048] Sample historical hydrological data in stages for training the LSTM time series prediction model;

[0049] Introduce noise interference and parameter perturbation in the reinforcement learning training to enhance the robustness of the system to environmental uncertainties. Preferably, the system can graphically present the path planning results of the unmanned ship on a visualization platform, and the platform has functions such as path planning playback, risk area marking, and path comparison display, which are convenient for operators to intuitively control the system operation status and assist in decision-making.

[0050] The unmanned ship path planning method in complex waters given in this case has the following beneficial effects:

[0051] 1. Strong dynamic adaptability

[0052] The present invention realizes real-time monitoring and response to different hydrological states (such as flood season, dry season, rapid flow, slow flow) by adopting a multi-agent reinforcement learning framework in global planning and introducing an improved A* algorithm and a dynamic adjustment mechanism of the risk compensation term λ in local planning. The system can generate a hydrological state vector based on the real-time water level, flow velocity, and obstacle density data collected by sensors, and then use a pre-trained LSTM model to accurately identify the current hydrological state, and accordingly dynamically adjust the weights of various items in the global reward function (such as safety, time, energy consumption, etc.) and the risk compensation parameters in local planning. This dynamic adaptation mechanism greatly improves the safety and flexibility of the unmanned ship operating in various waters, enabling the system to quickly make adjustments when encountering sudden environmental changes and ensuring that the navigation decision always conforms to the actual environmental conditions.

[0053] 2. Multi-level coordination

[0054] The present invention fully considers the complementary roles of global and local planning. It not only uses deep reinforcement learning to generate the overall optimal strategy at the global level but also introduces an improved A* algorithm at the local level for fine obstacle avoidance. Global planning realizes a comprehensive modeling of the complex water area environment, task urgency, and ship state through multi-agent collaborative learning, while local planning carefully solves the specific obstacle distribution and local risks. The organic combination of the two enables the unmanned ship to complete long-distance navigation based on the global optimal path and can also avoid local obstacles in a timely manner during actual navigation, ensuring the scientific nature of the overall path planning and the high precision of the execution process.

[0055] 3. Balancing energy consumption and efficiency

[0056] The present invention introduces multi-objective indicators such as energy consumption and time in the design of the reward function. By dynamically adjusting the corresponding weights, the system strives to shorten the navigation time and reduce energy consumption while ensuring safety. Specifically, when the water area environment permits, the system will appropriately increase the weight of the time term γ to promote a more compact path planning, thereby reducing the ship's operating distance and working hours; while in high-risk areas, the proportion of the safety term β is increased to ensure that the unmanned ship can avoid potential hazards and prevent unnecessary energy consumption waste. This multi-objective trade-off design enables the unmanned ship to achieve the best balance among safety, speed, and energy conservation during the task execution process.

[0057] 4. Real-time online optimization

[0058] Through the built-in closed-loop feedback mechanism and online reinforcement learning update, the present invention realizes the ability of the unmanned ship to respond to environmental changes in real time during task execution. During the operation of the unmanned ship, the path execution effect and environmental feedback data are uploaded to the cloud control system in real time, and the soft update and experience replay techniques are used to fine-tune the policy network and local planning parameters online, thereby gradually improving the performance of the model. The long-term data accumulation can also be used to further train the LSTM prediction model, making the future path planning more forward-looking and accurate. This continuous online optimization ability not only enhances the system's ability to respond to emergencies but also ensures the stability and robustness of long-term operation.

[0059] 5. Wide range of applications

[0060] The path planning method of the present invention is not only applicable to river channels but also can cover lakes, coastal areas, and other complex water environments, with good versatility. The multi-agent reinforcement learning framework and dynamic parameter adjustment mechanism enable the system to flexibly configure and optimize the path planning strategy according to different application scenarios (such as environmental monitoring, pollution prevention and control, emergency avoidance, etc.); at the same time, the LSTM model trained by phased sampling based on historical data and the improved local A* algorithm also enable the present invention to achieve efficient, stable, and low-energy navigation control in the face of complex situations such as hydrological changes, flow rate fluctuations, and uneven obstacle distributions. Therefore, this method has high engineering application promotion value and multi-task adaptation ability. Brief description of the drawings

[0061] Figure 1 It is a flow chart of a path planning method for an unmanned ship in complex waters according to the present invention;

[0062] Figure 2 It is a working principle diagram of the unmanned ship according to the present invention;

[0063] Figure 3 It is a path planning diagram of the unmanned ship during the dry season generated by the path planning method for an unmanned ship in complex waters with deep reinforcement learning and improved A* algorithm;

[0064] Figure 4 It is a path planning diagram of the unmanned ship during the flood season generated by the path planning method for an unmanned ship in complex waters with deep reinforcement learning and improved A* algorithm. Specific implementation manners

[0065] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0066] Embodiment 1:

[0067] Autonomous Path Decision-making Process of Intelligent Unmanned Vessel in Urban River Environment Monitoring Project - Adaptive Adjustment Mechanism Based on Seasonal Hydrological Conditions with Deep Reinforcement Learning

[0068] 1. Task Initialization and Equipment Startup: Construction of State Perception Vector

[0069] After receiving the task instruction, the unmanned vessel immediately starts the initialization module. A series of embedded hardware components on the hull begin to operate in coordination: First, the RTK-GPS and IMU units assist the unmanned vessel to complete precise positioning and attitude initialization in the river; meanwhile, the flow velocity sensors at the stern of the vessel start to record the current lateral and longitudinal flow velocities of the water body; the water level detector continuously measures the average water surface height between the two banks of the river and encodes this information into the hydrological state vector. Each piece of state data is uniformly encapsulated as:

[0070] E t = level , flow , obs

[0071] These data are uploaded to the cloud in real time through the wireless communication module for subsequent judgment of the current hydrological environment.

[0072] 2. Cloud Identification and Policy Loading: Classification and Judgment of Hydrological States

[0073] In the cloud, a deep reinforcement learning policy network integrated with the TD3 algorithm is running, and a specially trained LSTM model is also loaded. The LSTM model will first process the uploaded state data and compare it with the historical hydrological database to determine whether the current is the "dry season" or the "flood season".

[0074] The said model takes the state vector of consecutive time steps as input. The state vector includes three types of information: water level W level , flow velocity V flow and obstacle density D obs . And based on the initial historical state, the initial hidden state h t and memory state c t of the model are constructed to capture the trend characteristics in the time series.

[0075] The operation structure of the LSTM unit is as follows:[[]]

[0076]

[0077] Among them:

[0078] ·x t is the input vector at the current moment;

[0079] ·h t-1 ,c​t-1 is the hidden state and memory state at the previous moment;

[0080] ·W i ,W f ,W g ,W o and U i ,U f ,U g ,U o are the weight matrices of the LSTM;

[0081] ·b i ,b f ,b g ,b o are the bias terms of each gating unit;

[0082] ·σ represents the sigmoid activation function, and tanh represents the hyperbolic tangent function.

[0083] In the model training stage, optimization is completed by minimizing the cross-entropy loss function, and the loss function is defined as follows:

[0084]

[0085] where y tc is the actual hydrological state label, is the state probability distribution predicted by the model. The state distribution output by the model is:

[0086]

[0087] where, represents the hydrological state probability distribution at the current moment, W y and b y are the weight and bias terms of the output layer respectively.

[0088] According to the state probability distribution result, the system makes the following classification judgments on the current hydrological state:

[0089] · When the water level is high and the flow velocity is large, it is judged as the flood season;

[0090] · When the water level is low and the flow velocity is gentle, it is judged as the dry season;

[0091] · If the flow velocity is greater than 3 m / s, it is judged as the rapid flow state;

[0092] · If the flow velocity is less than 1 m / s, it is judged as the slow flow state.

[0093] The system identifies based on the current hydrological state. After completing the classification judgment, it feeds the category label back to the reward function dynamic adjustment module and starts the policy adaptive update process.

[0094] 3. TD3 Policy Network: Dynamic Policy Adjustment Based on Hydrological State

[0095] The system identifies based on the current hydrological state. After completing the classification judgment, it feeds the category label back to the reward function dynamic adjustment module and initiates the policy adaptive update process. The TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm, as the core policy generation module of the system, after receiving the hydrological state label output by the LSTM, performs context awareness of the current environment based on this label and dynamically adjusts the path planning policy accordingly.

[0096] S1. State Embedding and Reward Function Adaptation

[0097] The hydrological state labels output by the LSTM (such as "dry season", "flood season", "rapid flow state") are encoded into the environmental context embedding vector z env , and jointly input with the state input s of the TD3 policy network t :

[0098]

[0099] Meanwhile, the system dynamically adjusts the reward function parameters based on the hydrological state label. For example, it increases the path smoothness penalty in the "rapid flow state" and enhances the weight of the obstacle avoidance success rate in the "dry season", thus guiding the policy to better fit the current hydrological characteristics.

[0100] S2. Policy Generation and Control Output

[0101] The TD3 policy network adopts a double Q-network structure and a delayed policy update mechanism. Based on the adjusted state input TD3 calculates the current optimal action:

[0102]

[0103] This action represents the steering, speed, or path direction instruction that the unmanned boat should execute in the current state. Finally, the control instruction is sent to the hull for execution in real time via the wireless module.

[0104] 4. Adaptive Adjustment of Reward Function (Dry Season): Task Efficiency First Strategy

[0105] After receiving the judgment signal, the global path planning module resets the current reward function parameters, and the reinforcement learning system reconfigures its policy objective. The reward function adopted at this time consists of multiple objectives:

[0106] R = αR target + βR safety + γR time + δRenergy +∈R multitask

[0107] The generated path planning map of the unmanned boat during the dry season is as follows Figure 3 shown. During the dry season, the system attaches more importance to task efficiency. Therefore, the reward weight γ related to time is increased to 1.0, while the safety weight β is slightly decreased to 0.8, and the energy consumption term δ remains the original value of 1.0. At the same time, the importance of multi-task switching is maintained at a neutral level, with ∈ being 0.5. The updated reward function immediately acts on the policy network to generate a new global path proposal. This strategy shows the characteristics of efficient response in actual performance: the unmanned boat bypasses the shoals, but does not over-avoid due to small obstacles on the shore. Instead, it chooses the path with the least cost to quickly cross the medium-risk area to ensure that the task is completed according to the scheduled time node.

[0108] 5. Adaptive adjustment of the reward function (during the flood season): Switching of the safety-first strategy

[0109] During the flood season, the water flow is rapid. As soon as the unmanned boat enters the operation area, the sensor immediately detects multiple abnormalities: the floating standard deviation of the RTK positioning module increases, the water velocity reported by the flow velocity probe exceeds 1.8 m / s, and the frequency of floating obstacles detected by the lidar along the shore increases significantly. All this information is re-encoded and sent to the cloud, triggering the recognition action of the LSTM model again. The system thus determines that it has entered the flood season and then resets the parameters in the reward function. Accordingly, the new weight settings are as follows:

[0110] α = 1.0, β = 1.8, γ = 0.6, δ = 0.8, ∈ = 0.3

[0111] The generated path planning map of the unmanned boat during the flood season is as follows Figure 4 shown. This configuration is dominated by the safety item, ensuring that the unmanned boat can give priority to a safe path planning under complex flow velocities and occlusion areas. Even if it sacrifices some efficiency and energy consumption control, it can avoid accidental yaw or collision in areas with sudden changes in flow velocity.

[0112] 6. Path planning execution and local fine-tuning: Global-local linkage navigation mechanism

[0113] It is worth mentioning that during the entire navigation process, the global planning module does not operate in isolation. After each path is generated, the on-board local computing unit calls the improved A* algorithm for local path optimization. The system takes the sequence of path points output by the policy network as the global path framework and combines local environmental data (such as the detection results of the lidar) in real time for micro-path adjustment. The local path planning uses the improved A* algorithm, and its cost function is defined as:

[0114] f(n) = g(n) + h(n) + L(n)

[0115] Among them:

[0116] · g(n): Actual cost, covering navigation distance, energy consumption estimation, and resistance model at the current speed;

[0117] · h(n): Heuristic function, predicting the minimum path distance from the current node to the target point;

[0118] · L(n): Risk compensation term, used to adjust the passing cost of the path in high-risk environmental areas.

[0119] 7. Local risk compensation weight calculation

[0120] In the path cost function, the risk compensation term L(n) is calculated as follows:

[0121]

[0122] Among them, Risk(n) is the local environmental risk value calculated by the lidar and sensor module, Risk max is the normalization reference value, and ε is a constant term to prevent the denominator from being zero.

[0123] According to different hydrological states, the system sets the dynamic adjustment rule of λ as follows:

[0124] · In the dry season environment, set λ ∈ [0.8, 1.0] to avoid excessive avoidance behavior and improve navigation efficiency;

[0125] · In the flood season or rapids state, set λ ∈ [1.5, 2.0] to strengthen risk punishment and prefer stable path areas. The system recalculates L(n) every time it plans a path node to ensure that the local path has real-time response capabilities to the surrounding environment

[0126] 8. Policy feedback and online update mechanism

[0127] During the path execution process, the system continuously collects task trajectory, policy output, and environmental state feedback data, and constructs experience samples to add to the experience replay pool. The reinforcement learning module fine-tunes the policy online through the dual-network structure and delayed update mechanism of the TD3 algorithm.

[0128] In addition, the system dynamically updates the policy target network parameters according to the result of each path execution, and adopts a soft update policy to avoid policy oscillation caused by gradient mutation. The specific parameter update formula is:

[0129] θ target ← τ·θ main +(1 - τ)·θ target

[0130] Among them, τ is the soft update coefficient, and its usual value range is 0.005 ≤ τ ≤ 0.010.

[0131] Through this online update mechanism, the system can adapt to more environmental change patterns during long-term operation, achieving the collaborative optimization of high stability and task completion rate.

[0132] 9. Visual Output Module for Unmanned Ship Path Planning

[0133] Preferably, after the system completes path planning and execution, the finally generated unmanned ship path planning map can be visually presented in the task management platform. This platform has functions such as path planning playback, risk area marking, and comparison of multiple path planning results, and can display information such as path execution trajectories, key node decision-making states, and environmental risk levels in a graphical manner.

[0134] Among them, the visual interface supports the superposition and comparison of historical path planning within different task cycles, and combines with the environmental change trend to realize the tracking and analysis of the evolution process of path strategies; at the same time, the system can also display hydrological feature maps such as flow velocity, water level, and obstacle distribution through layer control methods, providing an auxiliary judgment basis for operators.

[0135] Through the design of the above visual module, the interpretability and practicality of path planning results can be significantly improved, facilitating the operator to intuitively control the system operation status and providing data support for subsequent strategy optimization and task adjustment.

Claims

1. An unmanned ship path planning method under complex water areas, characterized in that, It includes the following steps: S1: Real-time collect water area environment data including water level, flow velocity, and obstacles through embedded sensors, construct a global water area grid map, and construct a hydrological state vector based on the collected data to provide basic data for subsequent path planning; S2: Take the constructed hydrological state vector as input, use a time series prediction model to analyze the collected water area environment data, and output the probability distribution of each hydrological state, so as to realize the division of different hydrological states in the monitored water area; S3: Generate a global path planning scheme based on the deep reinforcement learning framework, and dynamically adjust according to the real-time hydrological state to ensure that the global strategy takes into account both efficiency while meeting safety constraints; S4: Use embedded sensors to detect local obstacles in real time, and adopt an improved path planning algorithm to generate local paths, and dynamically adjust according to the real-time hydrological state, so as to fully consider local risk factors during path planning; S5: The unmanned ship executes navigation according to the global path planning and local path fine-tuning results, and at the same time collects environmental data in real time. Through a closed-loop feedback mechanism, the local planning results are transmitted to the global decision-making module to update the hydrological state vector and adjust the subsequent planning strategy.

2. The unmanned ship path planning method under complex waters according to claim 1, wherein, The hydrological state vector E t is collected including real-time water level, flow velocity and obstacles, and the LSTM model, a time series prediction model, is used to predict the environmental state in combination with historical data to adjust the global and local path planning strategies in advance.

3. The method for path planning of an unmanned ship under complex water areas according to claims 1 to 2, characterized in that, Before path planning in S2, it is necessary to judge the hydrological state of the current water area through the LSTM model. The LSTM model takes the data vector of continuous time steps as input, and each data vector contains three types of information: water level, flow velocity, and obstacle density, and constructs the initial hidden state and memory state of the model based on the initial hydrological state vector of the current area to capture the trend characteristics in the time series; The model receives labeled historical data during the training phase to learn the correspondence between various hydrological states and input features, and outputs the probability distribution of the states during the prediction phase; among them, when the water level is high and the flow velocity is large, it is judged as the flood season, and when the water level is low and the flow velocity is gentle, it is judged as the dry season. If the flow velocity is greater than 3m / s, it is judged as the rapid flow state, and if the flow velocity is less than 1m / s, it is judged as the slow flow state; the judgment result of the final hydrological state is used to dynamically adjust the reward function and risk weight parameters in the path planning module.

4. The method for path planning of an unmanned ship under complex water areas according to claim 1, characterized in that, Using a multi-agent reinforcement learning framework, each agent corresponds to a different task. Define the state vector as: S t = [P ship , V ship , F static , F dynamic , W env , D task ​ Where: ·P ship Indicates the ship's position; ·V ship represents the ship speed; ·F static and F dynamic represent static and dynamic environmental risks, respectively; ·W env represents environmental weights (such as water level, flow velocity); ·D task Indicates the task urgency; Design the action space, including heading angle adjustment, speed control, and task scheduling instructions, and define the multi-objective reward function as: R = αR target + βR safety + γR time + δR energy + ∈R multitask And dynamically adjust according to the hydrological environment. To enable the unmanned ship to obtain safe and efficient decisions in different hydrological environments, the reward function parameters are adjusted in real time according to the hydrological state in the global strategy: During the flood season, due to the sharp increase in water flow, unstable water surface, and obvious changes in the riverbed, the safety term weight β will be dynamically adjusted in the path planning system, and the time term weight γ and energy consumption term weight δ will be reduced to guide the unmanned ship to preferentially choose low-risk paths and improve navigation safety; further, during the dry season, due to the decrease in water level, the increase in obstacles but the slow flow velocity, the time term weight γ will be dynamically adjusted in the path planning system, and the safety term weight β will be reduced to strengthen the path execution efficiency and improve the task scheduling response speed; Furthermore, in the rapids area, due to the relatively high local flow velocity and severe hydrodynamic interference, the weights δ of the energy consumption term and β of the safety term will be increased in the path planning system, while the weight ∈ of the multi-task term will be decreased to enhance the navigation stability and energy consumption control ability; Furthermore, in the slow flow area, due to the stable flow velocity and environment, the weight ∈ of the multi-task term will be increased in the path planning system, and the weights of the other terms will be kept balanced to achieve efficient switching between tasks and overall optimization of path planning.

5. The unmanned ship path planning method under complex waters according to claim 1, characterized in that, The improved A* algorithm adopts the following cost function: f(n) = g(n) + h(n) + L(n) where f(n) is the total cost of node n, g(n) is the actual cost from the starting point to the current node n, h(n) is the heuristic estimated cost from the current node to the target node, and L(n) is the risk compensation term, expressed as: Among them, Risk(n) is the risk value of node n, and Risk max is the maximum risk value among the current path nodes. ∈ is a small constant to prevent division-by-zero errors, and λ is a risk adjustment coefficient. The value of λ is dynamically adjusted according to the real-time hydrological state and is used to balance the requirements of efficiency and safety in path planning.

6. The method for path planning of an unmanned ship under complex water areas according to any one of claims 1 to 3, characterized in that, An LSTM time series prediction model is used to identify and classify the hydrological state in real time, and the TD3 algorithm is combined to update the reward function coefficient online to achieve smooth transition and adaptive adjustment of the reward parameters.

7. The unmanned ship path planning method under complex waters according to any one of claims 1 to 5, characterized in that, The system includes an information collection module, a global decision-making module, a local path planning module and a closed-loop feedback module. The information collection module is responsible for collecting hydrological environment data and obstacle information to provide real-time perception input; the global decision-making module conducts multi-objective strategy optimization based on the deep reinforcement learning algorithm and generates a global optimal path planning strategy according to the current environmental conditions and task objectives; the local path planning module is based on the improved A* algorithm and optimizes the path planning by dynamically adjusting the risk weights to respond to changes in the complex water area environment in real time; the closed-loop feedback module feeds back the local path planning results to the global decision-making module to realize online adaptive training of the path planning and continuously improve the response ability and accuracy of the system. In this way, the system can achieve dynamic environment adaptation, improvement of path planning accuracy and optimization of task execution efficiency, ensuring the efficient and safe operation of the unmanned ship in the complex water area environment.

8. The unmanned ship path planning method under complex waters according to claim 7, characterized in that, The closed-loop feedback mechanism includes: during the navigation of the unmanned ship, the local path planning results are compared with the sensor collected data in real time, and the learning rate and reward parameters of the deep reinforcement learning model are adjusted using the feedback information, so as to ensure safe and efficient navigation under various hydrological states.

Citation Information

Cited By

  • Robust autonomous navigation and obstacle avoidance method for underwater detection robot

    CN120521616A

  • Screw ship unloader work order processing method and system based on reinforcement learning

    CN120725633A

  • A spiral ship unloader operation work order processing method and system based on reinforcement learning

    CN120725633B

  • Ship bottom fouling organism curved surface monitoring method based on three-dimensional reconstruction and adaptive scanning

    CN120931681A

  • Complex sea condition unmanned ship path planning method and system based on cost grid map

    CN121113089A