Autonomous path planning method for dredger based on reinforcement learning
By constructing a digital twin scenario and a multimodal perception system based on reinforcement learning, the problems of low efficiency, high energy consumption, and large ecological disturbance in traditional dredging vessel operations have been solved, and efficient and safe path planning and multi-vessel collaborative operations have been achieved.
Patent Information
- Application Number
- CN202511081576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Traditional dredging operations suffer from problems such as low path planning efficiency, high energy consumption, significant ecological disturbance, insufficient obstacle detection accuracy, and conflicts in multi-objective optimization, making it difficult to operate efficiently and safely in complex underwater environments.
A reinforcement learning-based approach is adopted to construct a high-fidelity digital twin scenario using multi-dimensional real-time data. The strategy network is trained by combining fluid dynamics constraints and a three-stage course learning strategy to generate hybrid control commands. Obstacles are identified through a multi-modal perception system, enabling dynamic path planning and multi-ship collaborative operations.
It improved dredging efficiency, reduced energy consumption and ecological disturbance, significantly reduced collision risk, and enhanced the efficiency and safety of path planning.
Smart Images

Figure CN120802954A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of dredging engineering, and more particularly to a self-path planning method for a dredging ship based on reinforcement learning. BACKGROUND
[0002] Traditional dredging ship operation has significant defects, such as the path planning mode dominated by manual experience is difficult to adapt to the dynamic changes of complex underwater environment, resulting in low operation efficiency; the redundant navigation caused by the fixed operation mode makes the equipment energy consumption account for more than 40% of the total operation cost, which is not economical; the disturbance of mechanical operation to the bottom mud easily causes the diffusion of suspended solids, which seriously threatens the integrity of sensitive ecological areas such as coral reefs and seagrass beds; at the same time, the environmental perception scheme based on a single sensor has the problem of insufficient data dimension, and the obstacle mis-detection rate of a single sensor (such as a sonar) caused by environmental interference is as high as 25%, and the lack of obstacle detection capability of underwater rocks, wrecks and other obstacles further aggravates the collision risk.
[0003] In addition, traditional path planning algorithms such as A* and Dijkstra can only achieve single objective optimization and cannot form an effective balance between dredging efficiency improvement and ecological disturbance control. At the same time, the reinforcement learning path planning scheme generally does not embed the fluid mechanics equation constraint, and the migration error of the strategy network between the virtual simulation environment such as Unity-Mujoco and the real ship control system is more than 30%, which is difficult to meet the operation requirements under complex hydrodynamic conditions. The above problems together cause the traditional dredging operation to face multiple challenges of efficiency, energy consumption, ecology and safety, and an intelligent path planning method that integrates multi-source sensing data and has physical constraint perception capability is urgently needed. SUMMARY
[0004] Therefore, the present application provides a self-path planning method for a dredging ship based on reinforcement learning, which can improve the path planning efficiency in a multi-ship cooperative operation scene and reduce the training cost.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme:
[0006] A self-path planning method for a dredging ship based on reinforcement learning, characterized in that it comprises the following steps:
[0007] S1, real-time acquisition of three-dimensional terrain data, water body reflection spectrum data, dredging pump power data and three-dimensional flow velocity data, after preprocessing, a multi-dimensional state space vector is obtained;
[0008] S2, constructing a digital twin scene based on a multi-dimensional state space vector, under the digital twin scene, training a policy network by using a three-stage curriculum learning strategy with an increasing environmental complexity, introducing a Navier-Stokes residual term constraint, so that the policy network generates a hybrid control instruction containing dredging pump control, ship body turning angle and advancing speed, and dynamically adjusts the path planning priority through a hierarchical reward function; migrating the trained policy network to a ship entity;
[0009] S3, generating discrete-continuous hybrid control instructions and global path points of the ship based on the policy network, identifying obstacles based on a target multi-modal perception system, evaluating collision risks through a dynamic probability model, and when there is a collision risk, reconstructing a local path;
[0010] S4, in a multi-ship cooperative operation scene, constructing a global optimization model, assigning initial sub-regions to each ship, using distributed iterative solving, when the trigger condition is reached, using an incremental updating strategy to dynamically reassign tasks to each ship until convergence, and outputting a cooperative optimization scheme for each ship.
[0011] Further, the acquisition process of various data in S1 includes:
[0012] Deploying a multi-beam side-scan sonar to the middle area of the ship bottom to generate real-time three-dimensional terrain data of the operation area;
[0013] Carrying a hyperspectral imaging system to the bow deck to collect water reflectance spectrum data in the preset area, clustering the water reflectance spectrum data through the DBSCAN algorithm to obtain a plurality of clusters, each cluster representing a region with similar water reflectance spectrum characteristics, and determining the sediment density level corresponding to each cluster according to the pre-defined sediment density level standard;
[0014] Real-time acquisition of dredging pump power data through a Hall sensor;
[0015] Measuring three-dimensional flow velocity data in the preset water depth range through an ADCP current meter.
[0016] Further, in S1, the pre-processing method of various data includes:
[0017] After spatio-temporal registration of the three-dimensional terrain data and the water reflectance spectrum data, noise suppression is performed;
[0018] Normalizing the dredging pump power and vector decomposing the three-dimensional flow velocity data;
[0019] The processed three-dimensional terrain data, water body reflection spectrum data, dredging pump power and three-dimensional flow velocity data are subjected to feature extraction based on a space-time convolutional neural network, and a multi-head parallel attention mechanism is introduced to fuse the data in each dimension, so as to obtain a multi-dimensional state space vector containing terrain complexity, sediment density grade, energy consumption value, water flow component and environmental dynamic parameter.
[0020] Further, in S2, the construction process of the digital twin scene includes:
[0021] A fluid dynamics model is constructed based on the Navier-Stokes equation, which is used to simulate a flow velocity range of 0-10 kn and support dynamic water flow fields of laminar flow, turbulent flow and multiple modes;
[0022] A sediment settling model is established by coupling the Stokes settling formula and the two-phase flow algorithm, and the particle size of the sediment particles is defined as 0.1-5 mm and the settling rate is 0.1-0.5 m / s;
[0023] A six-degree-of-freedom ship dynamics model is constructed to match the propeller thrust-power curve with the actual ship parameters;
[0024] A three-dimensional terrain model is constructed, and the terrain roughness is subdivided into 10 levels, wherein level 0 is a smooth hard bed, and level 10 has a concave-convex amplitude of ≥2 m and a slope standard deviation of >15°;
[0025] Real-time multi-dimensional data is input into the corresponding model to construct a real-time digital twin scene.
[0026] Further, in S2, the process of training the strategy network using a three-stage curriculum learning strategy includes:
[0027] Basic terrain training: configure the digital twin scene as a flat terrain without water flow disturbance, and only enable the dredging amount reward item in the reward function, and the training target is to maintain a dredging efficiency of not less than 800 m 3 / h;
[0028] Directional water flow adaptation: introduce a 5 kn constant flow field in the digital twin scene, the flow direction is randomly initialized, and the terrain complexity is increased to level 3, and an energy consumption negative penalty item and an energy consumption benchmark value are added in the reward function, and the training target is to maintain a dredging efficiency of not less than 600 m 3 / h and an energy consumption fluctuation rate of less than 15% under the condition of directional water flow;
[0029] Complex interference reinforcement: introduce random eddy current and undercurrent in the digital twin scene, increase the ecological constraint, activate the ecological disturbance penalty item, and the training target is to maintain a dredging efficiency of not less than 400 m 3 / h under the condition of interference, and the ecological disturbance area is less than 30%.
[0030] Further, the expression of the layered reward function is:
[0031]
[0032] wherein V actual is the actual dredging volume, E current is the real-time energy consumption, A polluted is the ecological disturbance area, V target , E baseline , A total are the target dredging volume, the baseline energy consumption and the total area of the operation area respectively; A, B and C are weight coefficients in the interval of 0-1.
[0033] Further, the target multi-modal perception system in S3 includes a phased array sonar and a binocular vision module; the phased array sonar is used to model the spatial position of the hard seabed target; the binocular vision module is used to identify the flexible obstacle by using the YOLOv8s network.
[0034] Further, in S3, the expression of the dynamic probability model is:
[0035]
[0036] wherein d is the real-time detected distance between the ship and the obstacle, d safe is the safety threshold distance, and k is the sensitivity adjustment coefficient; P collision is the collision probability, and three levels of early warning are triggered according to the value of P collision .
[0037] When there is a collision risk, a dynamic window method is used to generate a candidate obstacle avoidance path with continuous curvature, and an optimal obstacle avoidance path is selected by optimizing the objective function, wherein the expression of the objective function is:
[0038] f(v,ω)=α·J dist +β·J vel +γ·J obs
[0039] wherein α, β and γ are weight coefficients; J dist is the path distance cost, J vel is the speed smoothing cost, and J obs is the obstacle avoidance cost.
[0040] Further, in S4, the objective function of the global optimization model is the minimum time taken by all ships to complete their respective tasks;
[0041] The constraint conditions include a region coverage constraint, a safety distance constraint and a load balancing constraint; wherein the region coverage constraint is that a union of all ship operation regions needs to completely cover a total operation region, and a coverage rate is 100%; the safety distance constraint is that a distance between any two different ships needs to be not less than 50 meters; and the load balancing constraint is that a difference between a maximum value and a minimum value of a single ship dredging amount is not more than 15% of an average dredging amount of all ships.
[0042] Further, S4 comprises:
[0043] dividing the operation region into n initial sub-regions, and assigning initial task boundaries to each ship;
[0044] each ship plans an optimal path in the assigned sub-region, calculates a local variable, and collects global constraint violation information, and updates the task boundary of each ship through a Lagrange multiplier until convergence; a convergence judgment condition is that a neighboring iteration task variation is less than 1%, a maximum iteration number reaches a preset value, or a global objective function variation rate is less than 0.5%;
[0045] when a path change of a certain ship caused by obstacle avoidance is more than 20% or it is detected that there is an unmarked obstacle in the task region, it is considered that a trigger condition is reached, at this time, an incremental updating strategy is used to reassign the affected sub-region, a hot start strategy is used to inherit the multiplier parameters of the last iteration, until convergence, and a collaborative optimization scheme is output.
[0046] According to the technical solution, compared with the prior art, the present application has the following beneficial effects:
[0047] 1. The present application constructs a high-fidelity digital twin scene based on multi-dimensional real-time data, and trains a strategy network using a three-stage curriculum learning strategy. During the training process, a hierarchical reward function mechanism is constructed by combining fluid mechanics constraint optimization and path energy consumption penalty mechanism, which realizes dynamic weight adjustment of dredging efficiency, energy consumption and ecological protection, reduces unit dredging energy consumption from 0.45L / ton to 0.35L / ton, with a reduction of 22%, saves 150 tons of diesel consumption per year for a single ship, and reduces CO2 emissions by 480 tons, solves the conflict of multi-objective optimization, and improves the efficiency of path planning.
[0048] 2. The present application trains the strategy network in a virtual simulation scene, and then migrates the trained strategy network to a real ship, which reduces the training cost of the real ship through virtual-real transfer learning and shortens the strategy deployment period by 60%.
[0049] 3、The application realizes the identification of hard targets and flexible targets such as fishing nets and submarine pipelines through the target multi-modal perception system, forms a multi-modal perception system with complementary rigid and flexible targets, the identification accuracy is above 95%, the collision risk can be significantly reduced, the collision accident rate is reduced from 2.1 times / month to 0.2 times / month, and the risk is reduced by 90%; combined with the millimeter-level trajectory adjustment technology, the full-scene safe operation of complex underwater environment is realized.
[0050] 4、In the multi-ship cooperative operation scene, the application distributes the operation areas of each ship based on the ADMM distributed optimization algorithm, and adopts an incremental updating strategy to dynamically re-distribute the tasks of each ship when the trigger condition is reached, thereby avoiding the high cost and low efficiency of global recalculation, focusing only on the affected local range, quickly responding to changes (such as path changes and obstacles), greatly reducing the calculation amount while ensuring system flexibility, improving overall operation efficiency, ensuring that the system can adapt to new situations in time and not waste resources in unaffected areas, and ensuring uninterrupted multi-ship cooperative operation. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.
[0052] Figure 1 The overall architecture diagram of the dredging ship autonomous path planning method based on reinforcement learning provided by the application is shown in the figure.
[0053] Figure 2 The multi-dimensional state space modeling schematic diagram provided by the application is shown in the figure.
[0054] Figure 3 The digital twin scene construction, course learning and migration process schematic diagram provided by the application is shown in the figure.
[0055] Figure 4 The flowchart of the reinforcement learning algorithm provided by the application is shown in the figure.
[0056] Figure 5 The multi-ship cooperative optimization flowchart provided by the application is shown in the figure. DETAILED DESCRIPTION
[0057] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0058] As shown in the drawings, Figure 1 The embodiment of the present application discloses a dredging ship autonomous path planning method based on reinforcement learning, comprising the following steps:
[0059] S1, real-time acquisition of three-dimensional terrain data, water body reflection spectrum data, dredging pump power data and three-dimensional flow velocity data, after preprocessing, a multi-dimensional state space vector is obtained;
[0060] S2, based on the multi-dimensional state space vector, a digital twin scene is constructed, under the digital twin scene, through incremental environmental complexity, a three-stage curriculum learning strategy is used to train a strategy network, a Navier-Stokes residual term constraint is introduced, the strategy network generates a mixed control instruction containing dredging pump control, ship turning angle and travel speed, and the path planning priority is dynamically adjusted through a hierarchical reward function; the trained strategy network is migrated to the ship entity;
[0061] S3, based on the strategy network, discrete-continuous mixed control instructions and global path points of the ship are generated, obstacles are identified based on a target multi-modal perception system, collision risk is evaluated through a dynamic probability model, and when there is collision risk, local path reconstruction is performed;
[0062] S4, in the multi-ship cooperative operation scene, a global optimization model is constructed, initial sub-regions are allocated to each ship, distributed iterative solution is used, when the trigger condition is reached, an incremental updating strategy is used to dynamically reassign tasks to each ship until convergence, and the cooperative optimization scheme of each ship is output.
[0063] The above steps will be further described below.
[0064] S1, initialization and data acquisition.
[0065] 1) Sonar system configuration: Reson SeaBat T50 multi-beam side-scan sonar is deployed in the middle area of the ship bottom, high-precision terrain scanning mode (resolution 0.5m, scanning range ±60°) is set, three-dimensional elevation grid of the operation area is generated in real time, elevation accuracy error ≤±0.1m, sampling frequency 10Hz. Sonar data is transmitted to the main control unit through Ethernet, and the sampling frequency is 10Hz.
[0066] 2) Multi-spectral perception module: Hyperspec III hyperspectral imaging system is mounted on the bow deck, covering the spectral range of 400-2500nm (spectral resolution 5nm), to collect water reflectance spectral data at a rate of 0.5 seconds per frame. After clustering by DBSCAN algorithm, several clusters are obtained, each representing an area with similar water reflectance spectral characteristics, which are closely related to sediment density. According to the pre-defined sediment density level standard (Level 1: loose, density 1.2-1.5g / cm 3 ; Level 5: dense, density >1.8g / cm 3 ), each cluster is analyzed: 2
[0067] Feature statistics: calculate the mean and variance of the spectral features of the sample points in each cluster, and analyze the correlation between these statistics and sediment density.2
[0068] Density calibration: by comparing with samples of known sediment density (for example, obtaining spectral data of different density sediments in laboratory conditions, establishing the mapping relationship between spectral features and sediment density), the corresponding sediment density range is calibrated for each cluster. For example, if the spectral features of a cluster are most similar to the spectral features of samples with a sediment density of 1.2-1.5g / cm 3 in the laboratory, it is classified as Level 1 (loose); if another cluster's spectral features match those of samples with a sediment density of >1.8g / cm 3 , it is classified as Level 5 (dense).2
[0069] Dynamic adjustment: Since the water environment is dynamic, the sediment density will also change with factors such as water flow and sediment input. Therefore, the DBSCAN algorithm can process newly collected spectral data in real time, dynamically adjust the clusters based on the latest data distribution, and thus dynamically classify the sediment density levels.
[0070] Through this closed-loop process of "data collection-feature statistics-density calibration-dynamic adjustment", the DBSCAN algorithm realizes the dynamic and adaptive classification of sediment density levels, and the classification accuracy is stable at more than 93% through actual measurement, providing a reliable quantitative analysis tool for water sediment movement monitoring and ecological environment assessment.
[0071] 3) Energy consumption monitoring unit: high-precision Hall sensor (accuracy ±0.5%) is used to collect real-time dredging pump power (range 50-500kW), data is transmitted through CAN bus, sampling frequency 20Hz.
[0072] 4) ADCP current meter: integrated Nortek Signature 1000 ADCP, measuring three-dimensional current velocity (u, v, w components, accuracy ±0.05 m / s) at 0-50 m water depth, data output frequency 2 Hz.
[0073] 5) Data preprocessing:
[0074] ① Temporal-spatial registration and noise suppression: The temporal-spatial registration relies on the SLAM (simultaneous localization and mapping) algorithm to build a unified coordinate system for multiple sensors: First, accurately extract seabed feature points (such as reef contour feature points, terrain fluctuation inflection points) from the sonar point cloud data, and simultaneously collect the water spectral feature region in the multispectral image. Through the Iterative Closest Point (ICP) algorithm, the optimal rigid transformation matrix between the two modal data is solved, realizing the sub-pixel level accurate alignment of the sonar three-dimensional point cloud coordinate system and the multispectral image pixel coordinate system. For the roll (±15°) and pitch (±10°) attitude deviation caused by ship motion, the SLAM algorithm solves the six-degree-of-freedom attitude parameters of the ship body in real time through the Extended Kalman Filter (EKF), and dynamically corrects the coordinate system conversion relationship based on the Euler angle compensation model, ensuring that the spatial consistency error of multi-source data in the UTM projection coordinate system is controlled within ≤5 cm.
[0075] After this processing, the spatial misalignment of cross-modal data caused by ship body sway is effectively eliminated, providing a unified spatial reference for high-precision fusion of terrain grid and sediment density distribution, and making the subsequent feature fusion accuracy improve by more than 40%.
[0076] In the noise suppression stage, the Kalman filter algorithm is used to recursively denoise the time series dynamic characteristics of the sonar data. Specifically, the original sonar elevation data can be processed by median filtering with a window size of 5x5 and a step size of 1 m to effectively eliminate isolated noise points with an error exceeding 0.3 m, while preserving key terrain detail features such as gullies and steep slopes.
[0077] For the frequency domain characteristics of spectral data, the Savitzky-Golay algorithm with a window length of 11 and a polynomial order of 3 is used for smoothing. Through the differential filtering strategy, the signal-to-noise ratio of both types of data is simultaneously improved to more than 20 dB.
[0078] ② Energy consumption data standardization: Based on the energy conservation model, the original power data of the sensor is normalized:
[0079] (unit: kW·h)
[0080] where P actual is the real-time power (kW), Δt = 10 s is the sampling interval, and P actual = 500 kW is the rated power.
[0081] ③ Water flow vector analysis: The three-dimensional flow velocity vector measured by ADCP is decomposed into:
[0082] Flow direction: 0°-360° (0° is the positive north direction, increasing clockwise), resolution 1°;
[0083] Flow velocity value: 0-10kn (1kn≈0.514m / s), quantization accuracy 0.1kn;
[0084] Vertical component: w direction velocity range ±2m / s, used to evaluate the water flow stratification effect.
[0085] 6) Multi-dimensional feature extraction and fusion:
[0086] The pre-processed multi-dimensional data is input into the spatio-temporal convolutional neural network, which includes: the input layer receives a 4-channel tensor composed of terrain elevation, sediment density, energy consumption and water flow; the spatio-temporal convolutional layer extracts the spatial-temporal correlation features in the data through a 3D convolution kernel of 3x3x3; the feature fusion layer uses a multi-head parallel attention mechanism to weight and fuse the four-dimensional data of terrain, sediment, energy consumption and water flow, and finally outputs a multi-dimensional feature vector, providing a high-robustness feature representation for subsequent analysis. The specific implementation process of the multi-head parallel attention mechanism in the four-dimensional data fusion is as follows:
[0087] Grouping and mapping of input features: The feature tensor output by the spatio-temporal convolutional layer covers four dimensions of terrain (including elevation, slope, curvature), sediment density, energy consumption (represented by pump power), and water flow (containing three components of flow velocity), with dimensions CxHxWxT (corresponding to channel number, height, width, and time step, respectively). In the grouping strategy, the four-dimensional data is divided into N parallel heads (N=8 here), and each parallel head independently processes the features of all four dimensions. Specifically, different linear transformations (weight matrices Map the input features X to different subspaces to generate the corresponding Query vector Key vector Value vector Where i represents the i-th parallel head.
[0088] Attention score calculation: In each parallel head, the correlation weight between the four-dimensional features is calculated through a specific formula, which is:
[0089]
[0090] Where, d k represents the dimension of the key vector, which is used for scaling operation to ensure the stability of the gradient, avoid the problem of gradient explosion or disappearance in the calculation process, and make the model training more stable and efficient.
[0091] Feature focus logic: Different parallel heads have different focuses on the correlation of various dimensional features. For example, in the case of elevation mutation in terrain features, some parallel heads may have strong correlation with the direction of flow velocity in water flow features; while other parallel heads may focus on the synergistic relationship between sediment density and energy consumption. Through this differentiated attention mechanism, the model can more comprehensively capture the complex relationships between multi-dimensional features, achieve deep integration and understanding of terrain, sediment density, energy consumption, water flow and other features, and improve the accuracy of representing the overall environmental state.
[0092] Weighted aggregation of value vectors: By applying the calculated attention weights to the value vectors V i , the weighted aggregation of value vectors is realized, thereby generating the fused features of each parallel head, with the specific formula being
[0093] Head i = Attention i ·V i
[0094] This process enables feature fusion to fully consider the correlation weights between different dimensional features, highlighting key information and enhancing the effectiveness of feature expression.
[0095] Multi-dimensional fusion example: In multi-dimensional fusion, different parallel heads exhibit specific synergistic optimization effects. For example, in terms of terrain and water flow synergy, a certain parallel head establishes a close correlation between "steep terrain" and "vertical flow velocity" with high attention weights, thereby optimizing the energy consumption in path planning for climbing, achieving more efficient operation; in terms of sediment and energy consumption optimization, another parallel head emphasizes the corresponding relationship between "dense sediment area" and "high pump power" to improve the efficiency of dredging operations, ensuring the rational use of resources and efficient task execution. Through this multi-dimensional weighted fusion mechanism, the model can integrate the advantages of various dimensional features to achieve more accurate and efficient decision support.
[0096] In the multi-head concatenation step, the outputs of all parallel heads are concatenated along the channel dimension, specifically represented as (MultiHead=Concat(Head1,Head2.....dots,Head N )). Then a linear projection operation is performed to reduce the dimension of the concatenated features to 128 dimensions through a fully connected layer, with the mathematical expression being (Z=
[0097] MultiHead·W O ), where (W O ) is the projection matrix, and the final output is The process reduces the feature dimension while retaining key information, both improving the model operation efficiency and ensuring the compactness and effectiveness of the feature expression.
[0098] 7) State vector construction and update: generate a multi-dimensional state vector containing terrain complexity (0-10 levels), sediment density (1-5 levels), energy consumption value, water flow u / v / w component, and environmental dynamic parameters, as shown in Figure 2 The environmental state is updated every 10 seconds.
[0099] The terrain complexity is a 4-dimensional vector containing terrain complexity score, elevation variance (terrain fluctuation), slope mean (overall inclination trend), and curvature extreme (local concave-convex feature). The sediment density is a 3-dimensional vector containing density level (1-5 levels), density gradient (adjacent area difference), and density stability (time series fluctuation). The internal loss value is a 3-dimensional vector containing normalized instantaneous power value, average energy consumption and standard deviation in the past 10 seconds. The water flow component contains a 6-dimensional vector containing real-time values of u, v, and w components, flow rate change rate (acceleration), and flow direction angle stability (variance). The environmental dynamic parameters contain a 2-dimensional vector including ecological disturbance index and obstacle proximity.
[0100] Finally, in the data packaging link, the pre-processed terrain grid, sediment density label, and flow velocity vector are integrated into a state vector containing multi-dimensional features through feature engineering technology, and the nanosecond-level transmission capability of gigabit Ethernet is used to deliver the state vector to the reinforcement learning decision module in real time.
[0101] S2, reinforcement learning training and online decision, the specific process is shown in Figures 3-4
[0102] 1) Digital twin scene construction:
[0103] Digital twin scene construction is based on multi-source real-time data to construct a high-fidelity simulation environment, providing physical constraints and dynamic inputs for reinforcement learning.
[0104] The present application constructs a high-fidelity physical constraint space through real-time data collection, ensuring the authenticity of the simulation environment. In the course of phased training, environmental parameters are actively configured to optimize the policy network in a controllable complexity progression. The two are decoupled through phase separation, ensuring the rigor of the training process and laying the foundation for dynamic adaptation after real ship deployment.
[0105] Specifically, the application ensures that the complexity of the training environment progresses in stages and is decoupled from the dynamic changes of real-time data through the architecture of "training phase preset environment parameters + real-time data calibration physical model". The core value of real-time data is to define the physically feasible region and initial state of the simulation environment, and it does not directly drive the training process, thereby ensuring the rigor of the training while improving the generalization ability of the strategy network to real scenarios.
[0106] The application ensures the consistency of the control instructions generated by the simulation environment and the control of the real ship dredging pump. The core logic is: calibrate the physical rules of the simulation model with real-time data, use transfer learning to bridge the simulation-reality gap, and dynamically correct the execution error through closed-loop feedback to finally realize the precise mapping of "virtual instructions-real ship execution".
[0107] When building a high-fidelity digital twin scene in the Unity-Mujoco joint simulation platform, a series of key technologies are adopted: through fluid dynamics modeling based on the Navier-Stokes equation, a dynamic water flow field that can simulate flow rates ranging from 0 to 10 kn (Reynolds number Re = 1 x 10 4 ~ 1 x 10 6 ) and supports laminar and turbulent multi-modal flow is constructed; with the aid of Stokes sedimentation formula and two-phase flow coupling algorithm, a sedimentation model is established to accurately define the sediment particle size as 0.1-5mm and the sedimentation rate as 0.1-0.5m / s; a six-degree-of-freedom rigid body motion model is integrated to carry out ship dynamics simulation, making the propeller thrust-power curve match the real ship parameters (maximum thrust 50kN, power response delay ≤0.2s); the terrain configurable design is realized, which subdivides the terrain roughness into 10 levels (level 0 is a smooth hard bottom, level 10 has a concave-convex amplitude ≥2m and a slope standard deviation >15°), thereby meeting the diversified simulation application needs.
[0108] The input parameters of each model come from real-time sensor data streams: the ADCP flow rate is injected into the fluid model after synchronization by PTP protocol (timing error <10ms), the hyperspectral sediment density is preprocessed by Savitzky-Golay filtering (window 11, order 3), and the sonar terrain data is denoised by Kalman filtering (Q=0.01, R=0.1). This model-driven architecture based on measured data forms a closed loop of "real-time sensor acquisition-data preprocessing-multi-physical field coupling-digital twin evolution", making the simulation scene respond to real environment changes with <200ms full-link delay, providing multi-dimensional dynamic decision-making basis for ship path planning including flow field, sedimentation, and terrain.
[0109] 2) Three-stage course learning strategy: The goal of the course learning strategy is to train the strategy network to adapt to multi-scenario operation requirements by increasing the complexity of the environment in stages. This module trains the strategy network in three stages to gradually increase the complexity of the environment:
[0110] ①Basic terrain training:
[0111] The scene is configured as a flat terrain (slope less than 5°) without water flow interference; the reward function only enables the dredging amount reward term (α = 1.0), aiming to quickly establish basic path planning capabilities; the training goal is to maintain a dredging efficiency of not less than 800 m 3 / h, with a convergence step number of not more than 500,000.
[0112] ②Directional water flow adaptation:
[0113] The scene is upgraded to introduce a 5kn constant flow field (flow direction randomly initialized), while the terrain complexity is increased to level 3; a negative energy consumption penalty term (β = 0.3) and an energy consumption baseline value are added to the reward function; the training goal is to maintain a dredging efficiency of not less than 600 m 3 / h and an energy consumption fluctuation rate of less than 15% under directional water flow conditions.
[0114] ③Complex interference reinforcement:
[0115] The scene is reinforced with superimposed random vortexes (diameter 10 to 50 m, rotation speed 0.5 to 2 m / s) and dark currents (flow velocity gradient 2 to 5 kn / m); the ecological constraint is increased, and the ecological disturbance penalty term (γ = 0.1) is activated, with a pollution area threshold A total = 10 4 m 2 ; during robustness verification, the strategy is required to maintain a dredging efficiency of not less than 400 m 3 / h and an ecological disturbance area of less than 30% under interference conditions.
[0116] 3) Strategy network inference phase: first, the Actor network generates discrete-continuous hybrid control instructions, including dredging pump start-stop state, precision ±0.5° ship body turning angle, and resolution 0.1 kn travel speed, and outputs global path points at a frequency of every 5 seconds; on this basis, the Critic network simultaneously corrects the fluid dynamics feasibility when evaluating the state value by introducing the Navier-Stokes residual term (the Navier-Stokes residual term reflects the deviation of the current fluid motion state from the basic mechanics of viscous fluid motion), thereby ensuring the physical consistency of the heading angle and the water flow vector, and providing state inputs that meet the real physical characteristics for reward function calculation.
[0117] 4) Hierarchical reward calculation link, real-time collection of multi-dimensional data such as dredging amount V actual , energy consumption E current , and ecological disturbance area a polluted , and then substituting them into the reward function, the expression is:
[0118]
[0119] wherein V actual is the actual dredging volume, E current is the real-time energy consumption, A polluted is the ecological disturbance area, V target , E baseline , A total are the target dredging volume, the baseline energy consumption and the total area of the operation region respectively; A, B, C are weight coefficients in the interval of 0-1. A represents the weighted degree of the ratio of the actual dredging volume V actual to the target dredging volume V target , reflecting the importance of dredging efficiency. The greater the value of A, the more significant the positive influence of dredging efficiency on the total reward; B is used to measure the punishment degree of the ratio of the current energy consumption E current to the baseline energy consumption E baseline . The greater the value of B, the more intense the negative influence of energy consumption higher than the baseline value on the total reward; C reflects the punishment weight of the ratio of the pollution area A polluted to the total area A total . The greater the value of C, the more prominent the negative influence of ecological disturbance on the total reward.
[0120] The formula quantifies and weights the dredging efficiency, energy consumption cost and ecological impact, and finally realizes the dynamic self-adaptive adjustment of the path planning priority, forming a closed-loop progressive decision-making process of "state perception-physical constraint evaluation-multi-objective reward calculation-control strategy optimization".
[0121] 5) Domain randomization migration: the goal of domain randomization migration is to migrate the policy network trained in simulation to the real ship, solving the "simulation-reality" domain difference problem. Random perturbation is performed on the water transparency (0.5-5m attenuation coefficient) and sensor noise (Gaussian white noise σ=0.05), and the simulation strategy parameters are migrated to the real ship through transfer learning, with an adaptation success rate of ≥85%.
[0122] The present application adopts a two-stage architecture of "pre-trained model migration + online dynamic calibration" to realize efficient reuse and fine adjustment of strategy parameters:
[0123] ① Initial parameter migration (cold start phase):
[0124] Model structure alignment: keep the policy network architecture of the simulation environment and the real ship controller consistent (both are 3-layer fully connected networks with 128 neurons per layer), ensuring the parameter dimension matching.
[0125] Key parameter mapping: The physical parameters learned in simulation, such as fluid resistance coefficient and sediment grabbing efficiency, are directly migrated to the Mujoco engine parameter table of the real ship dynamics model. The weights of the hierarchical reward function (a / b / g) are normalized, and the initial values are dynamically adjusted based on the ecological sensitivity level of the real ship operation area.
[0126] Zero-shot testing: When the real ship is first deployed, the automatic control mode is turned off, and the historical sensor data of the real ship (including 1000 complex water condition samples) are played back offline to verify whether the control instructions (such as steering angle and pump power) output by the policy network meet the principles of fluid mechanics (error ≤ 15%).
[0127] ② Online calibration optimization (adaptive phase):
[0128] Physical constraint closed loop: The IMU on the real ship measures the roll / pitch angles of the ship body in real time, and compares them with the predicted values of the simulation model. If the deviation is > 5°, the boundary conditions of the Navier-Stokes equation are modified (such as increasing the ship body resistance coefficient);
[0129] Reward function closed loop: The actual dredging volume, energy consumption, and ecological disturbance area are calculated every hour, and the deviation from the target value is calculated to automatically adjust the hierarchical reward weights a / b / g, forming a "measurement-evaluation-reoptimization" cycle.
[0130] ③ Real ship adaptation:
[0131] Space-time calibration technology: PTP (Precision Time Protocol) is used to achieve nanosecond-level synchronization between the simulation clock and the real ship sensor clock, ensuring that the 32-dimensional state vector (including terrain, flow rate, and energy consumption data) input to the policy network has a timing error of < 10ms; SLAM algorithm is used to establish the mapping relationship between the real ship coordinate system and the simulation global coordinate system, and the spatial alignment error between the measured sonar point cloud and the simulation terrain grid is ≤ 10cm.
[0132] Hardware-in-the-loop testing: Before the real ship is deployed, the simulation scene is connected to the real ship controller through the ROS2 middleware for 12 hours of continuous hardware-in-the-loop testing (HIL Test) to verify the reliability of the bidirectional transmission of sensor data flow and control instruction flow (packet loss rate < 0.1%).
[0133] Fault tolerance mechanism design: If the migrated policy has an exception on the real ship (such as failing to avoid obstacles for 3 consecutive times), it automatically switches to "simulation prediction mode": the Unity-Mujoco model is used to predict the ship's state in real time for the next 5 seconds, and generate emergency control instructions, with a maximum fault tolerance time of 30 seconds.
[0134] S3, dynamic obstacle avoidance execution:
[0135] 1) The policy network is used to generate discrete-continuous hybrid control instructions for the ship, including dredge pump control, hull turning angle and travel speed, and global path points.
[0136] 2) Obstacle detection and classification: In the environmental perception layer, the phased array sonar array scans the 200m area in front to construct a 3D point cloud model of rigid obstacles such as reefs and sunken ships with a resolution of 0.5m, achieving accurate modeling of the spatial position of hard targets on the seabed; at the same time, the binocular vision module collects image data with a resolution of 1920x1080, and through the integration of the YOLOv8s model, it identifies flexible targets such as fishing nets and seabed pipelines, with an average precision (mAP@0.5) of more than 95% at an IOU threshold of 0.5, thus forming a multi-modal perception system that complements rigid and flexible targets.
[0137] 3) Collision risk assessment: A dynamic probability model is constructed based on the logistic function, with the expression:
[0138]
[0139] where the safety distance threshold d safe = 50m, the sensitivity adjustment coefficient k = 0.5, and the parameter d represents the distance between the detected ship and the obstacle (in meters). According to the calculation results, a three-level warning mechanism is triggered (yellow warning: 0.3≤P collision <0.6; orange warning: 0.6≤P collision <0.8; red warning: P collision ≥0.8).
[0140] 4) Path re-planning stage:
[0141] When there is a collision risk, the dynamic window method is used to generate a candidate obstacle avoidance path with continuous curvature, and the optimal obstacle avoidance path is selected through optimization of the objective function, where the expression of the objective function is:
[0142] f(v,ω)=α·J dist +β·J vel +γ·J obs
[0143] where α, β, γ are weight coefficients; J dist is the path distance cost, J vel is the speed smoothing cost, and J obs is the obstacle avoidance cost. α is the path distance cost weight, reflecting the importance of the current path distance to the target path (or endpoint), β is the speed smoothing cost weight, reflecting the smoothness of the control speed change, and γ is the obstacle avoidance cost weight, indicating the priority of the obstacle avoidance behavior.
[0144] The local path reconstruction is completed within 200ms, ensuring that the ship performs obstacle avoidance actions with ≤0.1kn speed fluctuations and ±1° steering error, achieving millimeter-level precision trajectory adjustment and smooth connection of the global path.
[0145] The process of the target function optimization method is as follows:
[0146] Candidate solution generation: First, a feasible speed and angular velocity combination window is generated according to the physical limitations of current speed and angular velocity (including maximum acceleration, braking distance, etc.); then, in the cost calculation stage, for each candidate speed combination, the corresponding path distance cost (such as the Euclidean distance from the current path to the target point or the path tracking deviation), the speed smoothing cost (i.e. the square of the speed change rate such as acceleration and angular acceleration, used to punish non-smooth motion) and the obstacle avoidance cost (usually inversely proportional to the distance to the nearest obstacle, such as the inverse of the square of the distance) are calculated.
[0147] Next, each cost term is substituted into the target function for weighted summation, and the total cost of each candidate solution is calculated, and the speed and angular velocity combination with the smallest total cost is selected as the final control instruction.
[0148] In addition, parameter tuning strategies as optimization means include empirical method (manually adjusting the weight of each cost term according to scene requirements, such as increasing the weight of obstacle avoidance cost in obstacle avoidance scenarios), automatic optimization (using genetic algorithms, particle swarm optimization and other algorithms to search for the optimal weight combination to maximize dredging efficiency or minimize energy consumption) and adaptive adjustment (adjusting the weight in real time according to environmental changes such as increased water flow and increased obstacle density, for example, increasing the weight of obstacle avoidance cost as the collision risk increases), so as to ensure that the entire control process maintains optimal performance in different scenarios.
[0149] S4, multi-ship cooperative operation optimization, the specific process is as shown in Figure 5 .
[0150] 1) Global optimization model modeling:
[0151] Objective function: the time taken by all ships to complete their respective tasks is minimized, expressed as:
[0152] (n is the number of ships, T i )
[0153] where T i represents the time taken by the i-th ship to complete a specific task (such as dredging, sailing, etc.).
[0154] Constraint conditions:
[0155] ① Area coverage constraint: (coverage rate 100%)
[0156] The union of all ship operation areas needs to completely cover the total operation area, ensuring no omission and a coverage rate of 100%, wherein represents the set of the operation areas of the n ships respectively; TotalArea represents the total operation area; and n represents the number of ships participating in the operation. i
[0157] ② Safety distance constraint:
[0158] For any two different ships i and j, the distance between them needs to be no less than 50 meters to ensure the safety of the operation and avoid collision risks.
[0159] ③ Load balancing constraint: [max i V i -min i V i ≤ 15% x Mean(V) (V i is the single-ship dredging amount)
[0160] represents the difference between the maximum and minimum single-ship dredging amounts, which needs to be no more than 15% of the average dredging amount of all ships, to ensure that the dredging tasks of each ship are relatively balanced and to avoid overwork or idleness of some ships. Wherein V i represents the dredging amount of the i-th ship.
[0161] 2) Distributed iterative solution:
[0162] Initialization: The central controller divides the operation area into n initial sub-areas (based on the Voronoi diagram), and each ship receives the initial task boundary (the initial task boundary refers to the geometric boundary of the initial operation sub-area allocated to each dredging ship through the Voronoi diagram division algorithm).
[0163] Iterative optimization: Each ship locally solves the sub-problem, plans the optimal path within the allocated area, and calculates the local variables (operation time, dredging amount). The central controller collects global constraint violation information (such as area overlap, insufficient safety distance), updates the task boundary of each ship through the Lagrange multiplier (in the distributed optimization of multi-ship cooperative operation, the Lagrange multiplier method converts global constraints into penalty terms of local optimization problems through a decomposition and coordination mechanism, gradually adjusts the task boundary to meet safety and efficiency requirements), until convergence (the maximum number of iterations is 50 times, or the global objective function changes by less than 0.5%) Convergence criterion: terminate when the maximum number of iterations is 50 times, or the global objective function changes by less than 0.5%.
[0164] 3) Dynamic task reallocation:
[0165] Trigger condition: When a ship changes its path due to obstacle avoidance by more than 20% (deviation distance from the original planned path > 100m), or detects the presence of unmarked obstacles (such as temporary fishing nets) in the task area, an incremental update strategy is used to reassign the affected sub-area, and a hot start strategy is used to inherit the multiplier parameters from the previous iteration until convergence, outputting the collaborative optimization scheme.
[0166] Incremental update is an efficient local optimization strategy, which means that the central controller only reassigns and adjusts the affected sub-area (such as within a radius of 500m) when facing task changes, rather than recalculating all regions globally. This approach avoids the high cost and low efficiency of global recalculation, focuses only on the affected local range, quickly responds to changes (such as path changes and obstacle appearance), ensures system flexibility while significantly reducing computational load, improves overall efficiency, and ensures that the system can adapt to new situations without wasting resources in unaffected areas.
[0167] The hot start strategy inherits the multiplier parameters from the previous iteration, reducing the convergence time by 70% compared to full update. For a fleet of 10 ships, the reallocation time is ≤1.5 seconds, ensuring uninterrupted multi-ship collaborative operation. The hot start strategy is an efficient method for optimizing the iteration process, which involves inheriting key parameters (such as multiplier parameters) from the previous iteration, avoiding the tedious process of calculating all parameters from scratch during full update. In the context of multi-ship collaborative operation, this strategy directly uses the multiplier parameters calculated in the previous iteration, eliminating the need to rederive the initial values of these parameters or perform full recalculation.
[0168] The embodiments in the specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0169] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A dredging vessel autonomous path planning method based on reinforcement learning, characterized in that: The following steps are involved: S1. Real-time acquisition of three-dimensional terrain data, water body reflectance spectrum data, dredging pump power data, and three-dimensional flow velocity data, and after pre-processing, obtain a multi-dimensional state space vector; S2. Build a digital twin scenario based on a multi-dimensional state space vector. In this digital twin scenario, use a three-stage curriculum learning strategy to train the policy network by increasing the complexity of the environment, introduce Navier-Stokes residual term constraints, and dynamically adjust the path planning priority through a hierarchical reward function; then migrate the trained policy network to the ship entity. S3: Generates discrete-continuous hybrid control instructions and global path points for the ship based on the policy network, identifies obstacles based on the target multimodal perception system, assesses collision risk through a dynamic probability model, and performs local path reconstruction when collision risk exists. S4. In the scenario of multi-ship collaborative operation, a global optimization model is constructed, an initial sub-area is allocated to each ship, and a distributed iterative solution is adopted. When the trigger condition is met, an incremental update strategy is used to dynamically redistribute tasks for each ship until convergence, and the collaborative optimization plan for each ship is output.
2. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: The process of obtaining various data in S1 includes: Deploy a multi-beam side-scan sonar to the mid-bottom area of the ship to generate three-dimensional terrain data of the operating area in real time; A hyperspectral imaging system was mounted on the bow deck to collect water reflectance spectral data from a preset area. The water reflectance spectral data were clustered using the DBSCAN algorithm to obtain several clusters. Each cluster represented an area with similar water reflectance spectral characteristics. The sediment density level corresponding to each cluster was determined based on a pre-defined sediment density level standard. The power data of the dredging pump is collected in real time through the Hall sensor; The ADCP current meter measures three-dimensional flow velocity data within a preset water depth range.
3. The autonomous path planning method for a dredging vessel based on reinforcement learning according to claim 1, characterized in that: In S1, various data preprocessing methods include: After the three-dimensional terrain data and water body reflectance spectrum data are temporally and spatially registered, noise suppression is performed; Normalize the dredging pump power and perform vector decomposition on the three-dimensional flow velocity data; Based on the spatiotemporal convolutional neural network, feature extraction is performed on the processed three-dimensional terrain data, water body reflectance spectrum data, dredging pump power and three-dimensional flow velocity data. A multi-head parallel attention mechanism is introduced to fuse the data of each dimension to obtain a multidimensional state space vector containing terrain complexity, sediment density level, energy consumption value, water flow component and environmental dynamic parameters.
4. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: In S2, the construction process of the digital twin scene includes: A fluid dynamics model is constructed based on the Navier-Stokes equations to simulate the flow velocity range of 0-10kn and support dynamic water flow fields of laminar and turbulent multi-modal flows; The sedimentation model was established by combining the Stokes sedimentation formula with the two-phase flow coupling algorithm, with the sediment particle size defined as 0.1-5 mm and the sedimentation rate as 0.1-0.5 m / s. Construct a six-degree-of-freedom ship dynamics model to match the thrust-power curve of the propeller to the actual ship parameters; A three-dimensional terrain model was constructed, and the terrain roughness was subdivided into 10 levels, where level 0 was a smooth and hard bottom bed, and level 10 was a concave-convex amplitude ≥ 2m and a slope standard deviation > 15°; Input real-time multi-dimensional data into the corresponding model to build a real-time digital twin scenario.
5. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: In S2, the process of training the policy network using the three-stage curriculum learning strategy includes: Basic terrain training: The digital twin scene is configured as a flat terrain with no water flow interference. The reward function only enables the silt removal reward item. The training target silt removal efficiency is not less than 800m 3 / h; Directional water flow adaptation: A 5 kn constant velocity field is introduced into the digital twin scene, the flow direction is randomly initialized, and the terrain complexity is increased to level 3. A negative energy penalty term and an energy consumption baseline value are added to the reward function. The training goal is to maintain a dredging efficiency of no less than 600 m under directional water flow conditions. 3 / h and the energy consumption fluctuation rate is less than 15%; Complex interference enhancement: random eddies and undercurrents are introduced into the digital twin scenario, ecological constraints are added, and ecological disturbance penalty items are activated. The training goal is to maintain a dredging efficiency of no less than 400m under interference conditions. 3 / h, and the ecological disturbance area is less than 30%.
6. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 5, characterized in that: The expression of the layered reward function is: Among them, V actual is the actual desilting volume, E current is the real-time energy consumption, A polluted is the ecological disturbance area, V target 、E baseline 、A total They are target desilting volume, benchmark energy consumption and total area of the operating area; A, B and C are weight coefficients adjustable in the range of 0-1.
7. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: The target multimodal perception system in S3 includes a phased sonar array and a binocular vision module; the phased sonar array is used to model the spatial position of hard targets on the seabed; the binocular vision module is used to identify flexible obstacles using the YOLOv8s network.
8. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: In S3, the expression of the dynamic probability model is: Where d is the distance between the ship and the obstacle detected in real time, d safe is the safety threshold distance, k is the sensitivity adjustment coefficient; P collision is the collision probability, according to P collision The value of triggers the third level warning; When there is a collision risk, the dynamic window method is used to generate candidate obstacle avoidance paths with continuous curvature, and the optimal obstacle avoidance path is selected by optimizing the objective function, where the objective function is expressed as: f(v,ω)=α·J dist +β·J vel +γ·J obs Among them, α, β, and γ are weight coefficients; J dist is the path distance cost, J vel is the speed smoothing cost, J obs Obstacle avoidance cost.
9. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, characterized in that: In S4, the objective function of the global optimization model is to minimize the sum of the time taken by all ships to complete their respective tasks; The constraints include area coverage constraint, safety distance constraint and load balancing constraint; among them, the area coverage constraint requires that the union of all ship operating areas must completely cover the total operating area, with a coverage rate of 100%; the safety distance constraint requires that the distance between any two different ships must be no less than 50 meters; the load balancing constraint requires that the difference between the maximum and minimum dredging volume of a single ship does not exceed 15% of the average dredging volume of all ships.
10. The method for autonomous path planning of a dredging vessel based on reinforcement learning according to claim 1, wherein S4 include: Divide the operation area into n initial sub-areas and assign initial task boundaries to each ship; Each ship plans the optimal path within its assigned sub-area, calculates local variables, collects global constraint violation information, and updates each ship's mission boundaries using Lagrange multipliers until convergence. Convergence is determined when the task change between iterations is less than 1%, the maximum number of iterations reaches a preset value, or the rate of change of the global objective function is less than 0.5%. When a ship's path changes by more than 20% due to obstacle avoidance, or an unmarked obstacle is detected in the mission area, the trigger condition is considered to have been met. At this time, an incremental update strategy is used to reallocate the affected sub-areas, and a hot start strategy is used to inherit the multiplier parameters of the previous iteration until convergence, and a collaborative optimization solution is output.
Citation Information
Patent Citations
Robot motion planning method based on digital twinning and reinforcement learning
CN115903825A
Unmanned ship adaptive wave interference control method based on digital twinning
CN119045503A
Generating complete three-dimensional scene geometries using machine learning
DE102023132572A1
Adaptive cruise control using future trajectory prediction for autonomous systems and applications
US20240059285A1
Cited By
Underwater operation path adaptive planning method and system combined with machine learning
CN121722144A
An obstacle grid map acquisition method and device, electronic equipment and storage medium
CN122486592A
An obstacle grid map acquisition method and device, electronic equipment and storage medium
CN122486592B