Unmanned ship path tracking control method based on DDPG-MPC cooperative control
The DDPG-MPC cooperative control method solves the problems of high accuracy and robustness of path tracking for unmanned surface vessels in complex marine environments. By constructing dynamic and environmental models and combining LSTM and error feedback mechanisms, efficient and accurate path tracking of unmanned surface vessels in complex environments is achieved.
Patent Information
- Application Number
- CN202510848669.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-31
AI Technical Summary
Existing unmanned surface vessel (USV) path tracking control methods suffer from insufficient high precision, robustness, and real-time performance in complex marine environments. In particular, they are not robust enough under wind, wave, and current coupling interference. Furthermore, they have long training times and high online computational burdens, making it difficult to meet the high-frequency precision control requirements of USVs.
A method based on DDPG-MPC cooperative control is adopted. By constructing an unmanned surface vessel dynamics model and an environmental disturbance model, the path tracking problem is transformed into a Markov decision process. The DDPG algorithm is used to generate control objectives, and the MPC controller is used to generate control signals. An LSTM layer and a dynamic weighting mechanism are embedded in the Critic network to optimize the training process, enhance the evaluation capability of time-series states and key features, and achieve precise control by combining error feedback mechanism.
It improves the robustness and control accuracy of unmanned surface vessels in complex marine environments, reduces oscillations and overfitting during training, achieves flexible and precise path tracking control, and improves training stability and efficiency.
Smart Images

Figure CN120871842A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned surface vessel (USV) path tracking control technology, and in particular to a USV path tracking control method based on DDPG-MPC cooperative control. Background Technology
[0002] With the widespread application of unmanned surface vehicles (USVs) in fields such as marine monitoring and resource exploration, their autonomous navigation capabilities, especially high-precision and robust path tracking control technology, have become a core research focus. However, the marine environment is highly dynamic, uncertain (such as multi-source coupling of time-varying wind, waves, and current interference) and complex (obstacles and channel constraints), posing severe challenges to the path tracking control of USVs: it is necessary to achieve high tracking accuracy, strong environmental adaptability, and real-time response capabilities while meeting the physical constraints of actuators such as rudder angle / thrust.
[0003] In the paper "Wen Jia, Liang Xifeng, Wang Yongwei. Path tracking control of rice pollinating robot based on DDPG+MPC [J]. Agricultural Mechanization Research, 2025, 47(06):18-25," the proposed DDPG+MPC hybrid control algorithm improves the accuracy and real-time performance of path tracking for rice pollinating robots by combining deep reinforcement learning (DDPG) with model predictive control (MPC). Its shortcomings include long training time and insufficient environmental modeling. Although DDPG optimizes the performance of MPC, this method may still be limited in dynamically changing and uncertain environments. Furthermore, it does not fully consider the challenges of unmanned surface vessels (USVs) in extreme marine environments. First, the environmental modeling problem is underestimated; agricultural environmental interference is relatively small, while the strong dynamics and multi-source interference of the marine environment are not fully considered, resulting in insufficient robustness of existing methods under wind, wave, and current coupling interference. Second, some hybrid strategies are not designed for continuous optimization of path tracking, making it difficult to meet the high-frequency and precise control requirements of USVs. Finally, existing methods face problems such as long DRL training time and high online computational burden, making them difficult to implement on resource-limited USV platforms.
[0004] In the paper “Li W., Zhang X. Data-driven model predictive control for underactuated USV path tracking with unknown dynamics[J]. Ocean Engineering, 2025, 333:121457,” a path tracking method combining model predictive control (MPC) and dynamic fuzzy neural network (DFNN) is proposed to address the problem of unknown dynamic parameters in unmanned surface vessels (USVs). However, the accuracy of this method is still affected by DFNN modeling errors, and the model framework relies on traditional dynamic models, limiting its practical application effectiveness. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control, thereby solving the technical problems of existing technologies in terms of high precision, strong robustness, real-time performance, and constraint handling in complex marine environments.
[0006] This invention provides a path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control, comprising the following steps:
[0007] Step 1: Construct an unmanned surface vessel (USV) dynamics model; construct an environmental disturbance model based on the time-varying marine environment; transform the USV path-tracking problem into a Markov decision process;
[0008] Step 2: Construct a path tracking control model including the DDPG algorithm and the MPC controller. Use the DDPG algorithm to generate the control target and the MPC controller to generate the control signal. Use the output of the DDPG algorithm as part of the input of the MPC controller, and use the output of the MPC controller and the feedback of the actuator as part of the input of the DDPG algorithm. Through the error feedback mechanism, guide the control signal output by the MPC controller to dynamically approach the control target output by the DDPG algorithm.
[0009] Step 3: Optimize the DDPG algorithm, specifically by embedding LSTM and a dynamic weighting mechanism between the input layer and the fully connected layer in the Critic network of the DDPG algorithm;
[0010] Step 4: Train the path tracking control model and use the trained path tracking control model to perform path tracking control of the unmanned surface vessel.
[0011] Furthermore, the output parameters of the DDPG algorithm include:
[0012] ψ ref =μ(s) t ;θ μ )
[0013] u ref =v(s t ;θ v )
[0014] In the formula, ψ ref Represents the desired heading angle; u ref μ(s) represents the desired velocity. t ) and v(s t ;θ v These are the heading angle and velocity settings output by the Actor network in the DDPG algorithm.
[0015] Furthermore, the loss function of the DDPG algorithm is:
[0016]
[0017] in,
[0018]
[0019] In the formula, For LSTM networks, state s t The hidden state h after processing t .
[0020] Furthermore, the DDPG algorithm includes a dual Critic network.
[0021] Furthermore, the parameter update formula for the Critic network in the DDPG algorithm is as follows:
[0022]
[0023] in,
[0024]
[0025] Furthermore, the weighted sampling formula for the Actor network in the DDPG algorithm is as follows:
[0026]
[0027] In the formula, δ t For TD error, δ t =r t +γQ(s t+1 ,a t+1 )-Q(s t ,a t );h t h represents the hidden state after processing by the LSTM network. t =LSTM(s1,a1,...,s) t ,a t ); α is the contribution factor of the hidden state to the weighting; ∈ is a constant to prevent division by zero error; ||h t || represents the norm of the hidden state.
[0028] Furthermore, a target network is set in both the Critic network and the Actor network of the DDPG algorithm. The target network is:
[0029] θ Q′ ←τθ Q +(1-τ)θ Q′
[0030] θμ′ ←τθ μ +(1-τ)θ μ′
[0031] In the formula, τ << 1 represents the soft update rate.
[0032] Furthermore, the dynamic model of the MPC controller is as follows:
[0033] x t+1 =f(x) t ,u t )+g(d env )
[0034] In the formula, x t Let be the state vector of the unmanned surface vessel at time t. u t For the control input at time t, d env For environmental disturbance, f(x t ,u t ) is a function describing the dynamics of the unmanned surface vessel; g(d) env ) is a perturbation model.
[0035] Furthermore, the objective function of the MPC is:
[0036]
[0037] In the formula, e path (k) represents the path error, e path (k)=ψ k -ψ ref ;e control (k) represents the control error, e control (k)=u k -u ref J DDPG-Reward This is the reward information output by DDPG.
[0038] Furthermore, the rolling optimization process of the MPC is as follows:
[0039]
[0040] In the formula, e path (k) represents the path error, e path (k)=ψ k -ψ ref ;e control (k) represents the control error, e control (k)=u k -u ref J DDPG-Reward This is the reward information output by DDPG.
[0041] Furthermore, in step 2, the adjustment formula for error feedback is:
[0042]
[0043] In the formula, The adjusted MPC control signal; u MPC (t) represents the current control signal calculated by MPC; u DDPG (t) represents the optimal control signal calculated by DDPG; α is the adjustment factor, which determines the strength of the control feedback adjustment.
[0044] The beneficial effects of this invention are:
[0045] This invention provides a highly adaptive high-level decision-guided MPC through DDPG, combines the model constraints and rolling optimization capabilities of MPC, and utilizes LSTM to enhance the Critic's ability to evaluate temporal states and key features, thus achieving more flexible and accurate path tracking control for unmanned surface vessels. The collaborative architecture and optimized design of DDPG and MPC in this invention effectively reduce oscillations and overfitting during DDPG policy training, significantly improving training stability and the path tracking robustness of the control system under complex disturbance environments.
[0046] This invention addresses the shortcomings of traditional path-tracking methods, such as insufficient control accuracy and poor adaptability to complex marine environments. It uses the desired heading angle and velocity output by the Direct Driving Predator (DDPG) as dynamic reference inputs for the Dynamic Path Control (MPC). By introducing an environmental disturbance term into the prediction model and comprehensively considering path error, control error, and reward information from the DDPG output in the objective function, the invention effectively improves the control accuracy and system robustness of the MPC in complex marine environments, achieving more flexible and precise path-tracking control for unmanned surface vessels.
[0047] This invention constructs a dynamically weighted experience replay mechanism based on TD error and LSTM hidden states, enabling priority sampling of key training samples. Compared with the traditional uniform experience replay method, this mechanism can more effectively distinguish the importance of experience samples, focusing on learning experiences with large errors, improving training efficiency and convergence speed, and reducing oscillations and overfitting during policy training.
[0048] This invention introduces a Long Short-Term Memory (LSTM) layer into the Critic network to capture the temporal dependencies in the state-action sequences of unmanned surface vessels (USVs), addressing the problem of insufficient modeling of long-term dependencies in standard DDPGs under dynamic and time-varying environments. Through the gating mechanism of LSTM, the Critic network can effectively retain historical information, providing more accurate Q-value estimates, and thus providing a more stable gradient direction for Actor policy updates, significantly improving training stability and path tracking robustness. Attached Figure Description
[0049] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings:
[0050] Figure 1 This is a schematic diagram of the process framework in a specific embodiment of the present invention;
[0051] Figure 2 This is a schematic diagram of the coordinate system of an unmanned surface vessel in a specific embodiment of the present invention;
[0052] Figure 3 This is a schematic diagram illustrating the straight path tracking effect in a specific embodiment of the present invention;
[0053] Figure 4 This is a schematic diagram of straight path tracking error in a specific embodiment of the present invention;
[0054] Figure 5 This is a schematic diagram illustrating the curve path tracking effect in a specific embodiment of the present invention;
[0055] Figure 6 This is a schematic diagram of curve path tracking error in a specific embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] The present invention will be further illustrated below with reference to specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Modifications to the present invention in various equivalent forms all fall within the scope defined by the appended claims.
[0058] This invention provides a path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control, comprising the following steps:
[0059] Step 1: Construct an unmanned surface vessel (USV) dynamics model; construct an environmental disturbance model based on the time-varying marine environment; transform the USV path-tracking problem into a Markov decision process;
[0060] As an underactuated system, the motion of an unmanned surface vessel (USV) can be described by a three-degree-of-freedom model, including lateral, longitudinal, and yaw motions. The dynamic equations of the USV can be expressed as:
[0061]
[0062] In the formula, M is the inertia matrix; C(v) is the Coriolis force and centrifugal force matrix; D(v) is the damping matrix; ν = [u, v, r] T The velocity vector includes longitudinal velocity u, lateral velocity v, and yaw rate r; τ = [τ u ,τ v ,τ r ] T For control input; τ env For environmental disturbance.
[0063] The specific calculations for each matrix are as follows:
[0064]
[0065] The individual elements in these matrices (such as X) u Y v (etc.) are hydrodynamic coefficients, representing the hydrodynamic characteristics of the USV under different conditions. The above equations describe the motion of the USV in water, including its dynamic behavior in the longitudinal, lateral, and yaw directions.
[0066] Considering the underactuated characteristics of the unmanned surface vessel, the control input is mainly generated by the thrust T and thrust angle δ of the thruster, specifically expressed as:
[0067]
[0068] In the formula, x l This indicates the distance between the point of action of the thruster and the center of mass of the unmanned surface vessel.
[0069] The kinematic model of the unmanned surface vessel describes its motion in a geographic coordinate system, and its form is as follows:
[0070]
[0071] In the formula, represents the position information of the unmanned surface vessel, including the lateral position x, the longitudinal position y, and the heading angle ψ. R(ψ) is the heading angle rotation matrix, which is used to transform the velocity from the hull coordinate system to the geographic coordinate system.
[0072] The specific expressions for the force and torque vectors generated by the propulsion system of the unmanned surface vessel are as follows:
[0073] τ=[τ u ,0,τ r ]
[0074] Therefore, a three-degree-of-freedom model of the unmanned surface vessel can be established. The scalar form of the three-degree-of-freedom model is as follows:
[0075]
[0076] In the formula, m 11 m 22 m 33 d represents the diagonal elements of the rigid body inertia matrix; 11 d 22 d 33 These are the diagonal elements of the damping matrix.
[0077] The external environmental disturbances experienced by unmanned surface vessels during operation can be described as follows:
[0078] τ env =α(t)·v wind (t)+β(t)·v wave (t)+γ(t·v current (t)
[0079] In the formula, α(t), β(t), and γ(t) represent time-varying weighting coefficients adjusted according to the actual environment; v wind (t), v wave (t) and v current (t) represent environmental disturbances caused by wind, waves, and current, respectively;
[0080]
[0081] In the formula, V w (t) represents the wind speed at time t; θ w (t) represents the wind direction angle at time t; A w φ represents wave amplitude; k represents wave number; ω represents wave angular frequency; φ represents initial phase. V is a unit vector representing the direction of wave propagation; c (t) represents the speed of the ocean current; This represents the unit vector indicating the direction of ocean currents.
[0082] By modeling the path tracking problem of unmanned surface vessels as a Markov decision process (MDP), a state space and a thrust / torque action space containing multi-dimensional parameters such as position, velocity, and environmental disturbances are constructed. A composite reward function that integrates stability, energy consumption, and heading error is designed to achieve high-precision and robust intelligent path tracking control in complex marine environments.
[0083] The state space can be represented as:
[0084] s t ={x t ,y t ,ψ t ,u t ,v t ,r t ,v e(t),e path}
[0085] In the formula, x t The x-axis position of the unmanned surface vessel at time t; y t ψ represents the y-axis position of the unmanned surface vessel at time t; t U represents the heading angle of the unmanned surface vessel at time t; t ,v t ,r t Let v represent the forward velocity, lateral velocity, and heading speed of the unmanned surface vessel at time t, respectively; e (t) represents the velocity of time-varying ocean environmental disturbances; e path This indicates the path tracking error.
[0086] The action space can be represented as:
[0087] a t ={τ u ,τ R}
[0088] In the formula, τ u τ represents the thrust acting longitudinally on the hull. R This represents the torque acting on the hull about its vertical axis.
[0089] To improve training efficiency, a composite reward function was designed, taking into account both control stability and reward function (r). snd ), smooth reward function (r) smooth ) and heading angle error reward function (r ψ The goal of the reward function is to incentivize the unmanned surface vessel (USV) to travel along a predetermined path and minimize path tracking error. The reward function takes the following form:
[0090] Control stability reward function (r) snd ):
[0091]
[0092] In the formula, σ snd Let N be the standard deviation of the most recent N heading commands of the unmanned surface vessel; cos(k1·σ snd ) represents the control stability criterion based on standard deviation, used to measure the error change of the most recent N instructions; The second derivative of the standard deviation; It is the cumulative change in error, that is, the magnitude of error change over a period of time.
[0093] Energy consumption optimization reward function (r) en ):
[0094]
[0095] In the formula, τ u and τ R These represent the thrust magnitudes of the left and right thrusters, respectively.
[0096] Heading angle error reward function (r ψ ):
[0097]
[0098] In the formula, d is the yaw distance; |ψ ref -ψ| represents the heading angle deviation that needs to be corrected; α = 0.9 and β = 0.1.
[0099] Final reward function (J) DDPG-Reward ):
[0100] The final reward function consists of multiple sub-reward items, calculated using a weighted summation method:
[0101] J DDPG-Reward =w1r snd +w2r en +w3r ψ
[0102] Step 2: Construct a path tracking control model including the DDPG algorithm and the MPC controller. Use the DDPG algorithm to generate the control target and the MPC controller to generate the control signal. Use the output of the DDPG algorithm as part of the input of the MPC controller, and use the output of the MPC controller and the feedback of the actuator as part of the input of the DDPG algorithm. Through the error feedback mechanism, guide the control signal output by the MPC controller to dynamically approach the control target output by the DDPG algorithm.
[0103] In the DDPG module, an LSTM layer is added after the input layer of the Critic network to process continuous state-action sequences. This enhances the model's ability to handle temporal dependencies and improves the agent's value evaluation and action selection. After processing the sequence data, the LSTM layer generates the hidden state h. t This will be compared with the current state-action pair (s) t ,a t Together, they are input into subsequent fully connected layers. The formula for calculating the target Q-value is:
[0104]
[0105] In the formula, This indicates that the LSTM network is in state s t The hidden state h after processing t ;
[0106] In the DDPG algorithm, the Actor network is responsible for generating the desired control signal. To ensure that the DDPG output adapts to the MPC input, the DDPG primarily outputs the following two key control parameters:
[0107] ψ ref =μ(s) t ;θ μ )
[0108] u ref =v(s t ;θ v )
[0109] In the formula, ψ ref Represents the desired heading angle; u ref μ(s) represents the desired velocity. t ) and v(s t ;θ v These are the heading angle and speed settings output by the Actor network of DDPG, respectively.
[0110] In the MPC control module, an error feedback adjustment strategy is adopted to guide the MPC to approach the optimal strategy of DDPG in real time; the control signal output by DDPG is transmitted to the MPC input, and the MPC adjusts the ψ value output by DDPG. ref and u ref To track the target, the optimal control signal generated by DDPG is used as the reference signal for MPC, ensuring that the MPC control signal is adjusted towards the optimal path tracking control strategy. The error between the MPC control signal and the DDPG control signal is calculated in real time, and this error is used to adjust the MPC optimization target. MPC performs real-time optimization at each time step, while simultaneously reducing control error by guiding the optimal signal from DDPG, ultimately improving control accuracy.
[0111] The predictive model used by MPC describes the dynamic behavior of the unmanned surface vessel (USV). Assume the USV's state at time t is x. t The control input is u t And there is environmental disturbance d env This affects the state of the system. This dynamic model can be represented as:
[0112] x t+1 =f(x) t ,u t )+g(d env )
[0113] In the formula, It is the state vector of the unmanned surface vessel at time t; It is the control input at time t; It is an environmental disturbance; f(x) t ,u t) is a function describing the dynamics of an unmanned surface vessel, usually representing the system's state transition equation; g(d env ) is a disturbance model used to represent environmental disturbances d env It affects the state transitions of the system.
[0114] Construct the objective function:
[0115]
[0116] In the formula, e path (k)=ψ k -ψ ref It is path error; e control (k)=u k -u ref It is to control error; J DDPG-Reward It is the reward information output by DDPG, used to guide MPC to optimize path tracing.
[0117] During the rolling optimization process, MPC obtains the optimal control input sequence {u0, u1, ..., u} by solving the following optimization problem. H-1}:
[0118]
[0119] Step 3: Optimize the DDPG algorithm, specifically by embedding LSTM and a dynamic weighting mechanism between the input layer and the fully connected layer in the Critic network of the DDPG algorithm;
[0120] In DDPG, an LSTM layer is embedded between the input layer and the fully connected layer of the Critic network to process continuous state-action sequences and capture the temporal dependencies of the unmanned surface vessel's motion. The input to the LSTM is the state-action pair (s) of the current time step and k previous steps. t ,a t The sequence length k is dynamically adjusted according to the environment. The output of LSTM generates the hidden state h. t The current state-action pair is concatenated with the current state-action pair and then input into the subsequent fully connected layer to enhance the temporal correlation of Q-value estimation.
[0121] In DDPG, an experience replay buffer is added to store pairs of states, actions, rewards, and next states. Two independent Critic networks (Q1, Q2) are deployed, and the minimum of these is taken as the target Q-value to alleviate overestimation. To improve the utilization of the experience buffer, a dynamic weighting mechanism can be introduced to improve the replay buffer. The design of the dynamically weighted replay buffer is no longer based solely on the current TD error, but instead uses LSTM to process the sequence of historical states and actions and generate the hidden state h. tThis hidden state can be used as part of the weighted TD error to adjust the sampling probability of each sample. By calculating the TD-error for each experience, experiences with larger TD-errors are given higher weights. LSTM is used to process historical state-action sequences, enhancing long-term dependency learning. Combining the TD error with the sample sampling weights improves the model's learning efficiency for key experiences and optimizes the policy training process. The new weighted sampling formula is as follows:
[0122]
[0123] In the formula, δ t It is the TD error; h t =LSTM(s1,a1,...,s) t ,a t ) represents the hidden state after processing by the LSTM network, indicating the temporal dependency between the state and action; α is the contribution factor of the hidden state to the weighting; ∈ is a constant to prevent division by zero errors; ||h t || is the norm of the hidden state, reflecting the strength of LSTM's memory of historical information.
[0124] Step 4: Train the path tracking control model and use the trained path tracking control model to perform path tracking control of the unmanned surface vessel.
[0125] like Figure 1 As shown, the framework mainly consists of two parts: the DDPG module and the MPC module, which work together to achieve efficient path tracking control. In the DDPG module, the pre-planned path includes historical state sequence processing, an LSTM module for temporal feature extraction, and optimization using TD and LD errors. An experience replay mechanism stores state, action, reward, and next state data for dynamically updating network parameters. Simultaneously, the state-action value function outputs policy actions through the Actor network, and the Critic network evaluates the action value, jointly optimizing the decision-making process. The MPC module performs rolling optimization based on time parameters T and δ, adjusts the executed actions in conjunction with the rudder angle-thrust controller, and corrects deviations in real time through the tracking error calculation module. Constraint handling ensures that control commands conform to the limitations of the unmanned driving mechanics model, guaranteeing the feasibility and stability of the system.
[0126] like Figure 2 As shown, a path tracking control simulation of an unmanned surface vessel (USV) was conducted based on a three-degree-of-freedom (3-DOF) model, establishing an inertial coordinate system and a body coordinate system. The simulation focuses on describing the longitudinal and lateral translations and the yaw rotation in the horizontal plane. This modeling method reduces complexity while maintaining accuracy, achieving motion parameter conversion through coordinate transformation. It provides an efficient testing platform for control algorithm verification and supports extension to six-degree-of-freedom analysis. The simulation model parameters of the USV are shown in Table 1 below.
[0127]
[0128] Table 1
[0129] In the simulation, the initial position and heading angle of the unmanned surface vessel (USV) were set to random values to enhance the algorithm's adaptability to different initial conditions. The single-straight-line path was mainly used to test the algorithm's path-keeping ability and control stability, while the multi-segment path was used to verify the USV's turning ability and dynamic adaptability on complex paths.
[0130] During training, the DDPG algorithm relies on an experience replay pool and a soft update strategy to improve training stability and efficiency. To better train the DDPG algorithm, the hyperparameter settings are shown in Table 2 below:
[0131]
[0132] Table 2
[0133] MPC further performs local optimization to control the unmanned surface vessel's heading, thereby improving the stability and accuracy of heading control. The MPC training parameters are shown in Table 3 below:
[0134]
[0135] Table 3
[0136] The experimental path design included both single-straight-line and multi-segment paths to comprehensively evaluate the performance of the path tracking algorithm under different scenarios. Through experiments using these two paths, the accuracy, response speed, and reliability of the path tracking algorithm in the face of sudden environmental changes can be fully verified, ensuring that the unmanned surface vessel (USV) can maintain efficient and stable performance in various complex tasks. The specific parameters and settings of the paths are shown in Table 4 below:
[0137]
[0138] Table 4
[0139] Based on the collaborative optimization principle of DDPG and MPC, the pseudocode of the DDPG-MPC workflow is shown in Table 5 below:
[0140]
[0141]
[0142] Table 5
[0143] Building upon the aforementioned algorithms, DDPG and MPC are combined to effectively improve the accuracy and stability of path tracking through collaborative optimization. DDPG is responsible for optimizing action selection in path tracking through deep reinforcement learning, particularly for efficient policy updates in complex environments. MPC, on the other hand, performs real-time optimization control, ensuring that the system adjusts at each time step based on the current state and target path through rolling optimization, optimizing heading angle and thrust input. The advantage of this invention lies in its combination of the adaptive capabilities of deep learning and the precise adjustment capabilities of model predictive control, achieving accurate tracking and control in complex dynamic environments.
[0144] like Figure 3 As shown, the black reference path represents the ideal path of the target. DDPG-MPC combines reinforcement learning and model predictive control, exhibiting lower bias and faster convergence speed. While the DDPG method can track the path, its convergence speed is slower than DDPG-MPC, and the bias increases in the later stages. MPC can track the target path to some extent, but its adjustment speed is slow. Traditional PID control performs the worst, with large path fluctuations, especially in the latter half of the path. The comparative results show that the method of this invention not only significantly outperforms traditional methods in terms of accuracy and convergence speed, but also provides more stable control performance in complex path tracking tasks.
[0145] like Figure 4 As shown, the three sub-figures illustrate the position error and heading angle error in the x and y directions, respectively. The first sub-figure shows that the error in the x direction fluctuates significantly initially but gradually decreases to zero; the second sub-figure shows that the error in the y direction oscillates to some extent but gradually stabilizes over time; the third sub-figure shows the heading error, which initially oscillates but then stabilizes. The upper right corner inset of each sub-figure magnifies the local oscillations for easier observation of error fluctuations. Overall, this invention gradually converges after initial fluctuations and eventually achieves stability.
[0146] like Figure 5 As shown, the black dashed line represents the desired path. DDPG-MPC performs best, closely following the desired path with a smooth and stable trajectory. DDPG exhibits some deviation, especially at turns. MPC performs reasonably well on straight sections, but its stability is slightly worse at turning points. PID shows the most significant deviation, with large overall trajectory fluctuations, particularly noticeable on winding paths. Overall, the comparison shows that the integrated DDPG-MPC controller has stronger path fitting capabilities and robustness, making it suitable for unmanned surface vessel control on complex paths. After initial fluctuations, this invention gradually converges and achieves stable path tracking, demonstrating the efficiency and effectiveness of the DDPG-MPC-based collaborative control method, especially in handling initial fluctuations and long-term stability.
[0147] like Figure 6 As shown, the three sub-graphs illustrate the position error and heading angle error in the x and y directions, respectively. Initially, all three dimensions exhibit significant fluctuations, but the errors converge rapidly within 50 seconds, eventually approaching zero, demonstrating the system's good convergence and stability. The embedded magnified view further reveals the details of error fluctuations within the 25–30 second interval, verifying that the control strategy maintains high tracking accuracy and robustness even under complex paths. This invention's method, by combining DDPG and MPC collaborative control, effectively reduces error fluctuations in complex path tracking tasks, exhibiting good convergence, stability, and high-precision tracking capabilities, proving the effectiveness of this method in dynamic environments.
[0148] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control, characterized in that, Includes the following steps: Step 1: Construct an unmanned surface vessel (USV) dynamics model; construct an environmental disturbance model based on the time-varying marine environment; transform the USV path-tracking problem into a Markov decision process; Step 2: Construct a path tracking control model including the DDPG algorithm and the MPC controller. Use the DDPG algorithm to generate the control target and the MPC controller to generate the control signal. Use the output of the DDPG algorithm as part of the input of the MPC controller, and use the output of the MPC controller and the feedback of the actuator as part of the input of the DDPG algorithm. Through the error feedback mechanism, guide the control signal output by the MPC controller to dynamically approach the control target output by the DDPG algorithm. Step 3: Optimize the DDPG algorithm, specifically by embedding LSTM and a dynamic weighting mechanism between the input layer and the fully connected layer in the Critic network of the DDPG algorithm; Step 4: Train the path tracking control model and use the trained path tracking control model to perform path tracking control of the unmanned surface vessel.
2. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, The output parameters of the DDPG algorithm include: ψ ref =μ(s t ;θ μ ) u ref =v(s t ;θ v ) In the formula, ψ ref Represents the desired heading angle; u ref μ(s) represents the desired velocity. t ) and v(s t ;θ v These are the heading angle and velocity settings output by the Actor network in the DDPG algorithm.
3. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, The loss function of the DDPG algorithm is: in, In the formula, For LSTM networks, state s t The hidden state h after processing t .
4. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, The DDPG algorithm includes a dual Critic network.
5. The unmanned surface vessel path tracking control method based on DDPG-MPC cooperative control as described in claim 4, characterized in that, The parameter update formula for the Critic network in the DDPG algorithm is as follows: in, 6. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, The weighted sampling formula for the Actor network in the DDPG algorithm is as follows: In the formula, δ t For TD error, δ t =r t +γQ(s t+1 ,a t+1 )-Q(s t ,a t );h t h represents the hidden state after processing by the LSTM network. t =LSTM(s1,a1,...,s) t ,a t ); α is the contribution factor of the hidden state to the weighting; ∈ is a constant to prevent division by zero error; ||h t || represents the norm of the hidden state.
7. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, In the DDPG algorithm, a target network is set in both the Critic network and the Actor network. The target network is: i Q′ ←tth Q +(1-τ)θ Q′ i μ′ ←tth μ +(1-τ)θ μ In the formula, τ << 1 represents the soft update rate.
8. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in claim 1, characterized in that, The dynamic model of the MPC controller is as follows: x t+1 =f(x t ,u t )+g(d env ) In the formula, x t Let be the state vector of the unmanned surface vessel at time t. u t For the control input at time t, d env For environmental disturbance, f(x t ,u t ) is a function describing the dynamics of the unmanned surface vessel; g(d) env ) is a perturbation model.
9. The unmanned surface vessel path tracking control method based on DDPG-MPC cooperative control as described in claim 1 or 8, characterized in that, The objective function of the MPC is: In the formula, e path (k) represents the path error, e path (k)=ψ k -ψ ref ;e control (k) represents the control error, e control (k)=u k -u ref J DDPG-Reward This is the reward information output by DDPG; The rolling optimization process of MPC is as follows: In the formula, e path (k) represents the path error, e path (k)=ψ k -ψ ref ;e control (k) represents the control error, e control (k)=u k =u ref J DDPG-Reward This is the reward information output by DDPG.
10. The path tracking control method for unmanned surface vessels based on DDPG-MPC cooperative control as described in any one of claims 1-9, characterized in that, In step 2, the adjustment formula for error feedback is: In the formula, The adjusted MPC control signal; u MPC (t) represents the current control signal calculated by MPC; u DDPG (t) represents the optimal control signal calculated by DDPG; α is the adjustment factor, which determines the strength of the control feedback adjustment.
Citation Information
Cited By
Unmanned ship formation adaptive control method based on DDPG
CN121578812A