Unmanned ship adaptive optimal interference control method based on reinforcement learning

By employing an adaptive optimal disturbance control method based on reinforcement learning, the problems of navigation stability and trajectory tracking of unmanned vessels in complex marine environments are solved. This method enables online estimation and suppression of unknown disturbances, thereby improving the control performance of unmanned vessels.

CN121325955AActive Publication Date: 2026-01-13LUDONG UNIVERSITY

Patent Information

Application Number
CN202511249740.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2026-01-13
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Traditional unmanned surface vessel (USV) control methods struggle to achieve ideal control performance when dealing with dynamic marine environmental disturbances and nonlinear systems, especially under complex disturbances such as waves, wind, and ocean currents, which affect navigation stability and path tracking accuracy.

Method used

An adaptive optimal disturbance control method based on reinforcement learning is adopted. By establishing the trajectory kinematics and dynamics model of the unmanned vessel, introducing a radial basis function neural network, designing an adaptive disturbance observer and a disturbance suppression controller, and using the adaptive vector backstepping method and reinforcement learning techniques, the online estimation and suppression of unknown disturbances can be achieved.

Benefits of technology

It improves the navigation stability and trajectory tracking accuracy of unmanned surface vessels in complex marine environments, enhances anti-interference capabilities, and improves control reliability, enabling unmanned surface vessels to track positions with the expected results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121325955A_ABST
    Figure CN121325955A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship adaptive optimal interference control method based on reinforcement learning, and particularly relates to the technical field of unmanned ship automatic control, and the method comprises the steps: building an unmanned ship trajectory kinematics model based on the position information and heading angle information of an unmanned ship under a geodetic coordinate system, and the corresponding speed information under an unmanned ship appendage coordinate system; the method comprises the following steps: establishing an unmanned ship trajectory dynamics model by considering the control operation of the unmanned ship and the time-varying environment interference problems of wind, waves, flow, unmodeled dynamics and the like in a marine environment in which the unmanned ship is located; introducing a radial basis function neural network based on a set unmanned ship trajectory mathematical model; designing a self-adaptive interference observer to estimate and offset time-varying environment interference in unmanned ship trajectory tracking; based on a radial basis function neural network and an interference observer, an unmanned ship adaptive interference suppression controller is designed by using an adaptive vector backstepping method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned ship automatic control, and particularly relates to an unmanned ship adaptive optimal disturbance control method based on reinforcement learning. BACKGROUND

[0002] As an important part of modern marine technology, unmanned surface vehicles (USVs) play an increasingly important role in resource exploration, maritime rescue, data collection and other fields. However, unmanned ships face complex marine environments and various disturbances when performing tasks, such as waves, wind and ocean currents, which seriously affect the navigation stability and path tracking accuracy of unmanned ships. Traditional unmanned ship control methods, such as PID control, sliding mode control and neural networks, can achieve trajectory tracking to some extent, but when dealing with dynamic changes in disturbances and nonlinear systems, it is often difficult to achieve ideal control results. SUMMARY

[0003] To this end, the present application provides an unmanned ship adaptive optimal disturbance control method based on reinforcement learning to solve the problems raised in the background art.

[0004] To achieve the above purpose, the present application provides the following technical solution: an unmanned ship adaptive optimal disturbance control method based on reinforcement learning, comprising: step 1, based on the position information and heading angle information of the unmanned ship in the earth coordinate system, and the corresponding velocity information in the unmanned ship body coordinate system, establishing an unmanned ship trajectory kinematics model; Step 2, considering the time-varying environmental disturbances such as wind, wave, current and unmodeled dynamics in the ocean environment where the unmanned ship is operating, establishing an unmanned ship trajectory dynamics model; Step 3, based on the set unmanned ship trajectory dynamics model, introducing a radial basis function (RBF) neural network; Step 4, designing an adaptive disturbance observer to estimate and cancel the time-varying environmental disturbances in the unmanned ship trajectory tracking; Step 5, based on the radial basis function (RBF) neural network and the disturbance observer, an adaptive disturbance suppression controller for the unmanned ship is designed using the adaptive vector backstepping method; Step 6, setting the desired trajectory of the unmanned ship, testing whether the optimal disturbance suppression can be achieved using the adaptive disturbance suppression controller for the unmanned ship, and completing the adaptive optimal control task of the unmanned ship trajectory.

[0005] Preferably, the unmanned ship trajectory kinematics model is specifically represented as: (1); (2); In formula (1), formula (2), is a position vector in the North-East coordinate system, which is composed of the actual position of the unmanned ship and the bow angle ; is a corresponding velocity vector in the appendage coordinate system, which is composed of the forward speed of the unmanned ship , the cross drift speed and the yaw angular velocity ; is defined as a rotation matrix, which represents the rotation matrix from the appendage coordinate system to the North-East coordinate system.

[0006] Preferably, the unmanned ship trajectory dynamics model is specifically represented as: (3); In formula (3), is an inertia matrix containing added mass; is a Coriolis centripetal force matrix; is a damping matrix; is a control vector provided by the propeller, which is composed of the longitudinal control force , the transverse control force and the yaw control moment ; is an equivalent time-varying force and moment vector acting on the ship body by the ocean environment such as wind, wave and current; is an equivalent time-varying force and moment vector acting on the ship body when the propeller of the unmanned ship fails, is a corresponding velocity vector in the appendage coordinate system, which is composed of the forward speed of the unmanned ship , the cross drift speed and the yaw angular velocity ; (4); In formula (4), is the mass of the unmanned ship; is the added mass caused by the motion of the unmanned ship; is the moment of inertia; is the distance between the ship center and the origin of the established coordinate system; (5); In formula (5), is the forward speed of the unmanned ship; are all damping coefficients.

[0007] Preferably, the specific process of step 3 is: A radial basis function (RBF) neural network is introduced, and its expression is as follows: wherein, is represented as an input vector, q represents a q-dimensional real number vector space, κ(z) = [κ1(z), …, κp(z)] represents a p-dimensional weight vector, and p represents a p-dimensional real number vector space. p (z)] T ∈ p is represented as a basis vector is represented as a weight vector, p > 1 is the number of nodes, and P represents a p-dimensional real number vector space.

[0008] Preferably, the specific process of step 4 is as follows: According to the unknown interference vector in formula (1) and (3) , the following adaptive interference observer is designed: (7); In formula (7), is the estimated value of the interference, is the gain matrix of the adaptive interference observer, is the auxiliary intermediate vector generated by formula (7); The estimation error vector of the interference observer is defined as : (8).

[0009] Preferably, the specific process of step 5 is as follows: Let the position error of the unmanned ship be , and the derivative of the error between positions is obtained: (9); In formula (9), is defined as a rotation matrix, which represents the rotation matrix from the appendage coordinate system to the north-east coordinate system, represents the expected position of the ship and the yaw angle ; The cost function is constructed as follows: (10); In formula (10), , wherein represents the distance between two variables; Replace the cross drift velocity vector with the required virtual control to define the Hamilton-Jacobi-Bellman equation as follows: (11); In formula (11), is represented as the gradient of the cost function with respect to , and is the estimated value of the virtual control law ; Based on reinforcement learning, the Critic network is used to evaluate the control performance, i.e.: (12); In formula (12), represents the cost function The gradient of the estimate value of , the gradient of the estimate value of represents the weight, is the time difference error, is the Critic network parameter, and the Actor network is used to perform control, as follows: (13); In formula (13), is the Actor network parameter, is the time difference error; The gradient of the cost function under the Critic network: (14); In formula (14), represents The estimate value of the adaptive law under the Critic network; The virtual intermediate vector of the unmanned ship under the Actor network: (15);

[0010] In formula (15), represents The estimate value of the adaptive law under the Actor network;

[0011] Wherein, the adaptive law under the Critic and Actor networks is: (16); In formula (16), represents the parameter of the Actor network, represents the parameter of the Critic network; Let the velocity vector error of the unmanned ship be , and thus, the derivative thereof is: (17); In formula (17), represents the control law of the unmanned ship, represents the external disturbance of the unmanned ship, and is constructed into a cost function, as follows: (18); In formula (18), represents the distance between two parameters; The Hamilton-Jacobi-Bellman equation is defined as follows: (19); In equation (19), is expressed as a cost function with respect to a gradient, denotes the estimated value of the unmanned ship control law; By adopting reinforcement learning technology, the Critic network is used to evaluate the control performance, i.e. (20); In equation (20), is expressed as a cost function with respect to a gradient of the estimated value, denotes the weight after updating, denotes the temporal difference error after updating, denotes the Critic network parameter after updating; The Actor network is used to perform control as follows: (21); In equation (21), denotes the estimated value of the control law , the estimated value of the disturbance , the estimated value of the disturbance , the estimated value of the disturbance , the estimated value of the disturbance , the estimated value of the disturbance , the estimated value of the disturbance , the estimated value of the disturbance (22); In equation (22), denotes the estimated value of the adaptive law after updating under the Critic network; An adaptive disturbance rejection controller for the unmanned ship under the Actor network is designed as follows: (23); In equation (23), denotes the estimated value of the adaptive law after updating under the Actor network; Wherein the adaptive law rule under the Critic and Actor networks is: (24); In equation (24), denotes the parameter of the Actor network after updating, represent the parameters of the Critic network after updating.

[0012] The control method of the application is optimal trajectory control for unmanned ships, uses an adaptive disturbance observer, solves the problem that a traditional disturbance observer needs prior information or model dynamic parameter structure, solves the online estimation and suppression problem of unknown external marine environment disturbance suffered by unmanned ship trajectory tracking; in addition, the adaptive optimal disturbance suppression method compensates for unknown frequency harmonic disturbance, overcomes the shortcoming that a traditional disturbance estimation needs prior information of external disturbance frequency, effectively enhances the anti-interference ability of the unmanned ship, improves the reliability of the unmanned ship control, and makes the tracking position of the unmanned ship achieve the expected effect.

[0013] The application considers the actual performance of the unmanned ship adaptive optimal disturbance suppression control based on reinforcement learning, has low cost, and is easy to implement in engineering. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a control flowchart of the application.

[0015] Figure 2 is a trajectory tracking diagram of the unmanned ship established by the embodiment of the application.

[0016] Figure 3 is a position tracking diagram of the unmanned ship established by the embodiment of the application.

[0017] Figure 4 is a speed tracking diagram of the unmanned ship established by the embodiment of the application.

[0018] Figure 5 is an adaptive neural network weight estimation diagram of the unmanned ship established by the embodiment of the application.

[0019] Figure 6 is a control law diagram of the unmanned ship established by the embodiment of the application.

[0020] Figure 7 is an adaptive disturbance suppression diagram of the unmanned ship established by the embodiment of the application.

[0021] Figure 8 is a Critic neural network weight updating diagram established by the embodiment of the application.

[0022] Figure 9 is an Actor neural network weight updating diagram established by the embodiment of the application. DETAILED DESCRIPTION

[0023] The following specific embodiments illustrate the embodiments of the present application, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosure. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0024] To improve the navigation stability of unmanned ships under interference conditions, those skilled in the art began to explore control strategies based on reinforcement learning (RL). Reinforcement learning is a machine learning method that learns optimal strategies by interacting with the environment, which has shown great potential in dealing with uncertainty and complexity. The core of reinforcement learning is that the agent learns through interaction with the environment, and constantly adjusts its behavior to maximize cumulative rewards. In the field of unmanned ship control, this means that unmanned ships can adapt to changing marine environments through reinforcement learning, automatically adjusting their control strategies to achieve optimal trajectory tracking and interference suppression. To deal with random interference, some studies have proposed adaptive control techniques that ensure the output error of the unmanned ship is always within a pre-specified range, improving the autonomous navigation capability of the unmanned ship in complex water environments. This adaptive control can update control parameters online to adapt to changes in the environment and the dynamic characteristics of the system.

[0025] Therefore, the reinforcement learning-based adaptive optimal interference control method for unmanned ships is a response to the need for improved navigation stability and trajectory tracking accuracy in the field of unmanned ship control, as well as the effectiveness of reinforcement learning techniques in dealing with complex environmental disturbances. These studies provide new control strategies for unmanned ships in the face of disturbances in marine environments, promoting the development and application of unmanned ship technology.

[0026] The reinforcement learning-based adaptive optimal interference control method for unmanned ships proposed by the present application is mainly aimed at optimal trajectory tracking control research for unmanned ships considering unknown external marine environment disturbances. By converting the mathematical model of unmanned ship motion and interference into a linear parameterized form, and integrating unknown exogenous systems, interference filters, neural networks and adaptive backstepping methods, the adaptive disturbance observer solves the problem of requiring prior information or model dynamic parameter structure for traditional disturbance observers. The designed adaptive disturbance suppression controller does not require any model parameters and model structure, but only relies on control input and position output signals, effectively enhancing the anti-interference ability of the unmanned ship, improving the reliability of the unmanned ship control, and achieving the expected effect of tracking the position of the unmanned ship.

[0027] As Figure 1As shown, step 1, based on the position information and the heading angle information of the unmanned ship in the earth coordinate system and the corresponding velocity information in the unmanned ship body coordinate system, an unmanned ship trajectory kinematics model is established, which is specifically represented as: (1); (2); In formula (1), (2), is a position vector in the north-east coordinate system, which is composed of the actual position of the unmanned ship and the heading angle ; is a corresponding velocity vector in the body coordinate system, which is composed of the forward speed of the unmanned ship , the lateral drift speed and the yaw angle speed ; is defined as a rotation matrix, which represents the rotation matrix from the body coordinate system to the north-east coordinate system.

[0028] Step 2, considering the problem of time-varying environmental disturbances such as wind, wave, current and unmodeled dynamics when the unmanned ship performs maneuvering operation and is in the marine environment, an unmanned ship trajectory dynamics model is established, which is specifically represented as: (3); In formula (3), is an inertia matrix containing added mass; is a Coriolis centripetal force matrix; is a damping matrix; is a control vector provided by the propeller, which is composed of the longitudinal control force , the lateral control force and the yaw control moment ; is an equivalent time-varying force and moment vector acting on the ship body by the marine environmental disturbances such as wind, wave and current. is an equivalent time-varying force and moment vector acting on the ship body when the propeller of the unmanned ship fails, is a corresponding velocity vector in the body coordinate system, which is composed of the forward speed of the unmanned ship , the lateral drift speed and the yaw angle speed .

[0029] (4); In formula (4), is the mass of the unmanned ship; is the added mass caused by the motion of the unmanned ship; is the moment of inertia; is the distance between the ship center and the origin of the established coordinate system.

[0030] In formula (5), u is the forward speed of the unmanned ship; X u ,Y υ ,Y r ,N υ ,N r are damping coefficients Step 3, based on the set unmanned ship trajectory dynamics model, a radial basis function (RBF) neural network is introduced, and the specific process is as follows: A radial basis function (RBF) neural network is introduced, and its expression is as follows: wherein, is an input vector, q represents a q-dimensional real number vector space, κ(z) = [κ1(z), …, κ p (z)] T ∈p represents a basis vector is a weight vector, p > 1 is the number of nodes, and P represents a p-dimensional real number vector space.

[0031] Step 4, an adaptive disturbance observer is designed to estimate and counteract the time-varying environmental disturbance in the unmanned ship trajectory tracking, and the specific process is as follows: According to the unknown disturbance vector in formula (1) and (3), the following adaptive disturbance observer is designed: (7); In formula (7), is the estimated value of the disturbance, is the gain matrix of the adaptive disturbance observer, is an auxiliary intermediate vector generated by formula (7); The estimation error vector of the disturbance observer is defined as : (8); Step 5, based on the radial basis function (RBF) neural network and the disturbance observer, an adaptive disturbance suppression controller for the unmanned ship is designed using the adaptive vector backstepping method, and the specific process is as follows: Let the position error of the unmanned ship be , and the error between the positions is derived to obtain: (9); In formula (9), is defined as a rotation matrix, which represents the rotation matrix from the appendage coordinate system to the north-east coordinate system, represents the desired position of the ship and the yaw angle ; a cost function is constructed as follows: (10); In formula (10), where denotes the distance between two variables; replace the cross drift velocity vector with the desired virtual control to define the Hamilton-Jacobi-Bellman equation as follows: (11); In formula (11), denotes the cost function with respect to the gradient, is the estimated value of the virtual control law .

[0032] By adopting reinforcement learning techniques, the Critic network is used to evaluate the control performance, i.e. (12); In formula (12), denotes the cost function with respect to the gradient of the estimated value, denotes the weight, is the time difference error, is the Critic network parameter, and the Actor network is used to perform control as follows: (13); In formula (13), is the Actor network parameter, is the time difference error; The gradient of the cost function under the Critic network is designed as: (14); In formula (14), denotes the estimated value of the adaptive law under the Critic network; The virtual intermediate vector of the unmanned ship under the Actor network is designed as: (15); In formula (14), denotes the estimated value of the adaptive law under the Critic network; where the adaptive law under the Critic and Actor networks is: (16); In formula (16), parameters of the Actor network, parameters of the Critic network; Let the velocity vector error of the unmanned ship be Thus, the derivative is (17); In equation (17), denotes the control law of the unmanned ship, denotes the external disturbance of the unmanned ship, and the cost function is constructed as follows: (18); In equation (18), denotes the distance between two parameters.

[0033] The Hamilton-Jacobi-Bellman equation is defined as follows: (19); In equation (19), denotes the cost function with respect to , denotes the estimated value of the control law of the unmanned ship.

[0034] By using reinforcement learning techniques, the Critic network is used to evaluate the control performance, i.e. (20); In equation (20), denotes the cost function with respect to , denotes the weight after updating, denotes the temporal difference error after updating, denotes the Critic network parameters after updating; The Actor network is used to perform control as follows: (21); In equation (21), denotes the estimated value of the control law , denotes the estimated value of the disturbance , denotes the Actor network parameters after updating, denotes the temporal difference error after updating; Next, the gradient of the cost function under the Critic network is designed as follows: (22); In formula (22), represents The estimated value of the adaptive law after updating under the Critic network; An adaptive disturbance rejection controller for the unmanned ship under the Actor network is designed: (23). In formula (23), represents The estimated value of the adaptive law after updating under the Actor network; Wherein the adaptive law rules under the Critic and Actor networks are: (24). In formula (24), represents the parameters of the Actor network after updating, represents the parameters of the Critic network after updating.

[0035] Step 6, set the desired trajectory of the unmanned ship, test whether the optimal disturbance rejection can be achieved by using the adaptive disturbance rejection controller for the unmanned ship, and complete the adaptive optimal control task of the unmanned ship trajectory: To verify the performance of the designed adaptive disturbance rejection controller for the unmanned ship, a 1:70 scale model ship CyberShip II (a test ship according to the scale of the supplied ship, with a length of 1.3m) is taken as the research object, and the dynamic parameters of the ship are: ; ; ; The desired trajectory of the unmanned ship is set as:

[0036]

[0037]

[0038] Suppose the external ocean environment disturbance vector of the unmanned ship is: ; Suppose the initial state of the unmanned ship is: , .

[0039] Take the gain parameters in the adaptive disturbance observer , design parameters , , , .​

[0040] To verify the effectiveness of the method of the application, simulation experiments are carried out from Figures 2-9 It can be seen that the control superiority of the method of the application, Figure 2 The position tracking graph of the unmanned ship shows that the proposed control strategy can overcome environmental disturbances and make the unmanned ship track the required tracking trajectory with arbitrary precision. Figure 3 The velocity tracking graph of the unmanned ship further shows that the unmanned ship can track the expected trajectory. Figure 4 The velocity tracking graph of the unmanned ship shows that the velocity of the unmanned ship is bounded and reasonable. Figure 5 The adaptive neural network weight estimation graph of the unmanned ship. Figure 6 The control law graph of the unmanned ship shows that the adaptive disturbance rejection controller of the unmanned ship can make the unmanned ship reach and track the desired trajectory with arbitrary precision. Figure 7 The adaptive disturbance rejection graph of the unmanned ship shows that the adaptive disturbance observer solves the online estimation and rejection problem of unknown external marine environment disturbances in the tracking of the trajectory of the unmanned ship. Figure 8 and Figure 9 The Critic and Actor neural network weight update graphs respectively show that the unmanned ship is trained in a short time and then stabilized to achieve tracking effect.

[0041] Although the application has been described in detail above with general description and specific embodiments, some modifications or improvements can be made on the basis of the application, which is obvious to those skilled in the art. Therefore, these modifications or improvements made on the basis of not deviating from the spirit of the application are within the scope of the application claimed.

Claims

1. A method for adaptive optimal disturbance control of unmanned surface vessels based on reinforcement learning, characterized in that: include: Step 1: Based on the unmanned vessel's position and heading angle information in the geodetic coordinate system, and the corresponding velocity information in the unmanned vessel's attached coordinate system, establish the unmanned vessel's trajectory kinematic model; Step 2: Considering the unmanned vessel's maneuvering operations and the wind, waves, currents, and unmodeled dynamic time-varying environmental disturbances in the marine environment, establish an unmanned vessel trajectory dynamics model; Step 3: Based on the established unmanned vessel trajectory dynamics model, a radial basis function neural network is introduced; Step 4: Design an adaptive disturbance observer to estimate and cancel time-varying environmental disturbances in unmanned vessel trajectory tracking; Step 5: Based on the radial basis function neural network and the interference observer, an adaptive vector backstepping method is used to design an adaptive interference suppression controller for unmanned vessels. Step 6: Set the desired trajectory of the unmanned vessel, and use the unmanned vessel adaptive interference suppression controller to test whether optimal interference suppression can be achieved, thus completing the unmanned vessel trajectory adaptive optimal control task.

2. The adaptive optimal disturbance control method for unmanned surface vessels based on reinforcement learning according to claim 1, characterized in that: The kinematic model of the unmanned vessel trajectory is specifically represented as follows: (1); (2); In equations (1) and (2), The position vector in the northeast coordinate system is determined by the actual position of the unmanned vessel. and heading angle constitute; This is the velocity vector in the attached coordinate system, which consists of the unmanned vessel's forward velocity. Horizontal drift speed and bow roll rate ; Defined as a rotation matrix, it represents the rotation matrix from the attached coordinate system to the northeast coordinate system.

3. The adaptive optimal disturbance control method for unmanned surface vessels based on reinforcement learning according to claim 1, characterized in that: The unmanned vessel trajectory dynamics model is specifically represented as follows: (3); In equation (3), The inertia matrix includes the added mass; Here is the Coriolis centripetal force matrix; Here is the damping matrix; The control vector provided for the thruster is determined by the oscillation control force. Horizontal control force and bow roll control torque composition; The equivalent time-varying force and moment vector of marine environmental disturbances on the ship hull; This represents the equivalent time-varying force and torque vector on the hull when the unmanned vessel's propulsion system malfunctions. This is the velocity vector in the attached coordinate system, which consists of the unmanned vessel's forward velocity. Horizontal drift speed and bow roll rate ; (4); In equation (4), For the quality of unmanned ships; The additional mass caused by the motion of the unmanned vessel; It is the moment of inertia; This is the distance between the ship's center and the origin of the established coordinate system; (5); In equation (5), The forward speed of the unmanned vessel; Both are damping coefficients.

4. The adaptive optimal disturbance control method for unmanned surface vessels based on reinforcement learning according to claim 2, characterized in that: Step 3 is as follows: A radial basis function neural network is introduced, and its expression is as follows: in, Let κ(z) be the input vector, q represent the q-dimensional real vector space, and κ(z) = [κ1(z), ..., κ2(z)]. p (z)] T ∈ p Represented as basis vectors It is represented as a weight vector, p>1 is the number of nodes, and P represents a p-dimensional real vector space.

5. The adaptive optimal disturbance control method for unmanned surface vessels based on reinforcement learning according to claim 3, characterized in that: Step 4 is as follows: Based on the unknown interference vector in equations (1) and (3) The following adaptive disturbance observer is designed: (7); In equation (7), This is an estimate of the interference. Here is the gain matrix of the adaptive interference observer. This is the auxiliary intermediate vector generated by equation (7); Define the estimation error vector of the disturbance observer as: : (8)。 6. The adaptive optimal disturbance control method for unmanned surface vessels based on reinforcement learning according to claim 4, characterized in that: Step 5 is as follows: Let the position error of the unmanned vessel be... Differentiating the error between positions, we obtain: (9); In equation (9), Defined as a rotation matrix, representing the rotation matrix from the attached coordinate system to the northeast coordinate system. Indicates the desired position of the ship and bow angle The cost function is constructed as follows: (10); In equation (10), ,in This represents the distance between two variables; The drift velocity vector Replace the Hamilton-Jacobi-Bellman equations with the desired virtual control, as shown below: (11); In equation (11), Represented as a cost function Compared to gradient, For virtual control laws The estimated value; Based on reinforcement learning, the Critic network is used to evaluate control performance, namely: (12); In equation (12), Representing the cost function Compared to The estimated value of the gradient, Indicates weight, For timing difference error, The parameters for the Critic network are as follows; the Actor network is used for execution control: (13); In equation (13), For Actor network parameters, This refers to timing difference error; Cost function under Critic network gradient: (14); In equation (14), express Estimated value of the adaptive law under Critic network; Virtual intermediate vector of unmanned vessel under Actor network: (15); In equation (15), express Estimated value of the adaptive law under the Actor network; The adaptive laws under Critic and Actor networks are as follows: (16); In equation (16), The parameters of the Actor network are represented. These represent the parameters of the Critic network; Let the velocity vector error of the unmanned vessel be... Therefore, differentiating it yields: (17); In equation (17), This represents the control law of the unmanned vessel. The external disturbances to the unmanned vessel are represented by the cost function, as follows: (18); In equation (18), Indicates the distance between two parameters; The Hamilton-Jacobi-Bellman equations are defined as follows: (19); In equation (19), Represented as a cost function Compared to gradient, This represents an estimate of the control law for the unmanned vessel; By employing reinforcement learning, the Critic network is used to evaluate control performance, namely: (20); In equation (20), Representing the cost function Compared to The estimated value of the gradient, This indicates the updated weights. This represents the updated timing difference error. This indicates the updated Critic network parameters; Actor networks are used for execution control, as follows: (21); In equation (21), Represents the control law The estimated value, Indicates interference The estimated value, This represents the updated Actor network parameters. This represents the updated timing difference error; Design the cost function under the Critic network. gradient: (22); In equation (22), express The estimated value of the adaptive law after the update under the Critic network; Design an adaptive interference suppression controller for unmanned surface vessels using an Actor network: (23); In equation (23), express The estimated value of the adaptive law after update under the Actor network; The adaptive laws under Critic and Actor networks are as follows: (24); In equation (24), This represents the parameters of the Actor network after the update. This indicates the parameters of the Critic network after the update.

Citation Information

Patent Citations

  • Self-adaptive control method for trajectory tracking of unmanned ship

    CN119311016A

  • Regulated performance optimal backstepping control method and system for unmanned ship formation

    CN119690071A

Cited By

  • Ship dynamic positioning anti-interference control method considering vertical dynamic optimization

    CN122284342A

  • A ship dynamic positioning anti-disturbance control method considering vertical dynamic optimization

    CN122284342B