Ship zero-sum game tracking control method based on numerical and physical dual-driven reinforcement learning

By employing a dual-drive reinforcement learning approach combining physical and data-driven models, feedforward and feedback controllers are designed and optimal feedback control is optimized. This addresses the challenges of constructing physical models and the insufficient generalization of data-driven methods in ship motion control, enabling efficient and stable tracking in complex marine environments.

CN121559889BActive Publication Date: 2026-04-07DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing ship motion control technologies, physical model construction is difficult and costly, and data-driven methods rely on high-quality data and have insufficient generalization ability, making it difficult to meet the safety and reliability requirements in complex marine environments.

Method used

A dual-drive reinforcement learning approach based on data and physics is adopted. A dual-drive model based on data and physics is constructed by combining a ship physics model with a recurrent neural network. A feedforward controller is designed using the Lyapunov backstepping method. The optimal feedback controller is optimized by combining zero-sum game theory and fuzzy logic system to achieve optimal control under the worst-case disturbance strategy.

Benefits of technology

It achieves efficient and stable tracking and control of ships in complex marine environments, enhances the adaptive capability and anti-interference performance to marine disturbances, and ensures stable tracking of target tracks under unknown conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121559889B_ABST
    Figure CN121559889B_ABST
Patent Text Reader

Abstract

The application discloses a ship zero-sum game tracking control method based on data-physical dual driving and reinforcement learning, which comprises the following steps: constructing a data-physical dual driving model; adopting a Lyapunov backstepping method to design a ship track tracking feedforward controller in layers; defining a performance index function and a Hamilton-Jacobi-Isaacs equation based on an error system and a zero-sum game theory, and designing an optimal feedback controller and a worst disturbance strategy; using a fuzzy logic system to approximate a solution of the HJI equation, and obtaining a Nash equilibrium of estimated values of outputs of the optimal feedback controller and the worst disturbance strategy; and fusing the feedforward controller, the optimal feedback controller and the worst disturbance strategy, and performing constraint processing to obtain a final ideal ship controller which can directly execute an actual control instruction, so that ship track tracking is realized. The application solves the problems of inherent defects of existing methods, such as high dependence on high-quality data, weak physical interpretability, insufficient generalization ability for complex working conditions and the like, and cannot meet the safety and reliability requirements of ship navigation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of ship motion control, and particularly relates to a ship zero-sum game tracking control method based on numerical and physical dual driving reinforcement learning. BACKGROUND

[0002] With the deep penetration and integration of electronic information technology in ships, the field of ship motion control is ushering in an unprecedented development opportunity, laying a solid technical foundation. Especially under the strong promotion of continuous iteration and upgrading of sensor technology and Internet of Things technology, ships have the ability to automatically perceive, accurately collect and real-time transmit multi-dimensional information such as their own motion state. Deeply integrating these real-time acquired multi-source heterogeneous information and data into the existing control technology system will become the core driving force for driving the innovation and upgrading of ship motion control technology, realizing intelligent transition, and providing new possibilities for precise control of ships in complex marine environments.

[0003] During ship navigation, ships need to continuously face the coupled influence of complex marine disturbances such as wind, waves and currents. The interaction between the controller and the external disturbance essentially constitutes a typical "zero-sum game" relationship - the controller takes minimizing tracking error as the core goal, while the disturbance takes maximizing error as the "opposing goal", which brings serious challenges to navigation safety and operation efficiency, and puts forward higher requirements for the anti-interference ability and dynamic countermeasures performance of the ship control system. The introduction of game theory provides a new perspective for solving this antagonistic control problem. The organic integration of game theory, optimal control theory and reinforcement learning can build a dynamic countermeasure optimization framework for the controller and the disturbance, and realize the solution of the optimal control strategy under the worst disturbance strategy. This integrated technology path provides a solid theoretical support and reliable technical guarantee for the intelligent transformation of ship motion control, helping ships to realize safe, efficient and precise autonomous navigation in complex marine environments.

[0004] At the same time, in the existing ship motion control technology, most schemes rely on physical models to build control frameworks. However, the core parameters of the ship physical model need to be obtained through a large number of high-cost, long-period pool tests and ship tests, and are affected by the complexity of marine environment and the coupling characteristics of ship dynamics, making it difficult to achieve absolute precision. This practical bottleneck greatly restricts the research and development efficiency and performance breakthrough of ship control technology. Although a few studies attempt to use data-driven methods to achieve control, such methods generally have inherent defects such as high dependence on high-quality data, weak physical interpretability, and insufficient generalization ability for complex working conditions, making it difficult to meet the safety and reliability requirements of ship navigation. Therefore, the organic integration of the physical laws contained in the ship physical model and the adaptive advantages of data-driven methods to build a control method with both physical interpretability and dynamic adaptive compensation ability can not only avoid the limitations of single modeling methods, but also improve control accuracy, and has important research value and engineering application prospects. SUMMARY

[0005] The application provides a ship zero-sum game tracking control method based on data-physical dual driving reinforcement learning to overcome the above technical problems.

[0006] To achieve the above-mentioned purpose, the technical scheme of the application is:

[0007] A ship zero-sum game tracking control method based on data-physical dual driving reinforcement learning, specifically comprising the following steps:

[0008] S1: obtaining a ship physical model considering unknown wind and wave flow disturbance, input constraint and unknown dynamics;

[0009] S2: constructing a data-physical dual driving model according to the ship physical model combined with a recurrent neural network;

[0010] S3: defining an ideal control input of the data-physical dual driving model based on Lyapunov backstepping method, obtaining a track tracking error according to the data-physical dual driving model based on a preset expected ship track, and constructing a ship virtual controller according to the track tracking error; obtaining a ship speed error according to the ship virtual controller to define an intermediate parameter for optimal feedback control design, obtaining a rewritten ship speed error derivative according to the intermediate parameter combined with the ideal control input after derivation of the ship speed error; designing a ship feedforward controller according to the rewritten ship speed error derivative; defining a Lyapunov function according to the track tracking error and the ship speed error, and obtaining a derivative of the Lyapunov function by derivation;

[0011] S4: defining an error system based on the derivative of the Lyapunov function; constructing a performance index function based on the error system combined with zero-sum game theory; constructing a Hamilton-Jacobi-Isaacs equation according to the performance index function to obtain an optimal feedback controller and a worst disturbance strategy;

[0012] S5: approximating the optimal performance index function by using a fuzzy logic system to obtain an output estimation value of the optimal feedback controller and the worst disturbance strategy, and constructing a neural network weight update rate of the fuzzy logic system according to the output estimation value of the optimal feedback controller and the worst disturbance strategy;

[0013] S6: constructing a ship final ideal controller according to the ship feedforward controller combined with the optimal feedback controller and the worst disturbance strategy based on the neural network weight update rate; constraining the control input of the ship final ideal controller through the input constraint in the ship physical model to obtain an actual control input, and realizing the ship zero-sum game tracking control based on data-physical dual driving reinforcement learning according to the actual control input.

[0014] Further, the expression of the ship physical model in S1 is:

[0015] ,

[0016] wherein: denotes the position vector of the ship in the inertial coordinate system , and the yaw angle and ; denotes the velocity vector of the ship in the body coordinate system , and the yaw rate and ; denotes the coordinate transformation matrix ; denotes the control vector of the ship which can be directly executed , and the yaw moment and ; denotes the disturbance vector of the ship in the body coordinate system , and the yaw moment and , satisfying ; denotes an upper bound of a positive constant; denotes the matrix of the ship weight inertia and hydrodynamic added inertia ; denotes the Coriolis matrix ; denotes the hydrodynamic damping parameter matrix ; denotes the transpose; denotes the real number space; denotes the first derivative of ;

[0017] The input constraint is defined as:

[0018] ,

[0019] wherein: denotes the control force of each degree of freedom of the ship physical model ; denotes the ideal control input of the ship in the longitudinal, lateral and yaw directions; denotes the maximum value of the control input of each degree of freedom of the ship; denotes the minimum value of the control input of each degree of freedom of the ship.

[0020] Furthermore, step S2 specifically includes the following steps:

[0021] S21: The nominal dynamic model obtained from the ship's physical model, without considering input constraints, is as follows:

[0022] ,

[0023] In the formula: Represents the ideal control input vectors for the ship's longitudinal, lateral, and bow directions, and ;

[0024] S22: Based on the Stone-Weilstras theorem, the nominal dynamic model of the ship is rewritten using RNN basis functions as follows:

[0025] ,

[0026] In the formula: Describes the ideal weights of the RNN basis functions and satisfies ; This represents the estimated values ​​of the RNN network weights; This represents the error between the ideal weights and the estimated weights of the RNN basis functions; Indicates reconstruction error and ; Describes a vector composed of RNN basis functions and ; Represents the basis functions of a monotonically increasing RNN;

[0027] S23: Based on the revised nominal dynamics model, construct the approximate dynamics system of the ship as follows:

[0028] ,

[0029] ,

[0030] In the formula: Represents the estimated value of the ship data physical dual-drive model and satisfies ; Indicates modeling error; This represents the reconstruction error compensation feedback term; This represents the gain coefficient of the linear feedback term. ; The gain parameter estimate of the nonlinear feedback term and ; The gain parameter represents the ideal nonlinear feedback term; This represents the gain parameter error of the nonlinear feedback term; express The first derivative;

[0031] S24: constructing a data-physical dual-driven model according to the ship approximate dynamics system and the ship physical model is:

[0032] .

[0033] Further, the S3 specifically includes the steps of:

[0034] S31: defining an ideal control input of the data-physical dual-driven model based on the Lyapunov backstepping method,

[0035] The expression of the ideal control input is:

[0036] + ,

[0037] In the formula: expresses the output of the feedforward controller; expresses the output of the optimal feedback controller; expresses the output of the worst interference strategy;

[0038] S32: obtaining a track tracking error of the ship according to the definition of the data-physical dual-driven model is:

[0039] ,

[0040] In the formula: expresses the track tracking error; expresses the track reference signal and ;

[0041] S33: obtaining a track tracking error derivative by deriving the track tracking error, and rewriting the track tracking error derivative into:

[0042] ,

[0043] In the formula: expresses the first derivative of ;

[0044] S34: designing a ship virtual controller according to the rewritten track tracking error derivative is:

[0045]

[0046] In the formula: expresses the control design parameter and ; expresses the output of the ship virtual controller;

[0047] S35: Obtain the ship speed error according to the ship virtual controller and the data physical double drive model, to define an intermediate variable for optimal feedback control design, and the expression is:

[0048] ,

[0049] ,

[0050] , ,

[0051] In the formula: represents the ship speed error; represents the intermediate variable;

[0052] S36: Derive the ship speed error to obtain the derivative of the ship speed error, and obtain the rewritten derivative of the ship speed error according to the intermediate variable combined with the ideal control input;

[0053] ,

[0054] In the formula: represents the first derivative of ; represents the first derivative of ;

[0055] S37: Design the ship feedforward controller according to the rewritten derivative of the ship speed error as:

[0056] ,

[0057] In the formula: represents the control design parameter and ; represents the output of the ship feedforward controller;

[0058] S38: Define the Lyapunov function according to the track tracking error and the ship speed error as:

[0059] ,

[0060] and derive the Lyapunov function to obtain the derivative of the Lyapunov function as:

[0061] .

[0062] Further, the S4 specifically includes the following steps:

[0063] S41: Define the term in the derivative of the Lyapunov function as the error system:

[0064] ,

[0065] S42: Based on the error system combined with the zero-sum game theory, a performance index function under the zero-sum game is defined as:

[0066] ,

[0067] In the formula: t represents a time parameter; γ represents a discount factor; u represents a permissible control input; , , P represents a positive definite gain matrix, and , , ;

[0068] S43: According to the Nash-Pontryagin maximum-minimum value principle, it is confirmed that the zero-sum game has a unique solution provided that there is a unique saddle point, so that the optimal performance index function is:

[0069] ,

[0070] S44: According to the performance index function, the Hamilton function of the error system is defined, and the expression is:

[0071] ,

[0072] In the formula: ∂H / ∂x represents the partial derivative of with respect to ; x represents the transpose of ;

[0073] Based on the Hamilton function, according to the optimal performance index function , the Hamilton-Jacobi-Eikonal equation is obtained as:

[0074] ,

[0075] In the formula: ∂H / ∂x represents the partial derivative of with respect to ;

[0076] S45: According to the performance index function, the Hamilton-Jacobi-Eikonal equation is constructed, and the gradient descent method is used to obtain the optimal feedback controller and the worst disturbance strategy as:

[0077] ,

[0078] .

[0079] Further, the S5 specifically comprises steps of:

[0080] S51: approaching the optimal performance index function by using a fuzzy logic system, and its expression is:

[0081] ,

[0082] In the formula: represents the ideal weight and satisfies ; represents the error of the neural network weight; represents the estimated value of ; represents the fuzzy base function; represents the approximation error;

[0083] S52: obtaining the partial derivative of the optimal performance index function after approximation with respect to according to S51 is:

[0084] ,

[0085] In the formula: represents the partial derivative of with respect to ; represents the partial derivative of with respect to ;

[0086] S53: determining the evaluation network according to the estimated value of the weight , and the estimated value of the partial derivative of the optimal performance index function according to S52 is:

[0087] ,

[0088] S54: according to the partial derivative of S52, the optimal feedback controller in S45, the worst disturbance strategy and the Hamilton-Jacobi-Essix equation in S44 are brought in to obtain:

[0089] ,

[0090] ,

[0091] ,

[0092] In the formula: represents the reconstruction residual and ​​​; represents an intermediate variable and ; represents a short form of

[0093] Similarly, the partial derivative estimate value obtained in S53 is substituted into the optimal feedback controller in S45 and the worst disturbance strategy and the Hamilton-Jacobi-Essix equation in S54 to obtain the output estimate value of the optimal feedback controller, the output estimate value of the worst disturbance strategy and the rewritten Hamilton-Jacobi-Essix equation as follows:

[0094] ,

[0095] ,

[0096] ,

[0097] S55: The output estimate value of the optimal feedback controller and the worst disturbance strategy is substituted into the error system in S41 to obtain a rewritten error system as follows:

[0098] ,

[0099] S56: Based on the rewritten error system, a neural network weight update rate of a fuzzy logic system corresponding to the minimum of the rewritten Hamilton-Jacobi-Essix equation is constructed by using the gradient descent method as follows:

[0100] ,

[0101] In the formula: represents the weight learning rate and .

[0102] Further, the S6 specifically comprises the following steps:

[0103] S61: Based on the neural network weight update rate, a ship final ideal controller is constructed according to the output estimate value of the optimal feedback controller and the worst disturbance strategy combined with the ship feedforward controller as follows:

[0104] ,

[0105] S62: The control input of the ship final ideal controller is constrained by the input constraint in the ship physical model to obtain an actual control input; and the ship zero-sum game tracking control based on the data-physical dual-driven reinforcement learning is realized according to the actual control input.

[0106] The ship zero-sum game tracking control method based on data-physical dual-driven reinforcement learning has the following beneficial effects:​​

[0107] 1. Through the deep fusion of physical models and ship speed and other navigation data, a data-physical double driving model of "physical mechanism + data compensation" is constructed. The model not only retains the interpretability and stability of the physical model, but also effectively avoids the problem of relying on high-quality data for data-driven methods. Without relying on accurate ship parameters, efficient and stable control can be achieved.

[0108] 2. Based on the data-physical double driving model, a Lyapunov backstepping method is used to design a ship feedforward controller, which realizes accurate tracking of the nominal dynamics of the ship and provides a solid foundation for the stability of the closed-loop system.

[0109] 3. Through the organic fusion of zero-sum game differential theory and adaptive dynamic programming based on fuzzy logic system, the optimal feedback controller and the worst disturbance strategy are simultaneously optimized to construct a dynamic confrontation framework of controller and disturbance. This framework can accurately capture the "zero-sum game" nature of the controller and complex disturbance, realize the optimal feedback control solution under the worst disturbance strategy, and significantly improve the adaptive ability and anti-interference performance of the ship to random disturbances in the ocean, ensuring stable tracking of the target track in complex scenarios where the physical constraints and dynamic characteristics of the ship are completely unknown.

[0110] 4. By fusing the tracking ability of feedforward control, the optimization ability of optimal feedback control, and the adaptive ability of worst disturbance strategy, the ideal control input is obtained, and the actual control instruction directly executable is generated after physical constraint processing. The present application deeply fuses data-physical double driving, zero-sum game and reinforcement learning, not only solves the precision and robustness problem of ship tracking control in complex marine environment, but also promotes the intelligentization and adaptive upgrade of ship tracking control system, and provides reliable technical support for ship autonomous operation. BRIEF DESCRIPTION OF DRAWINGS

[0111] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0112] Figure 1 The flowchart of the ship zero-sum game tracking control method of the present application is shown in the figure.

[0113] Figure 2 The error diagram of the actual model and the data-physical double driving model in the present embodiment is shown in the figure.

[0114] Figure 3 The position coordinates of the ship in the present embodiment are shown in the figure. Simulation diagram of the tracking effect;

[0115] Figure 4 In this embodiment, the ships are respectively in Simulation results of tracking the reference signal above;

[0116] Figure 5 This is a simulation diagram showing the convergence result of the weight norm of the fuzzy logic system in this embodiment;

[0117] Figure 6 This is a simulation diagram showing the convergence results of the RNN weight norm and feedback term parameters in this embodiment;

[0118] Figure 7 This is a simulation diagram showing the results of the ship's forward force, lateral drift force, and bow roll moment controller in this embodiment;

[0119] Figure 8 This is a simulation diagram showing the results of the worst-case interference strategy for ships in this embodiment. Detailed Implementation

[0120] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0121] This embodiment provides a zero-sum game tracking control method for ships based on dual-driven reinforcement learning of numbers and objects, such as... Figure 1 As shown, the specific steps include:

[0122] S1: Obtain a three-degree-of-freedom ship physical model considering unknown wind, wave, and current disturbances, input constraints, and unknown dynamics; specifically, the ship physical model is the ship's kinematics and dynamics model, and its expression is:

[0123] ,

[0124] In the formula: Indicates the ship in the inertial coordinate system coordinate, Coordinates and bow roll angle The position vector formed and ; Represents the forward speed of the ship in the attached coordinate system. Horizontal drift speed and bow roll rate The velocity vector formed and ; denotes the coordinate transformation matrix and ; denotes the control vector composed of the actual controllable forward force , the lateral drift force and the yaw moment and ; denotes the disturbance vector composed of the wind and current induced lateral force , the longitudinal force and the yaw moment and , satisfying ; denotes an upper bound positive constant; denotes the matrix composed of the ship weight inertia and hydrodynamic added inertia and ; denotes the Coriolis matrix and ; denotes the hydrodynamic damping parameter matrix and ; denotes the transpose; denotes the real number space; denotes the first derivative of ; and

[0125] and the matrices , , , are respectively expressed in the following forms:

[0126] , ,

[0127] ,

[0128] ,

[0129] wherein: denotes the ship mass; denotes the moment of inertia; denotes the component of the axis direction in the body coordinate system; denotes the hydrodynamic derivative of each direction in the body coordinate system;

[0130] The input constraint, i.e. the physical constraint of the actual controller, is defined as:

[0131] ,

[0132] wherein: representing the control force of each degree of freedom of the ship physical model, and ; representing the ideal control input of the ship in the longitudinal, lateral and bow directions; representing the maximum value of the control input of each degree of freedom of the ship; representing the minimum value of the control input of each degree of freedom of the ship.

[0133] S2: According to the ship physical model combined with the recurrent neural network, a data-physical dual driving model is constructed, and in this embodiment, based on the ship physical model, the speed and attitude data are collected by the Doppler log, and the data-physical dual driving model is constructed combined with the recurrent neural network, which specifically includes the following steps:

[0134] S21: According to the ship physical model, the nominal dynamics model without considering input constraints is obtained as:

[0135] ,

[0136] In the formula: representing the ideal control input vector of the ship in the longitudinal, lateral and bow directions, and ;

[0137] S22: Based on the Stone-Weierstrass theorem, the nominal dynamics model of the ship is rewritten as:

[0138] ,

[0139] In the formula: representing the ideal weight of the RNN base function, , and satisfy ; representing the estimated value of the RNN network weight; representing the error between the ideal weight of the RNN base function and the weight estimate value; representing the reconstruction error, and ; representing the vector composed of the RNN base function, and ; representing the monotonically increasing RNN base function, that is, satisfying: ; representing the base function independent variable; representing a normal number;

[0140] S23: According to the estimated value of the RNN weight , the data-physical dual driving model is determined, and according to the rewritten nominal dynamics model, the approximate dynamics system of the ship is constructed as:

[0141] ,

[0142] ,

[0143] In the formula: Represents the estimated value of the ship data physical dual-drive model and satisfies ; Indicates modeling error; This represents the reconstruction error compensation feedback term; This represents the gain coefficient of the linear feedback term. ; The gain parameter estimate of the nonlinear feedback term and ; The gain parameter represents the ideal nonlinear feedback term; This represents the gain parameter error of the nonlinear feedback term; express The first derivative;

[0144] This embodiment also includes the handling of modeling errors. Differentiating, we get:

[0145] ,

[0146] In the formula: ;

[0147] The derivative of the modeling error Perform stability analysis and design estimates for the RNN weights. Gain parameter estimates of the nonlinear feedback term The update law is used to obtain the gain parameters of the RNN weights and nonlinear feedback terms;

[0148] The expression for the update law of the estimated weights of the RNN is:

[0149] ,

[0150] In the formula: ; This represents the learning rate of the RNN. ;

[0151] The gain parameter estimate of the nonlinear feedback term The expression for the update law is:

[0152] ,

[0153] In the formula: This represents a predefined positive definite matrix;

[0154] S24: constructing a three-degree-of-freedom data-physical double-driven model without considering input constraints but considering unknown external environmental disturbances according to the ship approximate dynamics system and the ship physical model is:

[0155] ;

[0156] S3: defining an ideal control input of the data-physical double-driven model based on the Lyapunov backstepping method, obtaining a track tracking error based on the preset expected ship track according to the data-physical double-driven model, and constructing a ship virtual controller according to the track tracking error; obtaining a ship speed error based on the ship virtual controller to define an intermediate parameter for optimal feedback control design, deriving the ship speed error to obtain a rewritten derivative of the ship speed error according to the intermediate parameter combined with the ideal control input; designing a ship feedforward controller according to the rewritten derivative of the ship speed error; defining a Lyapunov function based on the track tracking error and the ship speed error, and deriving to obtain the derivative of the Lyapunov function;

[0157] Specifically, it includes the following steps:

[0158] S31: defining an ideal control input of the data-physical double-driven model based on the Lyapunov backstepping method,

[0159] The expression of the ideal control input is:

[0160] + ,

[0161] In the formula: represents the output of the feedforward controller; represents the output of the optimal feedback controller; represents the output of the worst disturbance strategy;

[0162] S32: obtaining a track tracking error of the ship according to the defined data-physical double-driven model is:

[0163] ,

[0164] In the formula: represents the track tracking error; represents the track reference signal and ;

[0165] S33: deriving the track tracking error to obtain the track tracking error derivative, and rewriting the track tracking error derivative as:

[0166] ,

[0167] In the formula: represents the first derivative of ;

[0168] S34: Design the virtual ship controller according to the rewritten track following error derivative as:

[0169]

[0170] wherein: represents the control design parameter and ; represents the output of the virtual ship controller;

[0171] S35: Obtain the ship speed error according to the virtual ship controller and the data physical double drive model to define the intermediate variable for optimal feedback control design, and the expression is:

[0172]

[0173]

[0174]

[0175] wherein: represents the ship speed error; represents the intermediate variable;

[0176] S36: Derive the ship speed error to obtain the ship speed error derivative, and rewrite the ship speed error derivative according to the intermediate variable combined with the ideal control input;

[0177]

[0178] wherein: represents the first derivative of ; represents the first derivative of ;

[0179] S37: Obtain the rewritten ship speed error derivative, and perform stability analysis on the rewritten ship speed error derivative to design the ship feedforward controller as:

[0180]

[0181] wherein: represents the control design parameter and ; represents the output of the ship feedforward controller;

[0182] S38: To prove the stability analysis in step S36, define the Lyapunov function for the track following error and the ship speed error as: ​​​​​​

[0183] ,

[0184] and the derivative of Lyapunov function is obtained by deriving the above track tracking error, ship speed error and virtual controller, feedforward controller is:

[0185] ;

[0186] S4: define the error system based on the derivative of Lyapunov function; construct the performance index function based on the error system combined with the zero-sum game theory; construct the Hamilton-Jacobi-Isaacs equation according to the performance index function to obtain the optimal feedback controller and the worst disturbance strategy;

[0187] Specifically, it includes the following steps:

[0188] S41: define the error system as follows according to the differential game theory combined with the last term in the derivative of Lyapunov function:

[0189] ,

[0190] S42: define the performance index function under the zero-sum game as follows based on the error system combined with the zero-sum game theory:

[0191] ,

[0192] In the formula: denotes the time parameter; denotes the discount factor; denotes the allowable control input; , , denotes a positive definite gain matrix and , , ;

[0193] S43: according to the Nash-Pontryagin maximum minimum value principle, confirm that the zero-sum game has a unique solution The condition is that there is a unique saddle point, so that the optimal performance index function is:

[0194] ,

[0195] S44: define the Hamilton function of the error system according to the performance index function, and its expression is:

[0196] ,

[0197] In the formula: denotes the partial derivative of with respect to denotes the transpose of

[0198] Based on the Hamilton function according to the optimal performance index function , the Hamilton-Jacobi-Isaacs equation (HJI) is obtained as:

[0199]

[0200] In the formula: denotes the partial derivative of with respect to

[0201] S45: According to the performance index function, the Hamilton-Jacobi-Isaacs equation is constructed, and the gradient descent method is used to obtain the optimal feedback controller and the worst disturbance strategy as:

[0202] ,

[0203] ,

[0204] S5: The optimal performance index function is approximated by using the fuzzy logic system to obtain the output estimation value of the optimal feedback controller and the worst disturbance strategy. In this embodiment, based on the above differential game framework, the fuzzy logic system is used to approximate the solution of the HJI equation to obtain the Nash equilibrium of the output estimation value of the optimal feedback controller and the worst disturbance strategy. The neural network weight update rate of the fuzzy logic system is constructed according to the output estimation value of the optimal feedback controller and the worst disturbance strategy.

[0205] Specifically, the steps include:

[0206] S51: The optimal performance index function is approximated by using the fuzzy logic system, and its expression is:

[0207] ,

[0208] In the formula: denotes the ideal weight and satisfies ; denotes the error of the neural network weight; denotes the estimation value of ; denotes the fuzzy basis function; denotes the approximation error;

[0209] S52: According to S51, the partial derivative of the approximated optimal performance index function with respect to is:

[0210] ​ ,

[0211] wherein: represents the partial derivative with respect to ; represents the partial derivative with respect to ;

[0212] S53: according to the weight of the estimated value of the partial derivative, the evaluation network is determined, and the estimated value of the partial derivative of the optimal performance index function is obtained according to S52

[0213] ,

[0214] S54: according to the partial derivative described in S52, the optimal feedback controller in S45 and the worst disturbance strategy and the Hamilton-Jacobi-Essix equation in S44 are substituted to obtain:

[0215] ,

[0216] ,

[0217] ,

[0218] wherein: represents the reconstruction residual and ; represents the intermediate variable and ; represents a short form of

[0219] Similarly, the estimated value of the partial derivative obtained in S53 is substituted into the optimal feedback controller in S45 and the worst disturbance strategy and the Hamilton-Jacobi-Essix equation in S54 to obtain the estimated value of the output of the optimal feedback controller, the estimated value of the output of the worst disturbance strategy, and the rewritten Hamilton-Jacobi-Essix equation as:

[0220] ,

[0221] ,

[0222] ,

[0223] S55: the estimated values of the outputs of the optimal feedback controller and the worst disturbance strategy are substituted into the error system in S41 to obtain the rewritten error system as:

[0224] ​ ,

[0225] S56: Construct the neural network weight update rate of the fuzzy logic system based on the rewritten error system, using the gradient descent method to minimize the rewritten Hamilton-Jacobi-Essacs equation ;

[0226] ,

[0227] In the formula: indicates the weight learning rate and ;

[0228] S6: Based on the neural network weight update rate, the output estimation value of the ship feedforward controller combined with the optimal feedback controller and the worst disturbance strategy is constructed to build the ship final ideal controller; the control input of the ship final ideal controller is constrained through the input constraint in the ship physical model to obtain the actual control input, and the ship zero-sum game tracking control based on data-physical dual-driven reinforcement learning is realized according to the actual control input. This embodiment fuses the feedforward controller, the optimal feedback controller and the worst disturbance strategy, and obtains the actual control instruction that can be directly executed through constraint processing, so as to realize ship track tracking, which specifically includes the following steps:

[0229] S61: Based on the neural network weight update rate, the output estimation value of the ship feedforward controller combined with the optimal feedback controller and the worst disturbance strategy is constructed to build the ship final ideal controller as follows:

[0230] ,

[0231] In the formula: indicates the RNN weight estimation value, which is obtained by processing the ship speed data collected by the Doppler log; indicates the track tracking error and indicates the speed error of each degree of freedom; , , , indicates the designed parameter; indicates the fuzzy basis function; indicates the coordinate conversion matrix;

[0232] S62: The control input of the ship final ideal controller is constrained through the input constraint in the ship physical model to obtain the actual control input; the ship zero-sum game tracking control based on data-physical dual-driven reinforcement learning is realized according to the actual control input, that is, the ship zero-sum game track tracking control under the condition that the random disturbance of the marine environment, the physical constraint and the dynamic characteristics of the ship are completely unknown.

[0233] The embodiment also includes step S7: using Lyapunov theory, proving that the designed ship zero-sum game tracking control method based on data-physical dual-driven reinforcement learning is stable in input to state, and all signals in the closed-loop system are ultimately uniformly bounded.

[0234] Specifically, in order to more smoothly prove the stability, the following assumptions are made here:

[0235] Assumption 1: the RNN reconstruction error is bounded, and the upper and lower bounds are functions of the modeling error , that is, the expression is .

[0236] Assumption 2: on the compact set , , , and are bounded, satisfying , , , , and and , where , , , , and are positive parameters.

[0237] Theorem 1: For the ship track tracking control system considering unknown wind and wave flow disturbance, input constraint and unknown dynamics, based on Lyapunov theory, under the action of the designed ship data-physical dual-driven model, RNN weight adaptive law, feedforward controller, output estimation value of optimal feedback controller, output estimation value of worst disturbance strategy and fuzzy logic system weight adaptive law, by selecting appropriate parameters, it can be guaranteed that all signals of the closed-loop system are ultimately uniformly bounded, and the tracking error converges in the optimal way.

[0238] Lemma 1: (Young’s inequality), for any , the following inequality holds:

[0239] ,

[0240] In the formula: , , and ;

[0241] Proof of Theorem 1:

[0242] S71: To prove that the ship data physics dual-driven model and the RNN weights are ultimately consistent and bounded, a Lyapunov function is constructed , and the Lyapunov function expression is:

[0243] ,

[0244] S72: Derive the Lyapunov function , and substitute the modeling error derivative , the RNN weight error derivative , and the gain parameter of the nonlinear feedback term , to obtain:

[0245] ,

[0246] Combining the monotonic increasing property of the RNN base function in step S22, assumption 1, and lemma 1 (Young's inequality), the following inequality can be obtained:

[0247] ,

[0248] Substitute the Lyapunov function derivative , to obtain:

[0249] ,

[0250] In the formula: represents the identity matrix of ;

[0251] Therefore, when a suitable value is selected, such that , the ship data physics dual-driven model and the RNN weights are ultimately consistent and bounded;

[0252] S73: Next, it is proved that the data physics dual-driven ship error system and the fuzzy logic system weights are ultimately consistent and bounded, according to the Hamilton function in step S54 , the following equation can be obtained:

[0253] ,

[0254] Therefore, substitute the equation of S73 into the fuzzy logic system weight adaptive law in step S55 , to obtain:

[0255] ,

[0256] S74: Based on the Lyapunov function of the feedforward control design in step S38 , a new Lyapunov function , the expression is:

[0257] ,

[0258] Taking the derivative of the above Lyapunov function , and substituting the relevant equations in steps S38, S41 and S73 into the derivative of the Lyapunov function , we can obtain:

[0259] ,

[0260] S75: According to assumption 2, take one item as an example to illustrate the scaling process

[0261] ,

[0262] The remaining items are processed in the same way, and we can obtain

[0263]

[0264] In the formula: ; ; ;

[0265] ; ,

[0266] Therefore, when is satisfied, all signals of the data-physical dual-driven ship track tracking control system can be ultimately consistent and bounded by selecting appropriate parameters.

[0267] Compared with the existing technology, the method described in the embodiment has the following advantages:

[0268] 1. The method described in the embodiment fuses the prior physical law contained in the ship physical model with the ship speed data collected by the sensor to construct a ship data-physical dual-driven model. This model not only retains the explainability and stability of the physical model, but also effectively avoids the problem of dependence on high-quality data in the data-driven method. Without accurate ship parameters, efficient and stable control can be achieved, and the engineering application threshold and implementation cost of the technology are greatly reduced.

[0269] 2. The method described in the embodiment constructs a dynamic confrontation mechanism of the controller and the disturbance through a zero-sum game framework, combines reinforcement learning and fuzzy logic system, realizes optimal control solution under the worst disturbance strategy, and significantly improves the adaptive ability of the ship to complex disturbances such as wind, waves and currents.

[0270] 3. The method described in this embodiment was simulated using the MATLAB platform. A zero-sum game optimal controller was designed to achieve ship track tracking in unknown and complex marine environments. The system controller is more in line with navigation practice, further verifying the effectiveness and rationality of the method described in this embodiment.

[0271] To verify the effectiveness of the method described in this embodiment, a computer simulation study was conducted using MATLAB. The parameter settings are as follows:

[0272] The simulation object selected is the CyberShipII from the Norwegian University of Science and Technology, with a weight of 23.8 kg, a length of 1.255 m, and a beam of 0.29 m. Parameters were set when designing the data-physics dual-drive model. , , , , , RNN weight initial values For the corresponding dimension, the zero matrix, basis functions Initial values ​​of feedback parameters When designing an ideal controller, the design parameters are... , , , , , The initial value of the weights in the fuzzy logic system is The basis functions are Center point selection The reference signal for the ship is set as follows: The initial state of the ship's control system is... .

[0273] The simulation results are shown in the figure: Figure 2 This is a schematic diagram of the error in the dual-driven model of the actual model and the data physics model. It can be seen that the error has successfully converged to the bounded neighborhood centered on the origin. Figure 3 For position coordinates The trajectory tracking effect on the plane shows that the output position signal can quickly track the reference signal; Figure 4 Then from position and heading The tracking effect is intuitively displayed in three dimensions; Figure 5 The figure shows the convergence curve of the weight norm of the fuzzy logic system. The results show that the weight norm can converge stably to a certain fixed value. Figure 6 The figure shows the convergence process of the RNN weight norm and feedback parameters. The weight norm and parameters can converge to stable values. Figure 7The change curve of the actual control input of the ship is a relatively smooth control signal. Figure 8 The change curve of the worst interference strategy. The simulation results show that the effectiveness of the method described in the embodiment is verified, and the zero-sum game controller designed based on the control system can well realize the ship track tracking, and the tracking error can quickly converge to a small residual set, which fully shows that the output of the data-physical double-driven ship tracking system has good tracking performance, further verifying the effectiveness and rationality of the method described in the embodiment.

[0274] In summary, the method described in the embodiment establishes a three-degree-of-freedom ship physical model considering input constraints and unknown wind and wave disturbance, unknown dynamics; through the Doppler log to collect key navigation data such as ship speed and attitude, combined with the data learning ability of recurrent neural network, a data-physical double-driven model with physical interpretability and data adaptive compensation ability is constructed; based on the model, the Lyapunov backstepping method is used to design the ship's feedforward controller in layers, to realize accurate tracking of the ship's nominal dynamics; then through the zero-sum game differential theory and adaptive dynamic programming based on fuzzy logic system, the optimal feedback controller and the worst interference strategy are optimized simultaneously, and a dynamic confrontation framework of the controller and the disturbance is constructed; finally, the ideal control input is obtained by fusing the feedforward control, the optimal feedback control and the worst interference strategy, and the directly executable actual control input is obtained after physical constraint, so as to realize the track tracking control of the ship under the condition of random disturbance of the marine environment, physical constraints and completely unknown dynamic characteristics of the ship. The method described in the embodiment breaks through the limitations of traditional single modeling method through deep integration of data and physics, and can realize efficient and stable control without relying on accurate ship parameters; in addition, with the help of the organic combination of zero-sum game and reinforcement learning, the adaptive ability and anti-interference performance of the ship to complex disturbance are significantly improved, which promotes the intelligent upgrading of the ship tracking control system, and provides a reliable technical solution for the accurate tracking control of the ship in offshore operation, ocean transportation and other multiple scenarios.

[0275] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A ship zero-sum game tracking control method based on dual-driven reinforcement learning of numbers and objects, characterized in that, The specific steps include: S1: Obtain a ship physical model that takes into account unknown wind, wave and current disturbances, input constraints and unknown dynamics; S2: Construct a data physics dual-drive model based on the ship physics model and a recurrent neural network; S3: Based on the Lyapunov backstepping method, define the ideal control input of the data-physical dual-drive model. Based on the preset desired ship trajectory, obtain the trajectory tracking error according to the data-physical dual-drive model, and construct a ship virtual controller based on the trajectory tracking error. Obtain the ship speed error from the ship virtual controller to define intermediate parameters for optimal feedback control design. After differentiating the ship speed error, obtain the rewritten ship speed error derivative based on the intermediate parameters and the ideal control input. Design a ship feedforward controller based on the rewritten ship speed error derivative. Define the Lyapunov function based on the trajectory tracking error and the ship speed error, and obtain the derivative of the Lyapunov function by differentiation. S4: Define the error system based on the derivative of the Lyapunov function; construct a performance index function based on the error system and zero-sum game theory; construct the Hamilton-Jacobi-Isax equation based on the performance index function to obtain the optimal feedback controller and the worst disturbance strategy. S5: Use the fuzzy logic system to approximate the optimal performance index function to obtain the output estimates of the optimal feedback controller and the worst interference strategy, and construct the neural network weight update rate of the fuzzy logic system based on the output estimates of the optimal feedback controller and the worst interference strategy. S6: Based on the neural network weight update rate, construct the ship's final ideal controller according to the output estimate of the ship's feedforward controller combined with the optimal feedback controller and the worst disturbance strategy; By constraining the control input of the ship's final ideal controller through input constraints in the ship's physical model, the actual control input is obtained, and ship zero-sum game tracking control based on data physics dual-driven reinforcement learning is realized based on the actual control input.

2. The ship zero-sum game tracking control method based on dual-driven reinforcement learning according to claim 1, characterized in that, The expression for the ship's physical model described in S1 is: In the formula: Indicates the ship in the inertial coordinate system coordinate, Coordinates and bow roll angle The position vector formed and ; Represents the forward speed of the ship in the attached coordinate system. Horizontal drift speed and bow roll rate The velocity vector formed and ; Denotes the coordinate system transformation matrix and ; This indicates the actual forward control force that the ship can directly execute. Lateral drift force and bow roll torque The control vector formed and ; This represents the lateral disturbance force on a ship caused by wind, waves, and currents in the attached coordinate system. Longitudinal interference force and bow disturbance torque The perturbation vector formed and ,satisfy ; express The upper bound of a positive constant; The matrix represents the ship's weight inertia and hydrodynamic additional inertia, and ; Describe the Coriolis matrix and ; Represents the hydrodynamic damping parameter matrix and ; Indicates transpose; Represents the space of real numbers; express The first derivative; The input constraints are defined as follows: In the formula: Represents the control forces of each degree of freedom in the ship's physical model and ; Indicates the ideal control inputs for the ship's longitudinal, lateral, and bow directions; This indicates the maximum value of the control input for each degree of freedom of the ship; This represents the minimum control input value for each degree of freedom of the ship.

3. The ship zero-sum game tracking control method based on dual-driven reinforcement learning according to claim 2, characterized in that, S2 specifically includes the following steps: S21: The nominal dynamic model obtained from the ship's physical model, without considering input constraints, is as follows: In the formula: Represents the ideal control input vectors for the ship's longitudinal, lateral, and bow directions, and ; S22: Based on the Stone-Weilstras theorem, the nominal dynamic model of the ship is rewritten using RNN basis functions as follows: In the formula: Describes the ideal weights of the RNN basis functions and satisfies ; This represents the estimated values ​​of the RNN network weights; This represents the error between the ideal weights and the estimated weights of the RNN basis functions; Indicates reconstruction error and ; Describes a vector composed of RNN basis functions and ; Represents the monotonically increasing basis functions of an RNN; S23: Based on the revised nominal dynamics model, construct the approximate dynamics system of the ship as follows: In the formula: Represents the estimated value of the ship data physical dual-drive model and satisfies ; Indicates modeling error; This represents the reconstruction error compensation feedback term; This represents the gain coefficient of the linear feedback term. ; The gain parameter estimate of the nonlinear feedback term and ; The gain parameter represents the ideal nonlinear feedback term; This represents the gain parameter error of the nonlinear feedback term; express The first derivative; S24: Based on the ship's approximate dynamics system and the aforementioned ship physical model, a data-physical dual-drive model is constructed as follows: 。 4. The ship zero-sum game tracking control method based on dual-driven reinforcement learning according to claim 3, characterized in that, S3 specifically includes the following steps: S31: Based on Lyapunov's backstepping method, define the ideal control input for the data-physical dual-drive model. The expression for the ideal control input is: + In the formula: Express the output of the feedforward controller; Express the output of the optimal feedback controller; This represents the output of the worst-case interference strategy; S32: Based on the defined data-physical dual-drive model, the ship's trajectory tracking error is: In the formula: Indicates the tracking error; Indicates the track reference signal and ; S33: Obtain the derivative of the trajectory tracking error by taking the derivative of the trajectory tracking error, and rewrite the derivative of the trajectory tracking error according to the data-physics dual-drive model as follows: In the formula: express The first derivative; S34: Based on the rewritten derivative of the trajectory tracking error, the ship virtual controller is designed as follows: In the formula: Indicates the control design parameters and ; This represents the output of the ship's virtual controller; S35: Obtain the ship speed error based on the ship virtual controller and the data physics dual-drive model to define intermediate parameters for optimal feedback control design. The expression for these parameters is: , In the formula: Indicates the error in ship speed; Indicates intermediate variables; S36: Obtain the derivative of the ship speed error by taking the derivative of the ship speed error, and obtain the rewritten derivative of the ship speed error based on the intermediate parameters and the ideal control input; In the formula: express The first derivative; express The first derivative; S37: Based on the rewritten derivative of the ship speed error, the ship feedforward controller is designed as follows: In the formula: Indicates the control design parameters and ; This indicates the output of the ship's feedforward controller; S38: Define the Lyapunov function based on track tracking error and ship speed error. for: And for Lyapunov functions Find the derivative of the Lyapunov function. for: 。 5. The ship zero-sum game tracking control method based on dual-driven reinforcement learning according to claim 4, characterized in that, S4 specifically includes the following steps: S41: The derivative of the Lyapunov function... The term is defined as an error system: S42: Based on the error system and zero-sum game theory, the performance index function under zero-sum game is defined as follows: In the formula: Indicates time parameters; Indicates the discount factor; Indicates the permissible control inputs; , , Denotes the positive definite gain matrix and , , ; S43: Based on the Nash-Pontryagin maximum-minimum principle, it is confirmed that the zero-sum game has a unique solution. The condition is that there exists a unique saddle point such that the optimal performance index function... for: S44: Based on the performance index function, define the Hamiltonian function of the error system, whose expression is: In the formula: express about The partial derivative; express transpose; Based on the Hamiltonian function and the optimal performance index function The Hamilton-Jacobi-Isax equations are obtained as follows: In the formula: express about The partial derivative; S45: Construct the Hamilton-Jacobi-Isax equation based on the performance index function, and use the gradient descent method to obtain the optimal feedback controller and the worst-case disturbance strategy: , 。 6. The ship zero-sum game tracking control method based on dual-drive reinforcement learning of numbers and objects according to claim 5, characterized in that, S5 specifically includes the following steps: S51: The optimal performance index function is approximated using a fuzzy logic system, and its expression is: In the formula: Represents the ideal weights and satisfies ; This represents the error in the weights of the neural network; express The estimated value; Represents fuzzy basis functions; Indicates the approximation error; S52: Obtain the approximate optimal performance index function based on S51. partial derivatives for: In the formula: express about The partial derivative; express about The partial derivative; S53: Based on weight The estimated value Determine the evaluation network and obtain the partial derivative estimate of the optimal performance index function based on S52. for: S54: Partial derivative as described in S52 Substituting this into the optimal feedback controller and worst-case disturbance strategy in S45 and the Hamilton-Jacobi-Isax equation in S44, we obtain: In the formula: Indicates the reconstructed residual and ; Indicates intermediate parameters and ; express The abbreviated form; Similarly, the partial derivative estimates obtained from S53 are... Substituting the optimal feedback controller and worst-case disturbance policy from S45 and the Hamilton-Jacobi-Isax equation from S54, we obtain the output estimate of the optimal feedback controller. The output estimate of the worst-case interference strategy And the rewritten Hamilton-Jacobi-Isaks equation is: S55: Substitute the output estimates of the optimal feedback controller and the worst-case disturbance strategy into the error system in S41 to obtain the rewritten error system: S56: Based on the rewritten error system, use gradient descent to construct the neural network weight update rate of the corresponding fuzzy logic system that minimizes the rewritten Hamiltonian-Jacobi-Isax equation. for; In the formula: Indicates the learning rate of the weights and .

7. The ship zero-sum game tracking control method based on dual-drive reinforcement learning of numbers and objects according to claim 6, characterized in that, S6 specifically includes the following steps: S61: Based on the neural network weight update rate, and according to the output estimates of the ship's feedforward controller combined with the optimal feedback controller and the worst-case disturbance strategy, the final ideal controller of the ship is constructed as follows: S62: Constrain the control input of the ship's final ideal controller by the input constraints in the ship's physical model to obtain the actual control input; realize zero-sum game tracking control of the ship based on data physics dual-drive reinforcement learning according to the actual control input.

Citation Information

Patent Citations

  • Unmanned surface ship optimal trajectory tracking control method based on reinforced learning method

    CN110018687A

  • Unmanned ship course tracking control method based on execution comment system reinforcement learning

    CN118466220A