Knowledge-driven reinforcement learning heading tracking control method for unmanned ships

Through a knowledge-driven reinforcement learning method, a switching controller suitable for different sea conditions is built, which solves the problem of poor tracking capabilities of traditional unmanned ships in complex environments, and achieves higher accuracy and reliability of heading tracking control.

CN119512205BActive Publication Date: 2025-05-13DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510089695.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

When traditional unmanned ships’ heading tracking control methods face different complex marine environments and unknown disturbances, it is difficult to maintain good heading tracking control effects, resulting in poor tracking capabilities.

Method used

Using a knowledge-driven reinforcement learning method, heading tracking control is realized by obtaining a mathematical model of unmanned ship heading control with uncertain interference terms and converting it into a second-order state space equation, a switching controller suitable for complex and general sea conditions is constructed to realize heading tracking control.

Benefits of technology

Effectively adapting to different sea conditions, improving the accuracy and reliability of unmanned ships' heading tracking control, and solving the problem of poor tracking capabilities of traditional methods in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119512205B_ABST
    Figure CN119512205B_ABST
Patent Text Reader

Abstract

The present invention discloses a knowledge-driven unmanned ship reinforcement learning heading tracking control method, including obtaining an unmanned ship virtual control law according to an unmanned ship heading tracking control model; obtaining a feedforward item and a feedback item of an unmanned ship heading tracking controller based on the unmanned ship virtual control law; constructing a feedforward optimization controller about the feedforward item; constructing a first switching controller suitable for unmanned ship heading tracking control under complex sea conditions according to the feedback item; using a backstepping method combined with an RBF neural network, constructing a second switching controller for coping with general sea conditions according to the feedback item; using a step function, designing a switching response controller for different sea conditions according to the first switching controller and the second switching controller, so as to realize knowledge-driven reinforcement learning unmanned ship heading tracking control. The method solves the problem that the traditional unmanned ship heading tracking control method is difficult to maintain a good heading tracking control effect when facing different complex marine environments and unknown disturbances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned ships, and in particular to a knowledge-driven unmanned ship reinforcement learning heading tracking control method. Background Art

[0002] In recent years, unmanned ships, as an emerging intelligent means of transportation, have received widespread attention because they can improve marine operation efficiency, reduce human operating errors and reduce labor costs.

[0003] The navigation control of unmanned ships, especially the heading tracking control, is one of its key technologies. Effective heading tracking control technology not only affects the navigation stability and safety of unmanned ships, but is also directly related to the accuracy and reliability of its mission execution. With the development of machine learning and artificial intelligence technology, some emerging technologies have been widely used in unmanned ship heading tracking control. Traditional unmanned ship heading tracking control methods such as PID control, fuzzy control and optimal control often fail to take into account the influence of disturbances in different complex marine environments and unknown disturbances when designing, resulting in unsatisfactory path tracking effects, which in turn leads to the inability to adapt well to different sea conditions, resulting in poor tracking capabilities of the ship heading tracking control system and difficulty in maintaining good heading tracking control effects. Summary of the invention

[0004] The present invention provides a knowledge-driven unmanned ship reinforcement learning heading tracking control method to overcome the above technical problems.

[0005] In order to achieve the above object, the technical solution of the present invention is:

[0006] A knowledge-driven unmanned ship reinforcement learning heading tracking control method specifically includes the following steps:

[0007] S1: Obtain a mathematical model for unmanned ship heading control with uncertain interference terms, considering that the unmanned ship sailing on the sea will be affected by environmental interference;

[0008] The unmanned ship heading control mathematical model is converted into a second-order state space equation applicable to different sea conditions to serve as the unmanned ship heading tracking control model;

[0009] S2: Obtain the unmanned ship virtual control law according to the unmanned ship heading tracking control model;

[0010] S3: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the feedforward term and feedback term of the unmanned ship heading tracking controller are obtained;

[0011] And construct a feedforward optimization controller with respect to the feedforward term;

[0012] S4: Based on the reinforcement learning method, a heading tracking optimization controller suitable for the heading tracking control of the unmanned ship under complex sea conditions is constructed according to the feedback item, and a first switching controller for coping with complex sea conditions is obtained according to the feedforward optimization controller and the heading tracking optimization controller;

[0013] S5: Using the backstepping method combined with the RBF neural network, a second switching controller for coping with general sea conditions is constructed according to the feedback term;

[0014] S6: Using a step function, a switching response controller for different sea conditions is designed according to the first switching controller and the second switching controller to realize knowledge-driven reinforcement learning unmanned ship heading tracking control.

[0015] Furthermore, the S1 specifically includes the following steps:

[0016] S11: Considering that unmanned ships sailing on the sea will be affected by various interferences, the mathematical model of unmanned ship heading tracking control with uncertain interference terms is expressed as

[0017] ,

[0018] Where: Indicates the heading angle of the unmanned ship; Indicates the actual rudder angle of the unmanned ship; represents the equivalent rudder angle caused by uncertain environmental disturbance; It represents the maneuverability index of the unmanned ship rudder gain; The time constant expressed as the ship's maneuverability; All are expressed as nonlinear coefficients of the ship's bow angular velocity; express The first derivative of is the ship’s bow angular velocity; express The second derivative of is the ship's bow angular acceleration;

[0019] S12: Order , , , and satisfy , the unmanned ship heading control mathematical model is converted into a second-order state space equation and used as the heading tracking control mathematical model. The expression of the heading tracking control mathematical model is:

[0020] ,

[0021] Where: Indicates uncertainty about environmental interference and , This means that there is interference from complex sea conditions; Indicates general sea environment disturbance; Indicates uncertainty about environmental interference The upper bound of , express A subset of and is a compact set; represents the set of all real numbers; represents all real vectors n Vieux-style space; represents control input; Indicates the heading of the unmanned ship; represents the bow angular velocity of the unmanned ship; Indicates the maneuverability index of the unmanned ship rudder gain Time constant with ship maneuverability The intermediate parameter quantity;

[0022] Furthermore, the S2 specifically includes the following steps:

[0023] S21: defining a model error according to the unmanned ship heading tracking control model;

[0024] And the model error includes the heading tracking error of the unmanned ship Steering angular velocity tracking error ;

[0025] The expression of the model error is:

[0026] ,

[0027] ,

[0028] Where: A reference signal indicating the heading of the unmanned ship; represents the virtual control law of the unmanned ship to be designed; represents the steering angular velocity tracking error and , It represents the steering angular velocity tracking error under the interference of complex sea conditions; It represents the steering angular velocity tracking error under the general sea condition interference; Represents the heading tracking error of the unmanned ship;

[0029] S22: Constructing Lyapunov function based on heading tracking error , whose expression is ,

[0030] And for the Lyapunov function Derivative, get the derivative of the Lyapunov function :

[0031] ;

[0032] S23: Construct derivatives to ensure Lyapunov functions A virtual control law that is less than or equal to 0, and the expression of the virtual control law is ,

[0033] Where: Represents the design parameters and satisfies .

[0034] Furthermore, the S3 specifically includes the following steps:

[0035] S31: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is:

[0036] ,

[0037] Where: Indicates a switching signal; It represents the control input of the switching controller in response to complex sea conditions; Indicates complex sea conditions interference;

[0038] S32: Switching controller to cope with complex sea conditions Defined as ,

[0039] Where: represents the feedforward term of the switching controller; represents the feedback term of the switching controller;

[0040] S33: Based on step S32, the steering angular velocity tracking error after derivation is rewritten, and its expression is:

[0041] ,

[0042] And based on the feedforward term in the rewritten derivative steering angular velocity tracking error, a feedforward optimization controller is constructed, and its expression is:

[0043] ,

[0044] Where: Represents the design parameters and satisfies .

[0045] Furthermore, the S4 specifically includes the following steps:

[0046] S41: Unknown terms in the mathematical model of unmanned ship heading tracking control With ship parameters , reconstruct the feedback item of the switching controller in step S33 to obtain a reconstructed feedback item;

[0047] The expression of the reconstructed feedback term is:

[0048] ,

[0049] Where: represents the reconstruction feedback term and ; express The first derivative of Represents the unknown term in the unmanned ship heading tracking control model ;

[0050] S42: Design a first cost function based on the reconstruction feedback term, and its expression is:

[0051] ,

[0052] Where: represents the first cost function; express The integrand of ; express The abbreviation of express The abbreviation of Indicates the switching signal for dealing with complex sea conditions The optimal controller under represents the positive definite term about the partial derivative of the first cost function;

[0053] S43: Solve and obtain the first cost function for the reconstruction feedback term The partial derivative of is substituted into the constructed HJB equation, which is expressed as

[0054] ,

[0055] ,

[0056] Where: Denotes the first cost function with respect to the reconstruction feedback term The partial derivative of represents the design constant, express The abbreviation of ; express The abbreviation of represents a positive definite constant;

[0057] S44: Using the optimal condition equation Solve the constructed HJB equation to obtain the switching signal in response to complex sea conditions The optimal controller under ;

[0058] S45: Use a single critic neural network to approximate the first cost function, and based on the approximated first cost function, obtain its partial derivative, which is expressed as

[0059] ,

[0060] ,

[0061] Where: Represents the weights of the neural network; Represents the activation function of the neural network; represents the approximation error of the neural network; express The partial derivative of express The partial derivative of

[0062] S46: Obtain partial derivatives based on step S45 Estimated value of , to obtain the optimal controller for a single ciritc network controller , whose expression is

[0063] ,

[0064] ,

[0065] ,

[0066] Where: express The expected value and ; Indicates expected value The error between the true value and

[0067] S47: Substituting the output of the single ciritc network controller into step S41, we can get

[0068] ,

[0069] S48: The The same as in step S46 Substituting into step S43 and combining with step S47, we can get

[0070] ,

[0071] And further simplifying it can get

[0072] ,

[0073] Where: represents the intermediate parameter quantity and

[0074] ;

[0075] S49: The and Substituting into step S43 and combining with step S47, we can get

[0076] ;

[0077] S410: According to steps S48 and S49, the variable error can be obtained , whose expression is

[0078] ,

[0079] Where: , represents the positive definite matrix about the neural network basis function;

[0080] S411: Constructing the first cost function The Lyapunov function of

[0081] ,

[0082] Combined with variable error Construct the adaptive rate of the heading tracking optimization controller, which is expressed as

[0083] ,

[0084] ,

[0085] ,

[0086] Where: , represents the design parameters;

[0087] S412: combining the adaptive rate with the single network controller of the optimal controller in step S46 , we can obtain the heading tracking optimization controller suitable for unmanned ship heading tracking control in complex sea conditions. ,

[0088] ,

[0089] S413: The step S412 As a switching controller in step S32 to deal with complex sea conditions of , and combined with the feedforward optimization controller, the first switching controller for dealing with complex sea conditions is obtained.

[0090] Furthermore, the S5 specifically includes the following steps:

[0091] S51: Based on the backstepping method and the unmanned ship virtual control law, combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is:

[0092] ,

[0093] Where: Indicates the switching signal corresponding to general sea conditions; Indicates the switching controller for dealing with general sea conditions; Indicates general sea disturbance;

[0094] S52: Transform the formula of step S51, and the transformation formula is:

[0095] ,

[0096] And use RBF neural network to approximate the transformation formula , can get

[0097] ,

[0098] S53: Based on step S52, a second switching controller for coping with general sea conditions is constructed, and its expression is:

[0099] ,

[0100] Where: express The expected value and ; Indicates expected value With the true value The error between Represents the weight of the RBF neural network, and the adaptive rate of the weight of the RBF neural network is , and Both represent adjustment parameters and satisfy , ; Represents the basis function of the RBF neural network; Represents design parameters.

[0101] Furthermore, in S6, a step function is used to design a switching response controller for different sea conditions according to the first switching controller and the second switching controller, and its expression is:

[0102] ,

[0103] Where: express The step function of time.

[0104] Beneficial effect: The present invention provides a knowledge-driven unmanned ship reinforcement learning heading tracking control method, which considers the unmanned ship sailing on the sea and is affected by different wave environment interferences to establish an unmanned ship heading control mathematical model with uncertain interference terms; and converts the unmanned ship heading control mathematical model into a second-order state space equation as the unmanned ship heading tracking control model; and obtains the unmanned ship virtual control law according to the unmanned ship heading tracking control model to construct a first switching controller for coping with complex sea conditions; and a second switching controller for coping with general sea conditions, and then uses a step function to design a switching response controller for different sea conditions according to the first switching controller and the second switching controller, and switches the two controller systems of the switching response controller under different sea conditions to achieve better adaptation to different sea conditions, thereby solving the problem that it is difficult to maintain a good heading tracking control effect when facing complex sea environments and unknown disturbances, and finally realizing knowledge-driven unmanned ship reinforcement learning heading tracking control. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0106] Figure 1 It is a flow chart of the knowledge-driven unmanned ship heading tracking control method of the present invention;

[0107] Figure 2 A simulation diagram of the first tracking result of the ship heading angle to the reference heading provided in this embodiment;

[0108] Figure 3 A simulation diagram of a second tracking result of the ship heading angle to the reference heading provided in this embodiment;

[0109] Figure 4 A first tracking error simulation diagram of the ship heading angle and the reference heading provided in this embodiment;

[0110] Figure 5 A second tracking error simulation diagram of the ship heading angle and the reference heading provided in this embodiment;

[0111] Figure 6 It is a simulation diagram of the convergence results of the neural network neuron weights under the general sea conditions provided in this embodiment;

[0112] Figure 7 This is a simulation diagram of the convergence results of the neural network neuron weights under complex sea conditions provided in this embodiment;

[0113] Figure 8 A first control effect simulation diagram of the switching response controller provided in this embodiment;

[0114] Fig. 9 A second control effect simulation diagram of the switching response controller provided in this embodiment;

[0115] Fig.10 This is a simulation diagram of the first sea condition effect when the unmanned ship in this embodiment is affected by environmental interference;

[0116] Fig.11 This is a simulation diagram of the second sea condition effect when the unmanned ship in this embodiment is affected by environmental interference. DETAILED DESCRIPTION

[0117] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0118] This embodiment provides a knowledge-driven unmanned ship reinforcement learning heading tracking control method, such as Figure 1 As shown, the specific steps include:

[0119] S1: Obtain a mathematical model for unmanned ship heading control with uncertain interference terms, considering that the unmanned ship sailing on the sea will be affected by environmental interference;

[0120] The unmanned ship heading control mathematical model is converted into a second-order state space equation applied to different sea conditions as an unmanned ship heading tracking control model, which specifically includes the following steps:

[0121] S11: Considering that unmanned ships sailing on the sea will be affected by various interferences, the mathematical model of unmanned ship heading tracking control with uncertain interference terms is expressed as

[0122] ,

[0123] Where: Indicates the heading angle of the unmanned ship; Indicates the actual rudder angle of the unmanned ship; represents the equivalent rudder angle caused by uncertain environmental disturbance; It represents the maneuverability index of the unmanned ship rudder gain; It is expressed as the time constant of the ship's maneuverability; All are expressed as nonlinear coefficients of the ship's bow angular velocity; express The first derivative of is the ship’s bow angular velocity; express The second derivative of is the ship's bow angular acceleration;

[0124] S12: Order , , , and satisfy , the unmanned ship heading control mathematical model is converted into a second-order state space equation and used as the heading tracking control mathematical model. The expression of the heading tracking control mathematical model is:

[0125] ,

[0126] Where: Indicates uncertainty about environmental interference and , This means that there is interference from complex sea conditions; Indicates general sea environment disturbance; Indicates uncertainty about environmental interference The upper bound of , express A subset of and is a compact set; represents the set of all real numbers; represents all real vectors n Vieux-style space; represents control input; Indicates the heading of the unmanned ship; represents the bow angular velocity of the unmanned ship; Indicates the maneuverability index of the unmanned ship rudder gain Time constant with ship maneuverability The intermediate parameter quantity;

[0127] S2: Obtain the unmanned ship virtual control law according to the unmanned ship heading tracking control model;

[0128] The specific steps include:

[0129] S21: defining a model error according to the unmanned ship heading tracking control model;

[0130] And the model error includes the heading tracking error of the unmanned ship Steering angular velocity tracking error ;

[0131] The expression of the model error is:

[0132] ,

[0133] ,

[0134] Where: A reference signal indicating the heading of the unmanned ship; represents the virtual control law of the unmanned ship to be designed; represents the steering angular velocity tracking error and , It represents the steering angular velocity tracking error under the interference of complex sea conditions; It represents the steering angular velocity tracking error under the general sea condition interference; Represents the heading tracking error of the unmanned ship;

[0135] S22: Constructing Lyapunov function based on heading tracking error , whose expression is

[0136] ,

[0137] And for the Lyapunov function Derivative, get the derivative of the Lyapunov function ,

[0138] ,

[0139] S23: Construct derivatives to ensure Lyapunov functions A virtual control law that is less than or equal to 0, and the expression of the virtual control law is

[0140] ,

[0141] Where: Represents the design parameters and satisfies ;

[0142] In addition, the virtual control law designed above is substituted into the derivative of the Lyapunov function in S22 , we can get:

[0143] ,

[0144] S3: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the feedforward and feedback items of the unmanned ship heading tracking controller are obtained, and a feedforward optimization controller for the feedforward item is constructed;

[0145] The specific steps include:

[0146] S31: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is:

[0147] ,

[0148] Where: Indicates a switching signal; It represents the control input of the switching controller in response to complex sea conditions; Indicates complex sea conditions interference;

[0149] S32: Control input of the switching controller to cope with complex sea conditions Defined as

[0150] ,

[0151] Where: represents the feedforward term of the switching controller; represents the feedback term of the switching controller;

[0152] S33: Based on step S32, the steering angular velocity tracking error after derivation is rewritten, and its expression is:

[0153] ,

[0154] And based on the feedforward term in the rewritten derivative steering angular velocity tracking error, a feedforward optimization controller is constructed, and its expression is:

[0155] ,

[0156] Where: Represents the design parameters and satisfies ;

[0157] S4: Based on the reinforcement learning method, a heading tracking optimization controller suitable for the heading tracking control of the unmanned ship under complex sea conditions is constructed according to the feedback item, and a first switching controller for coping with complex sea conditions is obtained according to the feedforward optimization controller and the heading tracking optimization controller;

[0158] The specific steps include:

[0159] S41: Unknown terms in the mathematical model of unmanned ship heading tracking control With ship parameters , reconstruct the feedback item of the switching controller in step S33 to obtain a reconstructed feedback item;

[0160] The expression of the reconstruction feedback term is:

[0161] ,

[0162] Where: represents the reconstruction feedback term and ; express The first derivative of Represents the unknown term in the unmanned ship heading tracking control model ;

[0163] S42: Design a first cost function based on the reconstruction feedback term, and its expression is:

[0164] ,

[0165] Where: represents the first cost function; express The integrand of ; express The abbreviation of express The abbreviation of Indicates the switching signal for dealing with complex sea conditions The optimal controller under represents the positive definite term about the partial derivative of the first cost function;

[0166] S43: Solve and obtain the first cost function for the reconstruction feedback term The partial derivative of is substituted into the constructed HJB equation, which is expressed as

[0167] ,

[0168] ,

[0169] Where: Denotes the first cost function with respect to the reconstruction feedback term The partial derivative of represents the design constant, express The abbreviation of ; express The abbreviation of represents a positive definite constant;

[0170] S44: Using the optimal condition equation Solve the constructed HJB equation to obtain the switching signal in response to complex sea conditions The optimal controller under

[0171] ,

[0172] S45: Use a single critic neural network to approximate the first cost function, and based on the approximated first cost function, obtain its partial derivative, which is expressed as

[0173] ,

[0174] ,

[0175] Where: Represents the weights of the neural network; Represents the activation function of the neural network; represents the approximation error of the neural network; express The partial derivative of express The partial derivative of

[0176] S46: Obtain partial derivatives based on step S45 Estimated value of , to obtain the optimal controller for a single ciritc network controller , whose expression is

[0177] ,

[0178] ,

[0179] ,

[0180] Where: express The expected value and ; Indicates expected value The error between the true value and

[0181] S47: Substituting the output of the single ciritc network controller into step S41, we can get

[0182] ,

[0183] S48: The The same as in step S46 Substituting into step S43 and combining with step S47, we can get

[0184] ,

[0185] And further simplifying it can get

[0186] ;

[0187] Where: represents the intermediate parameter quantity and

[0188] ;

[0189] S49: The and Substituting into step S43 and combining with step S47, we can get

[0190] ,

[0191] S410: According to steps S48 and S49, the variable error can be obtained , whose expression is

[0192] ,

[0193] Where: , represents the positive definite matrix about the neural network basis function;

[0194] S411: Constructing the first cost function The Lyapunov function of

[0195] ,

[0196] Then there must exist a positive definite matrix , so that

[0197] ;

[0198] Combined with variable error Construct the adaptive rate of the heading tracking optimization controller, which is expressed as

[0199] ,

[0200] ,

[0201] ,

[0202] Where: , represents the design parameters;

[0203] S412: For the adaptive rate designed in S411, according to , we can get:

[0204] ,

[0205] To be used for proof of subsequent stability analysis;

[0206] S413: combining the adaptive rate with the single network controller of the optimal controller in step S46 , we can obtain the heading tracking optimization controller suitable for unmanned ship heading tracking control in complex sea conditions. ,

[0207] ;

[0208] S414: The step S413 As a switching controller in step S32 to deal with complex sea conditions of , and combined with the feedforward optimization controller, a first switching controller for coping with complex sea conditions is obtained;

[0209] S5: Using the backstepping method combined with the RBF neural network, a second switching controller for coping with general sea conditions is constructed according to the feedback term;

[0210] The specific steps include:

[0211] S51: Based on the backstepping method and the unmanned ship virtual control law, combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is:

[0212] ,

[0213] Where: Indicates the switching signal corresponding to general sea conditions; Indicates the switching controller for dealing with general sea conditions; Indicates general sea disturbance;

[0214] S52: Transform the formula of step S51, and the transformation formula is:

[0215] ,

[0216] And use RBF neural network to approximate the transformation formula , can get

[0217] ;

[0218] S53: Based on step S52, a second switching controller for coping with general sea conditions is constructed, and its expression is:

[0219] ,

[0220] Where: express The expected value and ; Indicates expected value With the true value The error between Represents the weight of the RBF neural network, and the adaptive rate of the weight of the RBF neural network is , and Both represent adjustment parameters and satisfy , ; Represents the basis function of the RBF neural network; represents the design parameters;

[0221] S6: Using the step function, a switching response controller for different sea conditions is designed according to the first switching controller and the second switching controller to realize the knowledge-driven reinforcement learning unmanned ship heading tracking control, specifically including

[0222] S61: For the controllers under two different sea conditions, namely the first switching controller and the second switching controller, the switching setting is realized, and its expression is:

[0223] ,

[0224] In this embodiment, the sea conditions at sea are generally divided into nine levels, 0-9. Above level 7, ships are prohibited from sailing. Therefore, level 4-6 sea conditions are regarded as bad sea conditions, and the interference is set to , treat sea conditions of level 0-3 as normal sea conditions, and set the interference to , and the switching signal under adverse sea conditions is defined as , the switching signal under general sea conditions is defined as ;

[0225] S62: Using the step function, a switching response controller for different sea conditions is designed according to the first switching controller and the second switching controller, and its expression is:

[0226] ,

[0227] Where: express This embodiment assumes that Before the time, the sea conditions were normal; assuming Therefore, this embodiment uses a step function to design a switching response controller for different sea conditions to achieve switching between the system and the controller under different sea conditions, thereby adapting to different sea conditions to achieve knowledge-driven unmanned ship reinforcement learning heading tracking control.

[0228] This example also includes step 7: using Lyapunov theory, prove the stability of the designed knowledge-driven unmanned ship reinforcement learning heading tracking control method, and all signals in the closed-loop system are ultimately uniformly bounded.

[0229] Specifically, in order to prove stability more smoothly, the following assumptions are made here:

[0230] Assumption 1: Assume that the interference caused by the environment is bounded and satisfies , ;

[0231] Theorem 1: For the knowledge-driven unmanned ship reinforcement learning heading tracking control method, based on Lyapunov theory, under the action of the designed virtual control law, feedforward controller, first switching controller and its adaptive rate, second switching controller and its neural network adaptive rate, by selecting appropriate parameters, it can be ensured that all signals in the closed-loop system are ultimately uniformly bounded, and the tracking error converges within a reasonable range.

[0232] Lemma 1: (Young's inequality), for any , the following inequality holds:

[0233] ,

[0234] in, And (b-1)(q-1)=1.

[0235] Lemma 2: According to , the following inequality holds:

[0236] ,

[0237] Next, we prove Theorem 1:

[0238] S71: First select the controller The Lyapunov function under the action is as follows:

[0239] ,

[0240] S72: Lyapunov function Taking the derivative, we get the following form

[0241] ,

[0242] S73: Approximate the transformed image in S31 and S52 using RBF neural network and S53 into S72, using Lemma 1 and according to The following form can be obtained

[0243] ,

[0244] Where: represents an intermediate variable and ;

[0245] S74: Substitute S41 and S412 into S73 for the simplified Lyapunov function above, and according to Lemma 2, we can get the following form

[0246] ;

[0247] S75: For the further simplified Lyapunov function above, the neural network approximation in S41 and S52 is , and using Substituting in, we can get:

[0248] ;

[0249] S76: Further simplifying the above Lyapunov function, we can get:

[0250] ,

[0251] S77: According to Lemma 1 and Lemma 2, the Lyapunov function in S76 can be further simplified to obtain:

[0252] ,

[0253] S78: Further simplifying the Lyapunov function in S77, we can get:

[0254] ,

[0255] In the formula, , if and only if or When , the above formula holds; by selecting appropriate parameters, the signal in the system can be made consistent and stable, and then Theorem 1 holds.

[0256] Compared with the existing technology, this embodiment has the following advantages:

[0257] 1. In this embodiment, two controllers are designed to be suitable for different sea conditions. The first switching controller is applied to In complex sea conditions, the second switching controller feedback controller is applied to Under normal sea conditions, it enables effective tracking and control according to sea conditions of varying complexity;

[0258] 2. This embodiment switches the controllers under two different sea conditions when the sea conditions change, and the control effect of the controller is good;

[0259] 3. This embodiment conducts simulation experiments based on the MATLAB platform and designs two sets of adaptive controllers for switching control to achieve tracking of the unmanned ship's heading under different sea conditions. Figures 2 to 11 As shown, it can be seen that the switching response controller implemented in this embodiment is more suitable for navigation practice under different sea conditions, and the path tracking effect achieved is more ideal, which improves the tracking ability of the ship heading tracking control system under different sea conditions. Figure 2 This is a simulation diagram of the first tracking result of the ship heading angle relative to the reference heading in the embodiment; Figure 3 It is a simulation diagram of the second tracking result of the ship heading angle to the reference heading in the embodiment; Figure 4 It is a simulation diagram of the first tracking error between the ship heading angle and the reference heading in the embodiment; Figure 5 A second tracking error simulation diagram of the ship heading angle and the reference heading in the embodiment; Figure 6 It is a simulation diagram of the convergence results of the neural network neuron weights under general sea conditions in the embodiment; Figure 7 The embodiment is a simulation diagram of the convergence results of the neural network neuron weights under complex sea conditions; Figure 8 It is a first control effect simulation diagram of the switching response controller in the embodiment; Fig. 9 A second control effect simulation diagram of the switching response controller provided in the embodiment; Fig.10 This is a simulation diagram of the first sea condition effect when the unmanned ship is affected by environmental interference in the embodiment; Fig.11 This is a simulation diagram of the second sea condition effect when the unmanned ship is affected by environmental interference in the embodiment, which further verifies the effectiveness and rationality of the knowledge-driven unmanned ship heading tracking control method provided by the present invention.

[0260] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A knowledge-driven unmanned ship reinforcement learning heading tracking control method, characterized in that: The specific steps include: S1: Obtain a mathematical model for unmanned ship heading control with uncertain interference terms, considering that the unmanned ship sailing on the sea will be affected by environmental interference; The unmanned ship heading control mathematical model is converted into a second-order state space equation applicable to different sea conditions to serve as the unmanned ship heading tracking control model; S2: Obtain the unmanned ship virtual control law according to the unmanned ship heading tracking control model; S3: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the feedforward term and feedback term of the unmanned ship heading tracking controller are obtained; And construct a feedforward optimization controller with respect to the feedforward term; S4: Based on the reinforcement learning method, a heading tracking optimization controller suitable for the heading tracking control of the unmanned ship under complex sea conditions is constructed according to the feedback item, and a first switching controller for coping with complex sea conditions is obtained according to the feedforward optimization controller and the heading tracking optimization controller; S5: Using the backstepping method combined with the RBF neural network, a second switching controller for coping with general sea conditions is constructed according to the feedback term; S6: Using a step function, a switching response controller for different sea conditions is designed according to the first switching controller and the second switching controller to realize knowledge-driven reinforcement learning unmanned ship heading tracking control.

2. According to claim 1, a knowledge-driven unmanned ship reinforcement learning heading tracking control method is characterized in that: The S1 specifically includes the following steps: S11: Considering that unmanned ships sailing on the sea will be affected by various interferences, the mathematical model of unmanned ship heading tracking control with uncertain interference terms is expressed as In the formula: In the formula: ψ represents the heading angle of the unmanned ship; δ represents the actual rudder angle of the unmanned ship; δ w represents the equivalent rudder angle caused by uncertain environmental interference; K represents the maneuverability index of the unmanned ship rudder gain; T represents the time constant of the ship's maneuverability; α and β both represent the nonlinear coefficients of the ship's bow angular velocity; represents the first derivative of ψ, which is the bow angular velocity of the ship; represents the second-order derivative of ψ, which is the ship's bow angular acceleration; S12: Let x 1i =ψ, And satisfy ||Δ i ||≤Δ * , the unmanned ship heading control mathematical model is converted into a second-order state space equation and used as the heading tracking control mathematical model. The expression of the heading tracking control mathematical model is: Where: Δ i Indicates uncertain environmental interference and i=1,2, i=1 indicates complex sea environment interference; i=2 indicates general sea environment interference; Δ * represents the upper bound of the uncertain environmental disturbance Δ; Ω represents R n is a subset of and is a compact set; R represents the set of all real numbers; R n represents the n-dimensional Euclidean space of all real vectors; U represents the control input; x 1i Indicates the heading of the unmanned ship; x 2i The bow angular velocity g of the unmanned ship represents an intermediate parameter quantity related to the unmanned ship rudder gain maneuverability index K and the time constant T of the ship maneuverability.

3. The knowledge-driven unmanned ship reinforcement learning heading tracking control method according to claim 2 is characterized in that: The S2 specifically includes the following steps: S21: defining a model error according to the unmanned ship heading tracking control model; The model error includes the unmanned ship's heading tracking error z1 and the steering angular velocity tracking error z 2i ; The expression of the model error is: z1=x 1i -y d z 2i =x 2i -α1 Where: y d represents the reference signal of the unmanned ship’s heading; α1 represents the unmanned ship’s virtual control law to be designed; z 2i represents the steering angular velocity tracking error and i=1,2,z 21 represents the steering angular velocity tracking error under the interference of complex sea conditions; z 22 represents the steering angular velocity tracking error under general sea conditions; z1 represents the heading tracking error of the unmanned ship; S22: Construct the Lyapunov function V1 based on the heading tracking error, and its expression is: And take the derivative of the Lyapunov function V1 to obtain the derivative of the Lyapunov function S23: Construct derivatives to ensure Lyapunov functions A virtual control law that is less than or equal to 0, and the expression of the virtual control law is Where: c1 represents the design parameter and satisfies c1≥0.

4. The knowledge-driven unmanned ship reinforcement learning heading tracking control method according to claim 3 is characterized in that: The S3 specifically includes the following steps: S31: Based on the unmanned ship virtual control law and combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is: Where: σ1 represents the switching signal; represents the control input of the switching controller in response to complex sea conditions; Δ1 represents the disturbance of complex sea conditions; S32: Control input of the switching controller to cope with complex sea conditions Defined as Where: u1 represents the feedforward term of the switching controller; u2 represents the feedback term of the switching controller; S33: Based on step S32, the steering angular velocity tracking error after derivation is rewritten, and its expression is: And based on the feedforward term in the rewritten derivative steering angular velocity tracking error, a feedforward optimization controller is constructed, and its expression is: Where: c2 represents the design parameter and satisfies c2≥0.

5. The knowledge-driven unmanned ship reinforcement learning heading tracking control method according to claim 4 is characterized in that: The S4 specifically comprises the following steps: S41: Based on the unknown term f(x in the mathematical model of unmanned ship heading tracking control 21 ) and the ship parameter g, reconstructing the feedback item of the switching controller in step S33 to obtain a reconstructed feedback item; The expression of the reconstructed feedback term is: Where: ζ represents the reconstruction feedback term and ζ = z 21 +α1; represents the first-order derivative of ζ; F(ζ) represents the unknown term f(x 21 ); S42: Design a first cost function based on the reconstruction feedback term, and its expression is: Where: J1(ζ) represents the first cost function; h1(ζ(τ),u2(ζ)) represents the integrand of J1(ζ); ζ represents the abbreviated form of ζ(τ); u2 represents the abbreviated form of u2(ζ); represents the optimal controller under the switching signal σ1 in response to complex sea conditions; Q(ζ) represents the positive definite term with respect to the partial derivative of the first cost function; S43: Solve and obtain the partial derivative of the first cost function with respect to the reconstruction feedback term ζ, and substitute it into the constructed HJB equation, which is expressed as Where: represents the partial derivative of the first cost function with respect to the reconstruction feedback term ζ; represents the design constant, R1 represents The abbreviation of express The abbreviation of λ f represents a positive definite constant; S44: Using the optimal condition equation Solve the constructed HJB equation to obtain the optimal controller under the switching signal σ1 in response to complex sea conditions, and its expression is S45: Use a single critic neural network to approximate the first cost function, and based on the approximated first cost function, obtain its partial derivative, which is expressed as Where: W c Represents the weight of the neural network; S c (ζ) represents the activation function of the neural network; ε c (ζ) represents the approximation error of the neural network; Represents ε c The partial derivative of (ζ); Indicates S c The partial derivative of (ζ); S46: Obtain partial derivatives based on step S45 Estimated value of To obtain the optimal controller single ciritc network controller Its expression is Where: W c The expected value and Indicates expected value The error between the true value and S47: Substituting the output of the single ciritc network controller into step S41, we can get S48: The The same as in step S46 Substituting into step S43 and combining with step S47, we can get And further simplifying it can get Where: e(t) represents the intermediate parameter and S49: The and Substituting into step S43 and combining with step S47, we can get S410: According to steps S48 and S49, the variable error e can be obtained c , whose expression is Where: represents the positive definite matrix about the neural network basis function; S411: Construct the Lyapunov function of the first cost function J1(ζ), which is expressed as Combined with the variable error e c Construct the adaptive rate of the heading tracking optimization controller, which is expressed as Where: α c ,α s represents the design parameters; S412: combining the adaptive rate with the single network controller of the optimal controller in step S46 The heading tracking optimization controller suitable for unmanned ship heading tracking control in complex sea conditions can be obtained. S413: The step S412 As a switching controller in step SS32 to deal with complex sea conditions ’s u2, and combined with the feedforward optimization controller, the first switching controller for coping with complex sea conditions is obtained.

6. The knowledge-driven unmanned ship reinforcement learning heading tracking control method according to claim 5 is characterized in that: The S5 specifically includes the following steps: S51: Based on the backstepping method and the unmanned ship virtual control law, combined with the steering angular velocity of the unmanned ship heading tracking control model, the steering angular velocity tracking error is derived, and its expression is: Where: σ2 represents the switching signal corresponding to general sea conditions; Indicates the switching controller for dealing with general sea conditions; Δ2 indicates general sea condition interference; S52: Transform the formula of step S51, and the transformation formula is: And use RBF neural network to approximate the transformation formula Be able to get S53: Based on step S52, a second switching controller for coping with general sea conditions is constructed, and its expression is: Where: W k The expected value and Indicates expected value and the true value W k The error between k Represents the weight of the RBF neural network, and the adaptive rate of the weight of the RBF neural network is Γ and σ are both adjustment parameters and satisfy Γ>0, σ>0; S(x 22 ) represents the basis function of the RBF neural network; c3 represents the design parameter.

7. The knowledge-driven unmanned ship reinforcement learning heading tracking control method according to claim 6 is characterized in that: In S6, a step function is used to design a switching response controller for different sea conditions according to the first switching controller and the second switching controller, and its expression is: Where: ε(t) represents the step function at time t.

Citation Information

Patent Citations

  • Vehicle autonomous limit driving planning control method and system based on reinforcement learning

    CN114348021A

  • Adaptive control system having direct output feedback and related apparatuses and methods

    WO2001092974A2