Unmanned Surface Vehicle Trajectory Tracking Control Method Based on Actor-Critic-Advantage Network

By introducing the dominant function estimation network into the Actor-Critic network, the Actor-Critic-Advantage network is formed, which solves the problem of insufficient accuracy and timeliness in unmanned craft trajectory tracking control, and realizes efficient and stable trajectory tracking control.

CN115793455BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211507269.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-05-30
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

The existing unmanned craft trajectory tracking control methods cannot guarantee accuracy and timeliness.

Method used

Using the method based on the Actor-Critic-Advantage network, the network is estimated by introducing the advantageous function, a new Actor-Critic-Advantage network is formed to perform unmanned craft trajectory tracking control. The specific steps include training a new network, updating the policy network using a single-step acquisition strategy gradient method, designing segmented reward functions based on the inverse step method, and introducing virtual control laws to improve training efficiency and accuracy.

Benefits of technology

The accuracy and timeliness of unmanned craft trajectory tracking control are achieved, the training convergence problem caused by inaccurate estimation of dominant functions is overcome, the sample utilization rate and training stability are improved, and the anti-interference ability and adaptability are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115793455B_ABST
    Figure CN115793455B_ABST
Patent Text Reader

Abstract

The present invention provides an unmanned boat trajectory tracking control method based on an Actor-Critic-Advantage network, including: introducing an advantage function estimation network on the basis of the Actor-Critic network to form a new Actor-Critic-Advantage network; training the new Actor-Critic-Advantage network for unmanned boat trajectory tracking control; the unmanned boat trajectory tracking training adopts a single-step acquisition strategy gradient method, and uses the output value of the advantage function estimation network to obtain the policy gradient to update the policy network; solving the virtual control law based on the backstepping method to design a piecewise reward function; introducing the virtual control law into the reward function to train the speed output of the unmanned boat to tend to the virtual control law.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned boats, and particularly to a method for controlling the trajectory tracking of an unmanned boat based on an Actor-Critic-Advantage network. Background Art

[0002] In recent years, with the depletion of land fuel resources, the strategic position of the ocean, which occupies about 71% of the earth's area, has been continuously improving. To fully explore and exploit marine resources, the development of marine equipment technology is essential. Marine intelligent equipment represented by unmanned boats (including underwater vehicles, underwater robots, surface unmanned boats, etc.) is the main carrier of current offshore operations.

[0003] In recent years, the application of unmanned boats has been increasing. Currently, unmanned boats have played an important role in military fields such as encirclement and suppression, expulsion, mine sweeping, and anti-submarine warfare, as well as in civil fields such as material supply, topographic surveying, sea rescue, and unmanned search.

[0004] However, the trajectory tracking control of unmanned boats often has problems in ensuring accuracy and timeliness. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for controlling the trajectory tracking of an unmanned boat based on an Actor-Critic-Advantage network to solve the problem that the existing trajectory tracking control of unmanned boats cannot ensure accuracy and timeliness.

[0006] To solve the above technical problems, the present invention provides a method for controlling the trajectory tracking of an unmanned boat based on an Actor-Critic-Advantage network, including:

[0007] Introduce an advantage function estimation network into the Actor-Critic network to form a new Actor-Critic-Advantage network;

[0008] Train the new Actor-Critic-Advantage network to perform trajectory tracking control of the unmanned boat;

[0009] Use the single-step acquisition policy gradient method to execute the trajectory tracking training of the unmanned boat, and use the output value of the advantage function estimation network to obtain the policy gradient to update the policy network;

[0010] Design a piecewise reward function based on the backstepping method to solve the virtual control law; and

[0011] Introduce the virtual control law into the reward function to make the speed output of the trained unmanned boat tend to the virtual control law.

[0012] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0013] Step 1: Build an environmental model for unmanned boat trajectory tracking, simulate the navigation of an unmanned boat in a real marine environment with interference, and randomly generate an expected trajectory with time constraints and the starting position and heading angle of the unmanned boat.

[0014] The unmanned boat trajectory tracking is independent of the vertical space. According to the established mathematical model of a three-degree-of-freedom underactuated unmanned boat, obtain the state information of the unmanned boat at any moment in the environmental model.

[0015] The kinematic model expression of the unmanned boat is:

[0016]

[0017] where η = [x, y, ψ] T , where (x, y) represents the position of the unmanned boat, and ψ represents the angle between the hull coordinate system and the earth coordinate system, that is, the heading angle of the unmanned boat; υ = [u, v, r] T , and their symbols respectively represent the longitudinal velocity, lateral velocity, and yaw angular velocity in the hull coordinate system, and R(η) represents the rotation matrix for transforming the hull coordinate system to the earth coordinate system.

[0018] The dynamic kinematic model expression of the unmanned boat is:

[0019]

[0020] where τ = [τ u , 0, τ r T , τ u 、τ r respectively represent the longitudinal thrust and turning moment of the unmanned boat in the environmental model, τ e = [τ eu , τ ev , τ er T , τ eu 、τ ev 、τ er respectively represent the disturbances exerted on the unmanned boat by the marine environment in the u, ν, and r directions, and M, C(υ), and D(υ) respectively represent the inertia matrix, centripetal force matrix, and damping matrix of the unmanned boat.

[0021] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0022] Step 2: Set the action space and state space.

[0023] According to the environmental model, the state space of the unmanned boat at the current moment t is designed as s t =[x e , y e , ψ, ψ d , u, ν, r], where x e , y e represent the position error of the tracking trajectory of the unmanned boat in the geodetic coordinate system in the environmental model, ψ, ψ d represent the current heading angle and the desired heading angle, and u, v, and r respectively represent the longitudinal speed, lateral speed, and heading angular velocity of the unmanned boat in the hull coordinate system;

[0024] The action space of the unmanned boat at the current moment t is a t =[τ u , τ r , which respectively represent the longitudinal thrust and steering moment of the unmanned boat in the environmental model, and are converted into as the input to interact with the environmental model.

[0025] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0026] Step 3: Design a new Actor-Critic-Advantage network, which includes an Actor network, a Critic network, and an Advantage network, corresponding to the decision-making network, the evaluation network, and the advantage function estimation network respectively. Its update method is specifically as follows:

[0027] Step 3.1: The input of the evaluation network is the state s t , and the output is the state value The evaluation network updates with the mean square of the difference between the estimated state value and the target state value as the loss function. Its loss function is:

[0028]

[0029] where γ is the discount factor, with a value range of [0, 1], and r t is the real-time reward obtained by the unmanned boat interacting with the environment, and the reward at the current moment is calculated through Step 4.

[0030] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes

[0031] Step 3.2: The estimation of the advantage function is obtained according to the generalized advantage estimation method as:

[0032]

[0033] Among them, λ is the error weight coefficient, which determines the deviation magnitude of the estimator, and its value range is [0, 1]; δ t is the time series difference error,

[0034] The generalized advantage function estimation method includes: by introducing the variance of the advantage function estimation, reducing δ t as the bias of the advantage function, the generalized advantage function estimation method calculates the estimated value of the advantage function, making the bias and variance better balanced;

[0035] According to the expression of the generalized advantage function estimation, the estimated value of the advantage function under the current state and action is related to the subsequent moment. The following derivation is carried out for the generalized advantage function estimation:

[0036]

[0037] The recursive formula for the advantage function estimation is obtained:

[0038] A(s t ,a t ) = δ t + γλA(s t+1 ,a t+1 )

[0039] Construct an advantage function estimation network to approximate the generalized advantage function estimation value A(s t ,a t ). The estimated value of the advantage function satisfies the following formula:

[0040]

[0041] The input of the advantage function estimation network is the state s t and the action a t in the environmental model, and the output is the estimated value A ω (s t ,a t ) of the advantage function. The advantage function estimation network updates with the mean square error of the difference between the estimated value A ω (s t ,a t ) of the advantage function and the target estimated value δ t + γλA ω (s t+1 ,a t+1 ) as the loss function. Its loss function is:

[0042]

[0043] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0044] Step 3.3: The input of the decision-making network is the state s in the environment model t , and the output is the action a of the unmanned boat t ~π(a t |s t ), where π(a t |s t ) represents the probability of executing action a t under state s t ;

[0045] When solving the policy gradient of the policy network, the true value of the advantage function is used:

[0046] A π (s t ,a t ) = Q π (s t ,a t ) - V π (s t )

[0047] where, A π (s t ,a t ) is the true value of the advantage function, Q π (s t ,a t ) is the true value of the action-state value function, and V π (s t ) is the true value of the state value function;

[0048] The update expression of the parameter θ of the decision-making network π θ (a|s) is:

[0049]

[0050] where, α is the learning rate of the decision-making network;

[0051] In the update of the parameters of the decision-making neural network, the true value of the advantage function is used. In actual operation, is used to replace the true value of the advantage function for calculation, and the state value t in δ is obtained through the evaluation network, and there is a deviation in using δ t as the advantage function;

[0052] The generalized advantage function estimation method is based on δ at subsequent moments t+l, obtained by summing from \(l = 0, 1, \cdots, \infty\). Therefore, by introducing the advantage function estimation network \(A\) ω \((s,a)\) to approximate the advantage function estimate \(A(s\) t ,a t );

[0053] So the decision neural network policy gradient is:

[0054]

[0055] Optionally, in the above-mentioned unmanned surface vehicle trajectory tracking control method based on the Actor - Critic - Advantage network, it further includes:

[0056] Step Four: Determine the reward function;

[0057] Unmanned surface vehicle trajectory tracking control means that the unmanned surface vehicle starts from an arbitrary position and reaches the desired trajectory with time constraints and heads along the desired trajectory with as small a tracking error as possible;

[0058] By designing the reward function to evaluate the state in the environmental model at the current moment, the control objective of unmanned surface vehicle trajectory tracking is completed;

[0059] Following the framework derived by the backstepping method, construct the coordinate transformation: \(z\) 1 =\(\eta-\eta\) d , \(z\) 2 =\(v - v\) d ; where, \(\eta\) d is the desired trajectory, \(\upsilon\) d is the virtual control law;

[0060] For the unmanned surface vehicle mathematical model mentioned in Step One, design the virtual control law \(\nu\) d , ensuring that the tracking error \(z\) 1 is small enough to achieve the trajectory tracking goal; design the Lyapunov function:

[0061]

[0062] Obviously, \(V\) 1 \(\geq0\), positive definite, take its first derivative:

[0063]

[0064] Ideally, the virtual control law \(\upsilon\) d is designed as:

[0065]

[0066] where \(k\) 1 is a symmetric positive definite matrix;

[0067] Ensure υ→υ d That is Then η→η d , to achieve the trajectory tracking control of the unmanned boat;

[0068] Take k 1 = diag(k 11 , k 22 , k 33 ) The virtual control law ν c The conversion process is expressed as a system of equations:

[0069]

[0070]

[0071]

[0072] Design the reward function

[0073]

[0074] Among them, l 1 , l 2 , l 3 are adjustment coefficients, and d e is the expected trajectory tracking error.

[0075] Optionally, in the above-mentioned unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0076] Step Five: Training the controller based on the new Actor-Critic-Advantage network;

[0077] Step 5.1: Construct the network in Step Two, including five network structures: the current evaluation network The target evaluation network The generalized advantage function estimation network A ω (s, a), the current policy network π θ (a|s), and the target policy network π θ′ (a|s). The current evaluation network and the target evaluation network have the same structure, and the current policy network and the target policy network have the same structure;

[0078] Step 5.2: Initialize the number of experience samples E obtained from the experience data buffer and the state of the unmanned boat in the model environment, and initialize the network model parameters ω, θ, copy the parameters of the current evaluation network to the target evaluation network Copy the parameters of the current policy network to the target policy network θ→θ′;

[0079] Step 5.3: According to the state s of the current unmanned boat t input it into the current policy network, and output the action a t ~π θ (a t |s t ), input it into the action converter to obtain the action Then input it into the environment model to obtain the state s at the next moment t+1 , and obtain the immediate reward r based on Step 4 t , finally input the state s t+1 into the current policy network;

[0080] Loop and execute Step 5.3, and store the experience samples (s t , a t , r t , s t+1 ) generated in each process into the experience data buffer.

[0081] Optionally, in the above unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0082] Step 5.4: When the number of samples in the experience data buffer exceeds E, terminate the loop of Step 5.3, and select the experience samples (s i , a i , r i , s i+1 ), i = 1,..., E, input the states s i , s i+1 into the current evaluation network and the target policy network respectively to obtain the state values Calculate and calculate the loss function for training the current evaluation network based on Step 2.1:

[0083]

[0084] Update the weight parameters of the current evaluation network based on this loss function;

[0085] Step 5.5: Input the action s i+1 into the target policy network to output the action a i+1 ~π θ (a i+1 |s i+1 ), obtain the action a i+1 , then input s i , a i and s i+1 , a i+1 into the advantage function estimation network respectively to obtain the advantage function estimation value Aω (s i ,a i ) and A ω (s i+1 ,a i+1 ) and calculate the loss function for training the advantage function estimation network based on Step 2.2:

[0086]

[0087] Optionally, in the unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, it further includes:

[0088] Update the weight parameters of the advantage function estimation network based on this loss function;

[0089] Step 5.6: For the training of the current policy network, use the policy gradient algorithm to update the parameters of the current policy network, and derive and calculate the policy gradient according to Step 2.3:

[0090]

[0091] Step 5.7: Update the parameters of the target evaluation network and the target policy network in a soft update manner:

[0092]

[0093] θ′ = τθ + (1 - τ)θ′

[0094] where τ is the soft update coefficient, and its value range is [0, 1];

[0095] Step 5.8: Loop through Steps 5.3 - 5.7, observe whether the training end condition is reached. If it is reached, end the training, and the policy network generates the optimal policy Using the optimal policy can control the unmanned boat at any initial position and complete the tracking of any desired trajectory.

[0096] The inventor of the present invention found through research that in the research on the unmanned boat trajectory tracking control method, the prior art often only focuses on the method of using the conventional Actor-Critic network to achieve unmanned boat trajectory tracking control, and has not deeply studied whether the Actor-Critic network can be improved to give full play to the advantages of the Actor-Critic network in the unmanned boat trajectory tracking control problem, thus resulting in limitations in the application of reinforcement learning in unmanned boat trajectory tracking control.

[0097] In addition, for reinforcement learning based on the generalized advantage function method, it can only obtain numerical values through a fixed multi-step operation, and then calculate the advantage function value at each moment by backward recursion from the last moment, and then update the policy network using policy gradients. This has led to the problems of difficult acquisition of policy gradients and large computational complexity in the existing technical methods. It is necessary to develop a method for updating the policy network by obtaining policy gradients in a single step for fast and effective training of the unmanned boat trajectory tracking.

[0098] Finally, the design of the existing reward function for unmanned boats often only considers the change in Euclidean distance information and ignores the change in speed information. In particular, the coupling of the position errors in the x and y directions of the Euclidean distance information makes it impossible to accurately evaluate the current state of the unmanned boat, resulting in poor training effects for the unmanned boat trajectory tracking. If the position error information is set as a sub-reward function and constitutes the reward function in the form of weight parameters, it will lead to a waste of a large amount of time in debugging the weight parameters of the reward function. Therefore, it is necessary to develop a piecewise reward function based on the backstepping method to solve the virtual control law, which can achieve the unmanned boat trajectory tracking error, converge in a timely and effective manner with small fluctuations, and meet the tracking error accuracy requirements.

[0099] Based on the above insights, the present invention provides a method for controlling the trajectory tracking of an unmanned boat based on an Actor-Critic-Advantage network, training a new Actor-Critic-Advantage network to achieve the trajectory tracking control of the unmanned boat: the new Actor-Critic-Advantage network introduces an advantage function estimation network on the basis of the Actor-Critic network, overcomes the difficulty of poor training convergence effect or even non-convergence caused by large estimation deviation due to inaccurate advantage function estimation, solves the problem of low sample utilization rate in the training process of the unmanned boat trajectory tracking, and ensures the smoothness of the training process of the unmanned boat trajectory tracking.

[0100] The trajectory tracking training of the unmanned boat in the present invention updates the policy network by obtaining policy gradients in a single step: no longer limited to the method of obtaining policy gradients in multiple steps in the Actor-Critic network by the generalized advantage function estimation method, making full use of the more easily obtained advantage function value to obtain the policy gradient to update the policy network, and quickly and effectively achieving the training goal of the unmanned boat trajectory tracking control.

[0101] The present invention designs a piecewise reward function based on the backstepping method to solve the virtual control law: overcomes the difficulty that the reward function designed based on position error cannot ensure the stability of the unmanned boat tracking error control method. The virtual control law is introduced into the reward function to train the speed output of the unmanned boat to tend to the virtual control law, ensuring that the unmanned boat tracking error converges in a timely, effective and small-fluctuation manner, and ensuring that the unmanned boat tracking error meets the tracking error accuracy requirements.

[0102] The present invention proposes a research on an unmanned surface vehicle trajectory tracking control method based on a novel Actor-Critic-Advantage network for the problem of unmanned surface vehicle trajectory tracking control, aiming to solve the problem that in the training process of unmanned surface vehicle trajectory tracking control based on the Actor-Critic network, the deviation of the advantage function estimation is too large, resulting in non-convergence or poor convergence results of the unmanned surface vehicle trajectory tracking training. Through the proposed improved method, the unmanned surface vehicle trajectory tracking control target can be quickly achieved, and at the same time, the anti-interference ability during trajectory tracking can be improved, with a high degree of adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] Figure 1 It is a description and model schematic diagram of the earth coordinate system Og-XgYg and the hull coordinate system Ob-XbYb of an unmanned surface vehicle according to an embodiment of the present invention;

[0104] Figure 2 It is a schematic diagram of the relationship between neural networks constructed in step two according to an embodiment of the present invention;

[0105] Figure 3 It is a schematic diagram of the unmanned surface vehicle achieving trajectory tracking control within a specified time of 30 s according to an embodiment of the present invention;

[0106] Figure 4 It is a schematic diagram of the effect of the controller trained by the novel Actor-Critic-Advantage network controlling the unmanned surface vehicle to track the desired trajectory according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0107] The present invention will be further described below in conjunction with the detailed description with reference to the accompanying drawings.

[0108] It should be noted that the components in the respective drawings may be exaggerated for illustration purposes and are not necessarily to scale correctly. In the respective drawings, the same or functionally identical components are provided with the same reference numerals.

[0109] In the present invention, unless otherwise specified, "arranged on...", "arranged above...", and "arranged on top of..." do not exclude the situation where there is an intermediate object between the two. In addition, "arranged on or above..." only represents the relative positional relationship between two components, and in certain cases, such as when the product direction is reversed, it can also be converted to "arranged under or below...", and vice versa.

[0110] In the present invention, the respective embodiments are only intended to illustrate the solutions of the present invention and should not be construed as restrictive.

[0111] In the present invention, unless otherwise specified, the quantifiers "a" and "one" do not exclude the scenario of multiple elements.

[0112] It should also be noted here that, in the embodiments of the present invention, for the sake of clarity and simplicity, only a part of the components or assemblies may be shown. However, those of ordinary skill in the art can understand that, under the teaching of the present invention, the required components or assemblies can be added according to the specific scenario requirements. In addition, unless otherwise specified, the features in different embodiments of the present invention can be combined with each other. For example, a certain feature in the second embodiment can be used to replace the corresponding or functionally identical or similar feature in the first embodiment, and the obtained embodiment also falls within the scope of disclosure or the scope of recording of this application.

[0113] It should also be noted here that within the scope of the present invention, terms such as "same", "equal", "equal to", etc. do not mean that the two values are absolutely equal, but allow a certain reasonable error. That is to say, these terms also cover "substantially the same", "substantially equal", "substantially equal to". By analogy, in the present invention, terms indicating directions such as "perpendicular to", "parallel to", etc. also cover the meanings of "substantially perpendicular to" and "substantially parallel to".

[0114] In addition, the numbering of the steps of each method of the present invention does not limit the execution order of the method steps. Unless otherwise specified, each method step can be executed in a different order.

[0115] The following further elaborates in detail on the unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network proposed by the present invention in combination with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings all adopt very simplified forms and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.

[0116] The purpose of the present invention is to provide an unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network to solve the problem that the existing unmanned boat trajectory tracking control cannot guarantee accuracy and timeliness.

[0117] To achieve the above purpose, the present invention provides an unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network, including: introducing an advantage function estimation network on the basis of the Actor-Critic network to form a new Actor-Critic-Advantage network; training the new Actor-Critic-Advantage network for unmanned boat trajectory tracking control; the unmanned boat trajectory tracking training adopts a single-step acquisition strategy gradient method, and uses the output value of the advantage function estimation network to obtain the policy gradient to update the policy network; solving the virtual control law based on the backstepping method to design a segmented reward function; introducing the virtual control law into the reward function to train the speed output of the unmanned boat to tend to the virtual control law.

[0118] The present invention proposes a research on an unmanned surface vehicle (USV) trajectory tracking control method based on a novel Actor-Critic-Advantage network for the problem of USV trajectory tracking control, aiming to solve the problem that in the training process of USV trajectory tracking control based on the Actor-Critic network, due to the too large deviation of the advantage function estimation, the USV trajectory tracking training does not converge or the convergence result is poor. Through the proposed improved method, the USV trajectory tracking control target can be quickly achieved, and at the same time, the anti-interference ability during trajectory tracking can be improved, with a high degree of adaptability.

[0119] The present invention trains a novel Actor-Critic-Advantage network to achieve USV trajectory tracking control: The novel Actor-Critic-Advantage network introduces an advantage function estimation network on the basis of the Actor-Critic network, overcomes the difficulty that the inaccurate estimation of the advantage function leads to poor training convergence effect or even non-convergence due to large estimation deviation, solves the problem of low sample utilization rate in the USV trajectory tracking training process, and ensures the smoothness of the USV trajectory tracking training process.

[0120] The USV trajectory tracking training of the present invention updates the policy network by adopting a single-step acquisition strategy gradient method: It is no longer limited to the method of adopting a multi-step acquisition strategy gradient in the Actor-Critic network by the generalized advantage function estimation method. By making full use of the more easily obtained advantage function values to obtain the policy gradient to update the policy network, the training goal of USV trajectory tracking control can be quickly and effectively achieved.

[0121] The present invention designs a piecewise reward function based on the backstepping method to solve the virtual control law: It overcomes the difficulty that the reward function designed based on the position error cannot ensure the stability of the USV tracking error control method. The virtual control law is introduced into the reward function to train the speed output of the USV to tend to the virtual control law, ensuring that the USV tracking error converges timely, effectively and with small fluctuations, and ensuring that the USV tracking error meets the tracking error accuracy requirements.

[0122] Figures 1 to 4 Embodiments of the present invention are provided. Step 1: Build an environmental model for USV trajectory tracking, simulate the navigation of the USV in a real marine environment with interference, and randomly generate an expected trajectory with time constraints and the starting position and heading angle of the USV.

[0123] The research on the USV trajectory tracking problem has nothing to do with the vertical space. Usually, the state information of the USV at any time in the environmental model is obtained according to the established three-degree-of-freedom underactuated USV mathematical model.

[0124] The description and model of the earth coordinate system Og-XgYg and the hull coordinate system Ob-XbYb of the unmanned boat are as follows Figure 1 as shown

[0125] The kinematic model expression of the unmanned boat is

[0126]

[0127] where η = [x, y, ψ] T , where (x, y) represents the position of the unmanned boat, and ψ represents the angle between the hull coordinate system and the earth coordinate system, that is, the heading angle of the unmanned boat. υ = [u, v, r] T , and their symbols respectively represent the longitudinal velocity, lateral velocity and yaw angular velocity in the hull coordinate system. R(η) represents the rotation matrix for transforming the hull coordinate system to the earth coordinate system

[0128] The dynamic kinematic model expression of the unmanned boat is

[0129]

[0130] where τ = [τ u , 0, τ r T , τ u and τ r respectively represent the longitudinal thrust and turning moment of the unmanned boat in the environmental model. τ e = [τ eu , τ ev , τ er T , τ eu , τ ev , and τ er respectively represent the disturbances exerted on the unmanned boat by the ocean environment in the u, ν, and r directions. M, C(υ), and D(υ) respectively represent the inertia matrix, centripetal force matrix, and damping matrix of the unmanned boat

[0131] Step 2: Set the action space and state space

[0132] According to the environmental model, the state space of the unmanned boat at the current time t is designed as s t = [x e , y e , ψ, ψ d , u, ν, r], where x e , y e represent the position error of the tracking trajectory of the unmanned boat in the earth coordinate system in the environmental model. ψ, ψ d represent the current heading angle and the desired heading angle. u, v, and r respectively represent the longitudinal velocity, lateral velocity, and yaw angular velocity of the unmanned boat in the hull coordinate system. The action space of the unmanned boat at the current time t is a​​t = [τ u , τ r respectively represent the longitudinal thrust and turning moment of the unmanned boat in the environmental model, and need to be converted by the action converter into as the input to interact with the environmental model.

[0133] Step 3: Design a new Actor-Critic-Advantage network. The Actor network, Critic network, and Advantage network respectively correspond to the decision-making network, evaluation network, and advantage function estimation network, and establish the above three networks. Their update methods are specifically as follows:

[0134] Step 3.1: The input of the evaluation network (Critic network) is the state s t , and the output is the state value The evaluation network updates with the mean square of the difference between the estimated state value and the target state value as the loss function. Its loss function is:

[0135]

[0136] where γ is the discount factor, with a value range of [0, 1], and r t is the real-time reward obtained by the unmanned boat interacting with the environment, and the reward at the current moment is calculated through Step 4.

[0137] Step 3.2: The estimation of the advantage function obtained according to the generalized advantage estimation method is:

[0138]

[0139] where λ is the error weight coefficient, which determines the deviation magnitude of the estimator, with a value range of [0, 1]. δ t is the temporal difference error,

[0140] The generalized advantage function estimation method reduces δ t to a certain extent as the deviation of the advantage function by introducing the variance of the advantage function estimation. The generalized advantage function estimation method calculates a better balance between the deviation and variance of the advantage function estimation value.

[0141] From the expression of the generalized advantage function estimation, it can be seen that the estimated value of the advantage function under the current state and action is related to subsequent moments. Therefore, the accurate estimated value of the advantage function cannot be calculated in actual operations. So, the following derivation is carried out for the generalized advantage function estimation:

[0142]

[0143] Obtain the recurrence formula for estimating the advantage function:

[0144] A(s t ,a t ) = δ t +γλA(s t+1 ,a t+1 )

[0145] By analogy with the update method of the evaluation network, construct an advantage function estimation network (Advantage network) to approximate the generalized advantage function estimation value A(s t ,a t ). The ideal advantage function estimation value should satisfy the following equation:

[0146]

[0147] The input of the advantage function estimation network is the state s t and the action a t in the environmental model, and the output is the advantage function estimation value A ω (s t ,a t ). The advantage function estimation network is updated using the mean square error of the difference between the advantage function estimation value A ω (s t ,a t ) and the target advantage function estimation value δ t +γλA ω (s t+1 ,a t+1 ) as the loss function. Its loss function is:

[0148]

[0149] Step 3.3: The input of the decision-making network (Actor network) is the state s t in the environmental model, and the output is the action a t ~π(a t |s t ), where π(a t |s t ) represents the probability of executing the action a t in the state s t .

[0150] When solving the policy gradient of the policy network, use the true value of the advantage function (i.e., the action value in the current state, the advantage relative to the average action value in the current state):

[0151] A π (s t ,a t ) = Qπ (s t ,a t )-V π (s t )

[0152] Among them, A π (s t ,a t ) is the true value of the advantage function, Q π (s t ,a t ) is the true value of the action-state value function, V π (s t ) is the true value of the state value function.

[0153] The parameter update expression of the decision-making network π θ (a|s) is:

[0154]

[0155] Among them, α is the learning rate of the decision-making network.

[0156] The true value of the advantage function is used in the parameter update of the above decision-making neural network, but the true value of the advantage function under the current policy cannot be accurately obtained. In actual operation, generally is used to calculate instead of the true value of the advantage function, but the state value t in δ is obtained through the evaluation network, there is a deviation, so there is a deviation in using δ t as the advantage function.

[0157] The generalized advantage function estimation method is obtained by summing δ t+l , l = 0, 1,..., ∞ at subsequent times. Therefore, by introducing the advantage function estimation network A ω (s,a) to approximate the advantage function estimation value A(s t ,a t )

[0158] So the policy gradient of the decision-making neural network is:

[0159]

[0160] Step 4: Determine the reward function;

[0161] The trajectory tracking control of the unmanned boat is that the unmanned boat starts from any position and reaches the expected trajectory with time constraints and sails along the expected trajectory with as small a tracking error as possible. Only by designing a suitable reward function to effectively evaluate the state in the environmental model at the current moment can the control goal of the unmanned boat trajectory tracking be completed.

[0162] Following the framework derived by the backstepping method, construct the coordinate transformation: z 1 = η - η d , z 2 = v - v d where η d is the desired trajectory, and υ d is the virtual control law.

[0163] For the unmanned surface vehicle mathematical model proposed in Step 1, design the virtual control law ν d based on the backstepping method to ensure that the tracking error z 1 is small enough to achieve the trajectory tracking goal. Design the Lyapunov function:

[0164]

[0165] Obviously, V 1 ≥ 0, positive definite. Take its first derivative:

[0166]

[0167] Ideally, the virtual control law υ d is designed as:

[0168]

[0169] where k 1 is a symmetric positive definite matrix.

[0170] Therefore, as long as υ → υ d , that is then η → η d , and the trajectory tracking control of the unmanned surface vehicle can be achieved.

[0171] Take k 1 = diag(k 11 , k 22 , k 33 ). The virtual control law ν c is transformed into a system of equations as:

[0172]

[0173]

[0174]

[0175] Design the reward function

[0176]

[0177] where l 1 , l 2 , l 3 are the adjustment coefficients, de is the expected trajectory tracking error.

[0178] Step Five: The controller training based on the novel Actor-Critic-Advantage network is as Figure 2 shown;

[0179] Step 5.1: Construct the network in Step Two, including five network structures: the current evaluation network the target evaluation network the generalized advantage function estimation network A ω (s,a), the current policy network π θ (a|s), the target policy network π θ′ (a|s). The current evaluation network and the target evaluation network have the same structure, and the current policy network and the target policy network have the same structure.

[0180] Step 5.2: Initialize E (the number of experience samples obtained from the experience data buffer) and the state of the unmanned boat in the model environment, and initialize the network model parameters ω,θ, copy the parameters of the current evaluation network to the target evaluation network copy the parameters of the current policy network to the target policy network θ→θ′;

[0181] Step 5.3: According to the current state s t of the unmanned boat, input it into the current policy network, and output the action a t ~π θ (a t |s t ). Input the action into the action converter to obtain the action and then input it into the environment model to obtain the state s t+1 at the next moment, and obtain the immediate reward r t based on Step Four. Finally, input the state s t+1 into the current policy network again. Loop and execute Step 5.3, and store the experience samples (s t ,a t ,r t ,s t+1 ) generated in each process into the experience data buffer.

[0182] Step 5.4: When the number of samples in the experience data buffer exceeds E, terminate the loop of Step 5.3, and select the experience samples (s i ,a i ,r i ,s i+1 ) from the experience buffer, i = 1,...,E, and input the states s i ,s i+1Input the current evaluation network and the target policy network respectively to obtain the state value Calculate And calculate the loss function for training the current evaluation network based on Step 2.1:

[0183]

[0184] Update the weight parameters of the current evaluation network based on this loss function.

[0185] Step 5.5: Input the action s i+1 into the target policy network to output the action a i+1 ~π θ (a i+1 |s i+1 ), and obtain the action a i+1 , then input s i , a i and s i+1 , a i+1 into the advantage function estimation network respectively to obtain the advantage function estimation values A ω (s i , a i ) and A ω (s i+1 , a i+1 ), and calculate the loss function for training the advantage function estimation network based on Step 2.2:

[0186]

[0187] Update the weight parameters of the advantage function estimation network based on this loss function.

[0188] Step 5.6: For training the current policy network, use the policy gradient algorithm to update the parameters of the current policy network, and derive and calculate the policy gradient according to Step 2.3:

[0189]

[0190] Step 5.7: Update the parameters of the target evaluation network and the target policy network in a soft update manner:

[0191]

[0192] θ′ = τθ + (1 - τ)θ′

[0193] where τ is the soft update coefficient, and its value range is [0, 1]

[0194] Step 5.8: Loop and execute Steps 5.3 - 5.7, observe whether the training end condition is reached. If it is reached, end the training, and the policy network generates the optimal policy Using the optimal strategy, the unmanned boat at any initial position can be controlled to complete the tracking of any desired trajectory.

[0195] In summary, the present invention has the following innovative points:

[0196] First, an advantage function estimation network is introduced into the Actor-Critic network. In step 2.2, the input of the advantage function estimation network is the state s in the environmental model t and the action a t , and the output is the estimated value A of the advantage function ω (s t , a t ). The advantage function estimation network updates according to the mean square value of the difference between the estimated value A of the advantage function ω (s t , a t ) and the target estimated value δ of the advantage function t +γλA ω (s t+1 , a t+1 ) as the loss function. Its loss function is:

[0197]

[0198] This improvement point proposes an unmanned boat trajectory tracking control method based on introducing an advantage function estimation network into the Actor-Critic network, which overcomes the problem of inaccurate advantage function estimation, that is, the large estimation deviation leading to poor training convergence effect or even non-convergence, solves the problem of low sample utilization rate in the unmanned boat trajectory tracking training process, and ensures the smoothness of the unmanned boat trajectory tracking training process.

[0199] Secondly, the output value of the advantage function estimation network replaces the advantage function value in the policy gradient; in step 2.3, the following decision neural network policy gradient is designed:

[0200]

[0201] The estimated value A of the advantage function ω (s t , a t ) replaces the advantage function value in the policy gradient in the Actor-Critic network of the generalized advantage function estimation method, so that it is easy to obtain the advantage function value at any time to participate in the calculation of the policy gradient, thereby updating the policy network.

[0202] It is no longer limited to the way of obtaining the policy gradient by taking multiple steps in the Actor-Critic network in the generalized advantage function estimation method. By making full use of the more easily obtained advantage function value to obtain the policy gradient to update the policy network, the training of the unmanned boat trajectory tracking control target is accelerated.

[0203] In addition, the reward function design of the virtual control law is introduced. In step 4, the reward function is designed as

[0204]

[0205] where l 1 , l 2 , l 3 are adjustment coefficients, and d e is the expected trajectory tracking error. The reward is described in the form of a piecewise function. The highest reward indicates that the trajectory tracking control error at the current moment meets the accuracy requirements. At moments when the accuracy requirements are not met, the value range of the reward is limited to (-1, 1]. The closer the speed and heading angular velocity of the unmanned boat are to the virtual control law, the higher the reward, ensuring that the training of the unmanned boat's trajectory tracking converges in the direction of the maximum expected discounted return.

[0206] It overcomes the difficulty that the reward function designed based on the Euclidean distance cannot ensure the stability of the unmanned boat's tracking error control. The virtual control law is introduced into the reward function, and the speed output of the trained unmanned boat tends to the virtual control law, ensuring that the unmanned boat's tracking error converges to zero. At the same time, the design of the piecewise reward function ensures that the accuracy requirements of the unmanned boat's tracking error are met.

[0207] For example: Initialize the starting position of the unmanned boat in the earth coordinate system as (2, -1), and the starting heading angle as -π / 2, that is, η 0 = [-1, 2, -π / 2] T , and the expected tracking trajectory with time constraints is:

[0208]

[0209] It is required that the unmanned boat realizes trajectory tracking control within the specified time of 30 s (as Figure 3 shown), and the trajectory error is controlled within 0.5 m. The effect diagram of the controller trained by the new Actor-Critic-Advantage network to control the unmanned boat to track the expected trajectory is as Figure 4 shown.

[0210] The above is only an individual case. More generally, the policy network trained by the unmanned boat trajectory tracking control method based on the new Actor-Critic-Advantage network can be used to control the unmanned boat to track any expected trajectory with time constraints from any position, and the tracking error is controlled within the required range, with high adaptability.

[0211] In summary, the above embodiments have described in detail different configurations of the unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network. Of course, the present invention includes but is not limited to the configurations listed in the above embodiments. Any content obtained by transformation based on the configurations provided in the above embodiments belongs to the scope protected by the present invention. Those skilled in the art can draw inferences by analogy according to the content of the above embodiments.

[0212] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0213] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes or modifications made by those of ordinary skill in the art of the present invention based on the above disclosure belong to the scope protected by the claims.

Claims

1. An unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network, characterized in that, it includes: introducing an advantage function estimation network into the Actor-Critic network to form a new Actor-Critic-Advantage network; training the new Actor-Critic-Advantage network for unmanned surface vehicle trajectory tracking control; using the single-step acquisition policy gradient method to perform unmanned surface vehicle trajectory tracking training, and obtaining the policy gradient update policy network by using the output value of the advantage function estimation network; designing a piecewise reward function based on the backstepping method to solve the virtual control law; and introducing the virtual control law into the reward function to make the speed output of the trained unmanned surface vehicle tend to the virtual control law.

2. The unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 1, characterized in that, it further includes: Step 1: Build an environmental model for unmanned surface vehicle trajectory tracking, simulate the navigation of the unmanned surface vehicle in a real marine environment with interference, and randomly generate an expected trajectory with time constraints and the starting position and heading angle of the unmanned surface vehicle; The unmanned surface vehicle trajectory tracking is independent of the vertical space. According to the established three-degree-of-freedom underactuated unmanned surface vehicle mathematical model, obtain the state information of the unmanned surface vehicle at any time in the environmental model; The kinematic model expression of the unmanned surface vehicle is: Among them, , where represents the position of the unmanned boat, represents the angle between the hull coordinate system and the earth coordinate system, that is, the heading angle of the unmanned boat; , and their symbols respectively represent the longitudinal velocity, lateral velocity and yaw angular velocity in the hull coordinate system, represents the rotation matrix for transforming the hull coordinate system to the earth coordinate system; The dynamic kinematic model expression of the unmanned surface vehicle is: Among them, , respectively represent the longitudinal thrust and turning moment of the unmanned boat in the environmental model, , respectively represent the disturbances exerted on the unmanned boat by the marine environment in the direction, respectively represent the inertia matrix, centripetal force matrix and damping matrix of the unmanned boat.

3. The unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 2, characterized in that, it further includes: Step 2: Set the action space and the state space; According to the environmental model, the state space of the unmanned boat at the current moment t is designed as , where represents the tracking trajectory position error of the unmanned boat in the geodetic coordinate system in the environmental model, represents the current heading angle and the desired heading angle, respectively represent the longitudinal speed, lateral speed and heading angular velocity of the unmanned boat in the body coordinate system; The action space of the unmanned boat at the current moment t is , representing the longitudinal thrust and steering moment of the unmanned boat in the environmental model respectively, and is converted into through the action converter and used as the input to interact with the environmental model.

4. The unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 3, characterized in that, it further includes: Step 3: Design a new Actor-Critic-Advantage network, which includes an Actor network, a Critic network and an Advantage network, corresponding to the decision-making network, the evaluation network and the advantage function estimation network respectively. Its update method is specifically as follows: Step 3.1: The input of the evaluation network is the state , and the output is the state value . The evaluation network updates using the mean squared error of the estimated state value and the target state value as the loss function. The loss function is: Among them, is the discount factor, with a value range of [0, 1], is the real-time reward obtained from the interaction between the unmanned boat and the environment, and the reward at the current moment is calculated through Step 4.

5. The unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 4, characterized in that, it further includes: Step 3.2: The estimation of the advantage function obtained according to the generalized advantage estimation method is: Among them, is the error weight coefficient, which determines the deviation magnitude of the estimator and takes values in the range [0, 1]; is the time series difference error, ; The generalized advantage function estimation method includes: by introducing the variance of the advantage function estimation, reducing the bias of the advantage function, the generalized advantage function estimation method calculates the estimated value of the advantage function, so that the bias and variance are better balanced; According to the expression of the generalized advantage function estimation, the estimated value of the advantage function under the current state and action is related to the subsequent moments. The following derivation is carried out for the generalized advantage function estimation: Obtain the recursive formula for the advantage function estimation: Construct an advantage function estimation network to approximate the generalized advantage function estimate , the advantage function estimate satisfies the following equation: The input of the advantage function estimation network is the state in the environment model and the action , and the output is the estimated value of the advantage function . The advantage function estimation network updates according to the mean square difference of the difference between the estimated value of the advantage function and the estimated value of the target advantage function as the loss function, and its loss function is: 。 6. The unmanned surface vehicle trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 5, characterized in that, it further includes: Step 3.3: The input of the decision-making network is the state in the environmental model , and the output is the action of the unmanned boat , where represents the probability of executing action in state ; using the true value of the advantage function when solving the policy gradient of the policy network: Among them, is the true value of the advantage function, is the true value of the action-state value function, is the true value of the state value function; Decision network parameters The update expression is as follows: Among them, is the learning rate of the decision network; The true value of the advantage function is used in the parameter update of the decision neural network. In actual operation, is used to calculate instead of the true value of the advantage function, The state value in is obtained through the evaluation network. There is a deviation in the advantage function with as the advantage function; The generalized advantage function estimation method is obtained by summing over subsequent time steps. Therefore, by introducing an advantage function estimation network to approximate the estimated value of the advantage function ; ; So the policy gradient of the decision neural network is: 。 7. The unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 6, characterized in that, it further includes: Step Four: Determine the reward function; The unmanned boat trajectory tracking control is that the unmanned boat starts from any position and reaches the desired trajectory with time constraints and sails along the desired trajectory with as small a tracking error as possible; The control objective of the unmanned boat trajectory tracking is completed by designing a reward function to evaluate the state in the environmental model at the current moment; Following the framework derived by the backstepping method, construct the coordinate transformation: ; where is the desired trajectory, is the virtual control law; For the unmanned surface vehicle mathematical model proposed in Step 1, design a virtual control law based on the backstepping method , to ensure that the tracking error is small enough to achieve the trajectory tracking goal; design the Lyapunov function: Obviously , it is positive definite. Take the first derivative of it: Ideally, the virtual control law is designed to be: wherein is a symmetric positive definite matrix; Guarantee That is then achieve the trajectory tracking control of the unmanned boat Take Virtual control law The conversion process equations are as follows: Design the reward function Among them, is the adjustment coefficient, is the expected trajectory tracking error.

8. The unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 7, characterized in that, it further includes: Step Five: Training of the controller based on the new Actor-Critic-Advantage network; Step 5.1: Construct the network in Step 2, including five network structures: the current evaluation network , the target evaluation network , the generalized advantage function estimation network , the current policy network , and the target policy network . The current evaluation network and the target evaluation network have the same structure, and the current policy network and the target policy network have the same structure; Step 5.2: Initialize the number of experience samples E obtained from the experience data buffer and the state of the unmanned boat in the model environment, and initialize the network model parameters with random weights , copy the parameters of the current evaluation network to the target evaluation network , copy the parameters of the current policy network to the target policy network ; Step 5.3: According to the current state of the unmanned boat Input it into the current policy network, and output an action Input the action into the action converter to obtain an action Then input it into the environment model to obtain the state at the next moment And obtain the immediate reward based on Step 4 Finally, input the state into the current policy network again; Loop and execute step 5.3 to store the experience samples generated by each process into the experience data buffer.

9. The unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 8, characterized in that, it further includes: Step 5.4: When the number of samples in the empirical data buffer exceeds E, terminate loop Step 5.3, and select empirical samples from the empirical buffer , and input the state into the current evaluation network and the target policy network respectively to obtain the state values , calculate , and calculate the loss function for training the current evaluation network based on Step 3.1: Update the weight parameters of the current evaluation network based on this loss function; Step 5.5: Input the action into the target policy network to output the action , obtaining the action . Then, input and into the advantage function estimation network respectively to obtain the advantage function estimation value , and calculate the loss function for training the advantage function estimation network based on Step 3.2: 。 10. The unmanned boat trajectory tracking control method based on the Actor-Critic-Advantage network according to claim 9, characterized in that, it further includes: Update the weight parameters of the advantage function estimation network based on this loss function; Step 5.6: For the training of the current policy network, use the policy gradient algorithm to update and train the parameters of the current policy network, and derive and calculate the policy gradient according to Step 3.3: Step 5.7: Update the parameters of the target evaluation network and the target policy network in a soft update manner: Among them is the soft update coefficient, and the value is ; Step 5.8: Loop through Steps 5.3 - 5.7 and observe whether the training end condition is met. If it is met, end the training and the policy network generates the optimal policy , and the optimal policy can be used to control the unmanned boat at any initial position to complete the tracking of any desired trajectory.

Citation Information

Patent Citations

  • Unmanned ship path tracking method based on reinforcement learning

    CN112947431A

  • Unmanned ship self-learning optimal tracking control method considering pose and speed limitation

    CN113189867A