An unmanned sailboat virtual anchoring path tracking control method

By employing adversarial inverse reinforcement learning algorithms and marine environmental data planning, the problem of path tracking accuracy of unmanned sailboats at anchor points was solved, achieving high-precision virtual anchoring control in windy and wave environments and improving the monitoring stability of unmanned sailboats in the marine environment.

CN116594383BActive Publication Date: 2025-12-05WUHAN UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310406676.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-12-05
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

The existing unmanned sailboats have low path tracking and control accuracy at anchor points, making it difficult to withstand wind and wave interference in the marine environment, resulting in unstable monitoring range and difficulty in achieving high-precision marine environmental monitoring.

Method used

By constructing an adversarial inverse reinforcement learning virtual anchoring path tracking algorithm, the desired cruise path is planned using marine environmental wind and current field data. Combined with the dynamic model of the unmanned sailboat and expert knowledge data, discriminator and generator modules are designed to achieve high-precision path tracking of the unmanned sailboat within the virtual anchoring range.

Benefits of technology

High-precision virtual anchoring control of unmanned sailboats was achieved in the marine environment, enabling stable tracking of the desired cruise path under wind and wave interference, thus improving the robustness and accuracy of path tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116594383B_ABST
    Figure CN116594383B_ABST
Patent Text Reader

Abstract

The application discloses an unmanned sailboat virtual anchoring path tracking control method, and belongs to the technical field of unmanned ship motion control, which comprises virtual anchoring range expected cruise path planning, expert knowledge data collection and processing and an adversarial reverse reinforcement learning virtual anchoring path tracking algorithm, wherein the virtual anchoring range expected cruise path planning comprises marine environment wind flow field data reading and expected cruise path planning, the expert knowledge data collection and processing comprises an expert knowledge data collection method and a data compression processing method, and the adversarial reverse reinforcement learning virtual anchoring path tracking algorithm comprises an unmanned sailboat action and state space, an adversarial reverse reinforcement learning path tracking algorithm and a path tracking reward function. Through the application, the unmanned sailboat can not only be virtually anchored within a certain range of an anchoring point, but also can overcome the influence of wind interference force and wave interference force in the marine environment, track an expected cruise path within the virtual anchoring range, and realize high-precision virtual anchoring control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned vessel motion control technology, and more specifically, relates to a virtual anchoring path tracking control method for unmanned sailboats. Background Technology

[0002] Virtual mooring path tracking control involves unmanned sailing vessels (USVs) tracking their path within a certain range of an anchorage point, ensuring they remain within the virtual mooring area. This facilitates long-term acquisition of marine environmental data and monitoring. Traditional marine environmental monitoring platforms typically involve constantly traversing anchorage targets or using monitoring buoys for fixed-point monitoring. These methods suffer from low mooring accuracy, poor resistance to wind and waves, and unstable monitoring coverage, making large-scale, high-precision marine environmental monitoring difficult. With the increasing prevalence of wind-powered USVs, their movement at sea is significantly affected by wind and wave interference, making virtual mooring path tracking control particularly challenging. Summary of the Invention

[0003] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention proposes a virtual anchoring path tracking and control method for unmanned sailboats. This method not only enables unmanned sailboats to virtually anchor within a certain range of the anchoring point, but also overcomes the influence of wind and wave interference in the marine environment, tracking the desired cruise path within the virtual anchoring range and achieving high-precision virtual anchoring control.

[0004] To achieve the above objectives, according to one aspect of the present invention, a virtual anchoring path tracking and control method for unmanned sailboats is provided, comprising:

[0005] S1: Set up a virtual anchorage circle based on the coordinates of the virtual anchorage point and the virtual anchorage radius. Read the meteorological NetCDF file to obtain marine environmental wind field and current field data. Read the wind field and current field data within the virtual anchorage circle to determine the starting point of the unmanned sailboat within the virtual anchorage circle. With the starting point of the unmanned sailboat as the center, construct a level set based on the wind field and current field data within the virtual anchorage circle to plan the desired cruise path of the unmanned sailboat in virtual anchorage.

[0006] S2: Place the unmanned sailboat in a water test field or marine environment, remotely control the unmanned sailboat through a shore-based machine to perform virtual anchoring path tracking, and collect control command data and unmanned sailboat status data. The control command data and unmanned sailboat status data are split according to data category and compressed into a dictionary-type data package along with marine environmental wind field and flow field data.

[0007] S3: Based on the performance parameters of the unmanned sailboat's control mechanism, design the unmanned sailboat's action space. Based on the quantity and type of the unmanned sailboat's state data and the ocean environment's wind field and flow field data, design the state space. Based on the unmanned sailboat's virtual anchoring expected cruise path, design the path tracking reward function based on the unmanned sailboat's virtual anchoring path tracking error and heading angle error. Design the discriminator and generator modules, and construct an adversarial inverse reinforcement learning virtual anchoring path tracking system.

[0008] In some alternative implementations, step S1 includes:

[0009] S11: Set up a virtual anchoring circle based on the virtual anchoring point coordinates and the virtual anchoring radius to determine the virtual anchoring cruising range of the unmanned sailboat, and establish a rectangular coordinate system with the virtual anchoring point coordinates as the origin;

[0010] S12: Read the meteorological NetCDF file, obtain the wind field and flow field data under the current marine environment, and read the wind field and flow field data within the virtual anchorage circle. Based on the wind field and flow field data within the virtual anchorage circle, calculate the force on the unmanned sailboat's sail, the wind interference force on the unmanned sailboat, and the wave interference force on the unmanned sailboat's hull.

[0011] S13: Input the forces on the sails of the unmanned sailboat, the wind interference forces on the unmanned sailboat, and the wave interference forces on the hull of the unmanned sailboat into the dynamic model of the unmanned sailboat to perform motion control on the unmanned sailboat;

[0012] S14: Based on the wind field and flow field data within the virtual anchorage circle, construct a level set centered on the starting point of the unmanned sailboat. The evolution velocity in the level set evolution equation is calculated from the ocean current velocity and the unmanned sailboat velocity, and plan the desired cruise path of the unmanned sailboat in the virtual anchorage.

[0013] In some alternative implementations, in step S14, by The evolution equation for the level set is obtained when v(x, y, t) = 0. When the direction of the unmanned sailboat's velocity v(x, y, t) is the direction of the level set gradient, the desired virtual anchoring cruise path of the unmanned sailboat is time-optimal. When the unmanned sailboat reaches the target point, all optimized waypoints are backtracked to obtain the time-optimal virtual anchoring cruise path of the unmanned sailboat. The backtracking equation is: Where η represents the level set function, v s (x, y, t) represents the ocean current velocity, t represents time, and (x, y) represents the coordinates at time t.

[0014] In some alternative implementations, step S2 includes:

[0015] S21: Place the unmanned sailboat in a water test area or marine environment, and remotely control the unmanned sailboat via a shore-based aircraft to perform multiple virtual anchoring path tracking tasks, collecting expert knowledge data τ E Among them, expert knowledge data τ E This includes unmanned sailboat maneuvering command data, unmanned sailboat status data, and marine environmental data;

[0016] S22: Classify expert knowledge data into action data A and state data S according to data type. Action data includes expert control commands, and state data includes unmanned sailboat state data and marine environment data. Expert control commands include the unmanned sailboat's rudder angle α at time t. r And sail angle data a s The unmanned sailboat state data includes the unmanned sailboat's forward speed u at time t. t Horizontal drift speed v t Bow roll rate r t Tracking distance error e dt , heading angle error e at The relative coordinates (x) of the virtual anchor point t y t (and marine environmental data, including the average wind speed at the location of the unmanned sailboat at time t) Wave force s f (x t y t );

[0017] S23: Categorize and compress expert knowledge data into dictionary-type data packages. Dictionary-type data includes action data and state data. The data storage format of the dictionary-type data packages is as follows:<S,A,S′> , where S′ is the state data at time t+1.

[0018] In some alternative implementations, step S3 includes:

[0019] S31: Based on the physical performance parameters of the rudder and sail of the unmanned sailboat, and based on the maximum rudder angle and sail angle, map the rudder angle and sail angle to the interval [-1, 1], set the action space dimension and value range, and determine the state space dimension based on the amount of unmanned sail state data.

[0020] S32: Based on the current position and heading of the unmanned sailboat (x) t y t , ψ t ) and virtual anchorage expected cruise path (px t py t ,pψ t ) Calculate the tracking distance error e dt and heading angle error e atThe path tracking reward function is designed based on the tracking distance error and heading angle error, x t y t , ψ t Let x and y represent the position coordinates of the unmanned sailboat at time t. t y t ) and heading ψ t , px t py t ,pψ t Let represent the expected position coordinates (px) of the unmanned sailboat at time t. t py t ) and heading pψ t ;

[0021] S33: Construct the saturation discriminant D1 θ,φ Unsaturated discriminator D2 θ,φ The discriminator's input data is dictionary-type data containing action data and state data, and its format is as follows:<S,A,S′> The discriminator module contains two sets of multilayer perceptrons (MLPs), namely the reward approximator g. θ (S, A) and shaper h φ (S), where θ are the reward approximator parameters, φ are the shaper parameters, and the output data of the discriminator module is the decoupled reward function f that evaluates the generator's output strategy. θ,φ (S, A, S′), when S′ can be uniquely determined by S and A, the discriminator has dynamic robustness, and this is the optimal discriminator;

[0022] S34: Construct the generator module. The input data for the generator module is state data S containing the state of the unmanned sailboat and the marine environment. The generator module contains two MLPs, namely the value network V. ω (S) and policy network Where ω and For network parameters, the generator module contains a generalized advantage estimation function and an action probability ratio function. The generalized advantage estimation function is calculated using the decoupled reward function output by the discriminator module. The action probability ratio function is input to the action probability of the current policy network and the action probability sampled in the buffer pool.

[0023] S35: Construct an adversarial inverse reinforcement learning path tracking algorithm. The adversarial inverse reinforcement learning path tracking algorithm includes a generator module and a discriminator module. The generator module is a policy network and the discriminator module is a decoupled reward function network. The path tracking of unmanned sailboats is achieved through iterative training.

[0024] In some alternative implementations, in step S32, the tracking distance error and heading angle error are: ||||2 represents the Euclidean norm;

[0025] The path tracking reward function is: n represents the maximum time.

[0026] In some alternative implementations, in step S33, the coupling reward function is: f θ,φ (S, A, S′)=gθ(S, A)+γh φ (S′)-h φ (S), where γ represents the shaper coefficient;

[0027] The optimal discriminator is: π(A|S) is the strategy of the generator module, D θ,φ (S, A, S′) represents the output of the discriminator;

[0028] The loss function of the optimal discriminator is: E D E π Let these represent the expectations of the discriminator and the policy, respectively.

[0029] In some alternative implementations, in step S34, the action probability function is: A t S t Let represent the motion and state of the unmanned sailboat at time t, respectively. Let represent the generator's old strategy and the probability of the unmanned sailboat choosing an action at time t, respectively.

[0030] The optimization objective of the generator's strategy is: π(A|S) represent the discriminator's output, the decoupling reward function, and the policy in state S, respectively.

[0031] In some alternative implementations, step S35 includes:

[0032] S351: Initialize expert knowledge data τ E Initialize the path tracing strategy π and two discriminators D1. θ,φ and D2 θ,φ Initialize the training time t∈{1,2,3…N}, where N is the maximum training time;

[0033] S352: Utilizing the strategy π at time t t Collect sample data τ t =(S t A t Using binary logistic regression to classify expert knowledge data τ E and sample data τ t The discriminator loss is calculated using the loss function of the optimal discriminator.

[0034] S353: Using discriminator D1 θ,φ and D2 θ,φ Calculate the return function, f θ,φ =0.5(log D1) θ,φ -log(1-D1 θ,φ ))+0.5(log D1 θ,φ -log(1-D1 θ,φ ));

[0035] S354: The reward f calculated by the discriminator θ,φ and sample τ t Optimize the strategy using the strategy optimization method, return to step S352, and iterate until time t = N.

[0036] According to another aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0037] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0038] 1. The present invention provides a virtual anchoring path tracking and control method for unmanned sailboats. The unmanned sailboats are virtually anchored within a certain range of the anchoring point, and can also overcome the influence of wind and wave interference in the marine environment. The method tracks the desired cruise path within the virtual anchoring range and achieves high-precision virtual anchoring control.

[0039] 2. Design a method for planning the desired cruise path within the virtual anchorage area. Based on ocean environmental wind and current field data, realize the optimal navigation path planning for unmanned sailboats within the virtual anchorage area.

[0040] 3. Construct an adversarial inverse reinforcement learning virtual anchoring path tracking algorithm, build a generator module and a discriminator module, and design a decoupled reward function network so that the algorithm output strategy is only related to the state, thereby improving the robustness of the algorithm and enabling the unmanned sailboat to maintain high-precision path tracking control under environmental disturbances such as larger wind, waves and currents. Attached Figure Description

[0041] Figure 1 This is a principle block diagram provided by an embodiment of the present invention;

[0042] Figure 2 This is a structural diagram of an adversarial inverse reinforcement learning virtual anchoring path tracking algorithm provided in an embodiment of the present invention;

[0043] Figure 3 This is a diagram illustrating the desired cruise path planning effect within a virtual anchorage area, provided by an embodiment of the present invention.

[0044] Figure 4 This is a diagram illustrating the effect of virtual anchoring path tracking and control for an unmanned sailboat provided in an embodiment of the present invention.

[0045] Figure 5 This is a diagram illustrating the effect of virtual anchoring path tracking and control for an unmanned sailboat, provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0047] This invention discloses a virtual anchoring path tracking and control method for unmanned sailboats, including desired cruise path planning within the virtual anchoring area, expert knowledge data collection and processing, and an adversarial inverse reinforcement learning virtual anchoring path tracking algorithm. The desired cruise path planning within the virtual anchoring area includes reading ocean environmental wind and current field data and planning the desired cruise path. The ocean environmental wind and current field data reading can read NetCDF format files, and the desired cruise path planning can plan the optimal navigation path for the unmanned sailboat within the virtual anchoring area based on the ocean environmental wind and current field data. The expert knowledge data collection and processing includes an expert knowledge data collection method and a data compression processing method. The expert knowledge data collection method acquires historical navigation status data of the unmanned sailboat remotely controlled by humans performing virtual anchoring path tracking. The data compression processing method decomposes the collected expert knowledge data into unmanned sailboat action and unmanned sailboat status data, and compresses the expert knowledge data and ocean environmental wind and current field data into a dictionary-type data file. The adversarial inverse reinforcement learning virtual anchoring path tracking algorithm includes an unmanned sailboat's action and state space, an adversarial inverse reinforcement learning algorithm, and a path tracking reward function. The action and state spaces include the unmanned sailboat's rudder and sail angles, its heading angle error, and path tracking error. The adversarial inverse reinforcement learning path tracking algorithm includes a discriminator module based on generative adversarial mechanisms and a generator module based on policy optimization. The path tracking reward function is used to evaluate the optimal navigation path for the unmanned sailboat, thereby iteratively training the adversarial inverse reinforcement learning path tracking algorithm. This invention's virtual anchoring path tracking control method for unmanned sailboats can also be extended to other surface vehicles with sail structures and multi-body mechanisms similar to unmanned sailboats. The principle block diagram of this invention is shown below. Figure 1 As shown, it includes the following steps:

[0048] S1: Set up a virtual anchorage circle based on the coordinates of the virtual anchorage point and the virtual anchorage radius. Read the meteorological NetCDF file to obtain marine environmental wind field and current field data. Read the wind field and current field data within the virtual anchorage circle to determine the starting point of the unmanned sailboat within the virtual anchorage circle. With the starting point of the unmanned sailboat as the center, construct a level set based on the wind field and current field data within the virtual anchorage circle to plan the desired cruise path of the unmanned sailboat in virtual anchorage.

[0049] S2: Place the unmanned sailboat in a water test field or marine environment, remotely control the unmanned sailboat through a shore-based machine to perform virtual anchoring path tracking, and collect control command data and unmanned sailboat status data. The control command data and unmanned sailboat status data are split according to data category and compressed into a dictionary-type data package along with marine environmental wind field and flow field data.

[0050] S3: Based on the performance parameters of the unmanned sailboat's control mechanism, design the unmanned sailboat's action space. Based on the quantity and type of the unmanned sailboat's state data and the ocean environment's wind field and flow field data, design the state space. Based on the unmanned sailboat's virtual anchoring expected cruise path, design the path tracking reward function based on the unmanned sailboat's virtual anchoring path tracking error and heading angle error. Design the discriminator and generator modules, and construct an adversarial inverse reinforcement learning virtual anchoring path tracking system.

[0051] In this embodiment of the invention, step S1 is implemented as follows:

[0052] S11: Set up a virtual anchoring circle based on the virtual anchoring point coordinates and the virtual anchoring radius to determine the virtual anchoring cruising range of the unmanned sailboat, and establish a rectangular coordinate system with the virtual anchoring point coordinates as the origin;

[0053] S12: Read the meteorological NetCDF file, obtain the wind field and flow field data under the current marine environment, and read the wind field and flow field data within the virtual anchorage circle. Based on the wind field and flow field data within the virtual anchorage circle, calculate the force on the unmanned sailboat's sail, the wind interference force on the unmanned sailboat, and the wave interference force on the unmanned sailboat's hull.

[0054] The forces acting on the sails of the unmanned sailboat are:

[0055]

[0056] X S Y S K S and N S These are the parameters representing the forces acting on the sail in the kinematic model of an unmanned sailboat. L represents the relative wind angle of the sail. s and D s These represent the lift and pull forces exerted by the wind on the sail, respectively, x s ys z s These represent the coordinates of the center of the sail's force on the axes of forward velocity u, lateral drift velocity v, and heave velocity w, respectively. They can be calculated using the following formula:

[0057]

[0058] The wind speed at the point where the sail is subjected to force. and To match the angle of attack α s The relevant lift-drag coefficients, which can be calculated through CFD simulation, A s For the sail area, ρ a This refers to air density.

[0059] The wind interference force experienced by the unmanned sailboat is:

[0060]

[0061] L is the length of the unmanned sailboat, A F A L C represents the frontal and side projected areas of the unmanned sailboat. X C Y C N H is the relative wind direction angle conversion factor. m X is the length of the lever arm of the lateral force. wind Y wind K wind N wind U represents the wind disturbance force in the direction of forward velocity, the wind disturbance force in the direction of lateral drift velocity, the torque of the wind disturbance force in the direction of forward velocity, and the torque of the wind disturbance force in the direction of settlement, respectively. aw and α aw These represent the wind speed and relative wind direction angle as projected onto the horizontal plane of the unmanned sailboat.

[0062] The wave disturbance force experienced by the hull of the unmanned sailboat is:

[0063]

[0064] s(x, y, t) represents the wave force at coordinate (x, y) at time t, B is the width of the unmanned sailboat, T is the draft of the unmanned sailboat, and ρ w Let g be the density of seawater, g be the acceleration due to gravity, χ be the difference between the heading angle and the wave angle of the unmanned sailboat, and X be the density of seawater. wave (t), Y wave (t), K wave (t), N wave(t) represent the wave disturbance forces experienced by the unmanned sailboat at time t in the direction of forward velocity, the direction of lateral drift velocity, the torque of the wave disturbance forces in the direction of forward velocity, and the torque of the wave disturbance forces in the direction of sinking, respectively. k This is the correction factor.

[0065] S13: Input the forces on the sails, wind interference forces, and wave interference forces on the hull of the unmanned sailboat calculated in the above steps into the dynamic model of the unmanned sailboat to perform motion control on the unmanned sailboat.

[0066] The dynamic model of the unmanned sailboat is as follows:

[0067]

[0068] Where m u m v m p and m r The parameters represent inertial parameters, where r and p represent the forward angular velocity and the lateral drift angular velocity, respectively, and a, b, c, and d are force restoring coefficients. and For the derivative, φ and These are the bow roll angle and the heel roll angle.

[0069] S14: Based on the wind and current field data within the virtual anchorage circle, a level set is constructed centered on the starting point of the unmanned sailboat. The evolution velocity in the level set evolution equation is determined by the ocean current velocity v. s The equations for calculating (x, y, t) and the velocity v(x, y, t) of the unmanned sailboat are as follows:

[0070]

[0071] η represents the level set function. When the direction of v(x, y, t) is the direction of the level set gradient, the path is time-optimal. When the unmanned sailboat reaches the target point, backtracking through all optimized waypoints yields the time-optimal path. The backtracking equation is:

[0072]

[0073] In this embodiment of the invention, step S2 is implemented as follows:

[0074] S21: Place the unmanned sailboat in a water test area or marine environment, and remotely control the unmanned sailboat via a shore-based aircraft to perform multiple virtual anchoring path tracking tasks, collecting expert knowledge data τ E Among them, expert knowledge data τ E This includes unmanned sailboat maneuvering command data, unmanned sailboat status data, and marine environmental data;

[0075] S22: Classify expert knowledge data into action data A and state data S according to data type. Action data includes expert control commands, and state data includes unmanned sailboat state data and marine environment data. Expert control commands include the unmanned sailboat's rudder angle α at time t. r And sail angle data a s The unmanned sailboat state data includes the unmanned sailboat's forward speed u at time t. t Horizontal drift speed v t Bow roll rate r t Tracking distance error e dt , heading angle error e at The relative coordinates (x) of the virtual anchor point t y t (and marine environmental data, including the average wind speed at the location of the unmanned sailboat at time t) Wave force s f (x t y t );

[0076] In this embodiment of the invention, the action data format is shown in Table 1 below:

[0077] Table 1

[0078]

[0079] The status data format is shown in Table 2 below:

[0080] Table 2

[0081]

[0082] S23: Using Python-based torch / numpy / pandas / xlwings / openpyxl function packages, categorize, compress, and save expert knowledge data into dictionary-type data packages. The file extensions for these dictionary-type data packages are .pth / .npy / .csv / .xls / .xlsx. The dictionary-type data includes action data and state data. The data storage format for these dictionary-type data packages is as follows:<S,A,S′> , where S′ is the state data at time t+1.

[0083] In this embodiment of the invention, step S3 is implemented as follows:

[0084] S31: Based on the physical performance parameters of the rudder and sail of the unmanned sailboat, and based on the maximum rudder angle and sail angle, map the rudder angle and sail angle to the interval [-1, 1], set the action space dimension and value range, and determine the state space dimension based on the amount of unmanned sail state data.

[0085] S32: Based on the current position and heading of the unmanned sailboat (x)t y t , ψ t ) and virtual anchorage expected cruise path (px t py t ,pψ t ) Calculate the tracking distance error e dt and heading angle error e at The path tracking reward function is designed based on the tracking distance error and heading angle error, x t y t , ψ t px represents the position coordinates and heading of the unmanned sailboat at time t, respectively. t py t ,pψ t Let represent the expected position coordinates and heading of the unmanned sailboat at time t, respectively.

[0086] Among them, the tracking distance error and the heading angle error are:

[0087]

[0088] ||||2 represents the Euclidean norm.

[0089] The path tracking reward function is:

[0090]

[0091] n Indicates the maximum time.

[0092] S33: Construct the discriminator module, which contains two discriminator modules: a saturation discriminator D1. θ,φ Unsaturated discriminator D2 θ,φ The input data for the discriminator module is dictionary-type data containing action data and state data, and its format is as follows:<S,A,S′> Each discriminator module contains two sets of multi-layer perceptrons (MLPs), namely the reward approximator g. θ (S, A) and shaper h φ (S), where θ are the reward approximator parameters, φ are the shaper parameters, and the output data of the discriminator module is the decoupled reward function f that evaluates the generator's output strategy. θ,φ (S, A, S′), when S′ can be uniquely determined by S and A, the discriminator has dynamic robustness, and this is the optimal discriminator;

[0093] The decoupling reward function is as follows:

[0094] f θ,φ (S, A, S′) = g θ (S, A) + γhφ (S′)-h φ (S) (10)

[0095] γ represents the shaper coefficient.

[0096] The optimal discriminator is:

[0097]

[0098] Where π(A|S) is the strategy of the generator module, and D θ,φ (S, A, S′) represents the output of the discriminator module.

[0099] The loss function of the optimal discriminator is:

[0100]

[0101] E D E π Let these represent the expectations of the discriminator and the policy, respectively.

[0102] S34: Construct the generator module. The input data for the generator module is state data S containing the state of the unmanned sailboat and the marine environment. The generator module contains two MLPs, namely the value network V. ω (S) and policy network Where ω and For network parameters, the generator module contains a generalized advantage estimation function and an action probability ratio function. The generalized advantage estimation function is calculated using the decoupled reward function output by the discriminator module. The action probability ratio function is input to the action probability of the current policy network and the action probability sampled in the buffer pool.

[0103] The action probability function is:

[0104]

[0105] A t S t Let represent the motion and state of the unmanned sailboat at time t, respectively. Let represent the generator's old strategy and the probability of the unmanned sailboat choosing an action at time t, respectively.

[0106] The optimization objective of the generator's strategy is:

[0107]

[0108] D θ,φ (S, A), f θ,φ (S, A) and π(A|S) represent the discriminator's output, the decoupling reward function, and the policy in state S, respectively.

[0109] S35: Construct an adversarial inverse reinforcement learning path tracking algorithm. The adversarial inverse reinforcement learning path tracking algorithm includes a generator module and a discriminator module. The generator module is a policy network and the discriminator module is a decoupled reward function network. The path tracking of unmanned sailboats is achieved through iterative training.

[0110] The adversarial inverse reinforcement learning path tracking algorithm flow is as follows: Figure 2 As shown, it includes the following steps:

[0111] S351: Initialize expert knowledge data τ E Initialize the path tracing strategy π and two discriminators D1. θ,φ and D2 θ,φ Initialize the training time t∈{1,2,3…N}, where N is the maximum training time.

[0112] S352: Utilizing the strategy π at time t t Collect sample data τ t =(S t A t Using binary logistic regression to classify expert knowledge data τ E and sample data τ t The discriminator loss is calculated using the loss function of the optimal discriminator described in step S33.

[0113] S353: Using discriminator D1 θ,φ and D2 θ,φ Calculate the return function, f θ,φ =0.5(log D1) θ,φ -log(1-D1 θ,φ ))+0.5(log D1 θ,φ -log(1-D1 θ,φ )).

[0114] S354: The reward f calculated by the discriminator θ,φ and sample τ t Optimize the strategy using the strategy optimization method, return to step S352, and iterate until time t = N.

[0115] from Figure 3 It can be seen that the horizontal set path planning method in step S14 plans the shortest path downstream.

[0116] from Figure 4 As can be seen, the virtual anchoring path tracking and control method for unmanned sailboats involved in this invention is effective in controlling the unmanned sailboat to navigate stably within the virtual anchoring area.

[0117] from Figure 5It can be seen that the adversarial inverse reinforcement learning virtual anchoring path tracking effect involved in step S3 is relatively small in both virtual anchoring path tracking distance error and heading error.

[0118] It should be noted that, depending on the implementation needs, the various steps / components described in this application can be broken down into more steps / components, or two or more steps / components or parts of the operation of steps / components can be combined into new steps / components to achieve the purpose of this invention.

[0119] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A virtual anchoring path tracking and control method for an unmanned sailboat, characterized in that, include: S1: Set up a virtual anchorage circle based on the coordinates of the virtual anchorage point and the virtual anchorage radius. Read the meteorological NetCDF file to obtain marine environmental wind field and current field data. Read the wind field and current field data within the virtual anchorage circle to determine the starting point of the unmanned sailboat within the virtual anchorage circle. With the starting point of the unmanned sailboat as the center, construct a level set based on the wind field and current field data within the virtual anchorage circle to plan the desired cruise path of the unmanned sailboat in virtual anchorage. S2: Place the unmanned sailboat in a water test field or marine environment, remotely control the unmanned sailboat through a shore-based machine to perform virtual anchoring path tracking, and collect control command data and unmanned sailboat status data. The control command data and unmanned sailboat status data are split according to data category and compressed into a dictionary-type data package along with marine environmental wind field and flow field data. S3: Based on the performance parameters of the unmanned sailboat control mechanism, design the unmanned sailboat action space. Based on the quantity and type of unmanned sailboat state data and marine environmental wind field and flow field data, design the state space. Based on the expected cruise path of the unmanned sailboat virtual anchoring, design the path tracking reward function based on the unmanned sailboat virtual anchoring path tracking error and heading angle error. Design the discriminator and generator modules, and construct adversarial inverse reinforcement learning virtual anchoring path tracking. The decoupling reward function is: f θ,φ (S,A,S′)=g θ (S,A)+γh φ (S′)-h φ (S), where γ represents the shaper coefficient, g θ (S,A) represents the return approximator, h φ (S) represents the shaper, θ represents the reward approximator parameter, and φ represents the shaper parameter.<S,A,S′> The input data for the discriminator consists of dictionary-type data containing action data and state data. The optimal discriminator is: π(A|S) is the strategy of the generator module, D θ,φ (S,A,S′) represents the output of the discriminator; The loss function of the optimal discriminator is: E D E π Let D1 represent the expectation of the discriminator and the expectation of the policy, respectively. θ,φ D2 is a saturation discriminator. θ,φ It is a non-saturated discriminator.

2. The method according to claim 1, characterized in that, Step S1 includes: S11: Set up a virtual anchoring circle based on the virtual anchoring point coordinates and the virtual anchoring radius to determine the virtual anchoring cruising range of the unmanned sailboat, and establish a rectangular coordinate system with the virtual anchoring point coordinates as the origin; S12: Read the meteorological NetCDF file, obtain the wind field and flow field data under the current marine environment, and read the wind field and flow field data within the virtual anchorage circle. Based on the wind field and flow field data within the virtual anchorage circle, calculate the force on the unmanned sailboat's sail, the wind interference force on the unmanned sailboat, and the wave interference force on the unmanned sailboat's hull. S13: Input the forces on the sails of the unmanned sailboat, the wind interference forces on the unmanned sailboat, and the wave interference forces on the hull of the unmanned sailboat into the dynamic model of the unmanned sailboat to perform motion control on the unmanned sailboat; S14: Based on the wind field and flow field data within the virtual anchorage circle, construct a level set centered on the starting point of the unmanned sailboat. The evolution velocity in the level set evolution equation is calculated from the ocean current velocity and the unmanned sailboat velocity, and plan the desired cruise path of the unmanned sailboat in the virtual anchorage.

3. The method according to claim 2, characterized in that, In step S14, by The evolution equation of the level set is obtained. When the direction of the unmanned sailboat's velocity v(x,y,t) is the direction of the level set gradient, the desired virtual anchoring cruise path of the unmanned sailboat is time-optimal. When the unmanned sailboat reaches the target point, all optimized waypoints are backtracked to obtain the time-optimal virtual anchoring cruise path of the unmanned sailboat. The backtracking equation is as follows: Where η represents the level set function, v s (x,y,t) represents the ocean current velocity, t represents time, and (x,y) represents the coordinates at time t.

4. The method according to claim 1, characterized in that, Step S2 includes: S21: Place the unmanned sailboat in a water test area or marine environment, and remotely control the unmanned sailboat via a shore-based aircraft to perform multiple virtual anchoring path tracking tasks, collecting expert knowledge data τ E Among them, expert knowledge data τ E This includes unmanned sailboat maneuvering command data, unmanned sailboat status data, and marine environmental data; S22: Classify expert knowledge data into action data A and state data S according to data type. Action data includes expert control commands, and state data includes unmanned sailboat state data and marine environment data. Expert control commands include the unmanned sailboat's rudder angle α at time t. r And sail angle data a s The unmanned sailboat state data includes the unmanned sailboat's forward speed u at time t. t Horizontal drift speed v t Bow roll rate r t Tracking distance error e dt , heading angle error e at The relative coordinates (x) of the virtual anchor point t ,y t (and marine environmental data, including the average wind speed at the location of the unmanned sailboat at time t) Wave force s f (x t ,y t ); S23: Categorize and compress expert knowledge data into dictionary-type data packages. Dictionary-type data includes action data and state data. The data storage format of the dictionary-type data packages is as follows:<S,A,S′> , where S′ is the state data at time t+1.

5. The method according to claim 4, characterized in that, Step S3 includes: S31: Based on the physical performance parameters of the rudder and sail of the unmanned sailboat, and based on the maximum rudder angle and sail angle, map the rudder angle and sail angle to the interval [-1,1], set the action space dimension and value range, and determine the state space dimension based on the amount of unmanned sail state data; S32: Based on the current position and heading of the unmanned sailboat (x) t ,y t ,ψ t ) and virtual anchorage expected cruise path (px t ,py t ,pψ t ) Calculate the tracking distance error e dt and heading angle error e at The path tracking reward function is designed based on the tracking distance error and heading angle error, x t ,y t ,ψ t Let x and y represent the position coordinates of the unmanned sailboat at time t. t ,y t ) and heading ψ t , px t ,py t ,pψ t Let represent the expected position coordinates (px) of the unmanned sailboat at time t. t ,py t ) and heading pψ t ; S33: Construct the saturation discriminant D1 θ,φ Unsaturated discriminator D2 θ,φ The discriminator's input data is dictionary-type data containing action data and state data, and its format is as follows:<S,A,S′> The discriminator module contains two sets of multilayer perceptrons (MLPs), namely the reward approximator g. θ (S,A) and shaper h φ (S), where θ are the reward approximator parameters, φ are the shaper parameters, and the output data of the discriminator module is the decoupled reward function f that evaluates the generator's output strategy. θ,φ (S,A,S′), when S′ can be uniquely determined by S and A, the discriminator has dynamic robustness, and this is the optimal discriminator; S34: Construct the generator module. The input data for the generator module is state data S containing the state of the unmanned sailboat and the marine environment. The generator module contains two MLPs, namely the value network V. ω (S) and policy network Where ω and For network parameters, the generator module contains a generalized advantage estimation function and an action probability ratio function. The generalized advantage estimation function is calculated using the decoupled reward function output by the discriminator module. The action probability ratio function is input to the action probability of the current policy network and the action probability sampled in the buffer pool. S35: Construct an adversarial inverse reinforcement learning path tracking algorithm. The adversarial inverse reinforcement learning path tracking algorithm includes a generator module and a discriminator module. The generator module is a policy network and the discriminator module is a decoupled reward function network. The path tracking of unmanned sailboats is achieved through iterative training.

6. The method according to claim 5, characterized in that, In step S32, the tracking distance error and heading angle error are: ‖‖2 represents the Euclidean norm; The path tracking reward function is: n represents the maximum time.

7. The method according to claim 6, characterized in that, In step S34, the action probability function is: A t S t Let represent the motion and state of the unmanned sailboat at time t, respectively. Let represent the generator's old strategy and the probability of the unmanned sailboat choosing an action at time t, respectively. The optimization objective of the generator's strategy is: D θ,φ (S,A), f θ,φ (S,A) and π(A|S) represent the discriminator's output, the decoupling reward function, and the policy in state S, respectively.

8. The method according to claim 7, characterized in that, Step S35 includes: S351: Initialize expert knowledge data τ E Initialize the path tracing strategy π and two discriminators D1. θ,φ and D2 θ,φ Initialize the training time t∈{1,2,3…N}, where N is the maximum training time; S352: Utilizing the strategy π at time t t Collect sample data τ t =(S t A t Using binary logistic regression to classify expert knowledge data τ E and sample data τ t The discriminator loss is calculated using the loss function of the optimal discriminator. S353: Using discriminator D1 θ,φ and D2 θ,φ Calculate the return function, f θ,φ =0.5(logD1) θ,φ -log(1-D1 θ,φ ))+0.5(logD1 θ,φ -log(1-D1 θ,φ )); S354: The reward f calculated by the discriminator θ,φ and sample τ t Optimize the strategy using the strategy optimization method, return to step S352, and iterate until time t = N.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Autonomous collision avoidance decision-making method for unmanned ship based on adaptive navigation situation learning

    CN109298712A

  • Unmanned ship formation path tracking method based on deep reinforcement learning

    CN111694365A