An intelligent action planning and trajectory control method for hovercraft based on reinforcement learning

Through the intelligent action planning and track control method of hovercraft based on reinforcement learning, decoupling the heading and lateral control channels, and designing an intelligent adaptive linear self-immune controller, the problem of high maneuverability and poor stability during the navigation of hovercraft is solved, and the safe and stable navigation of hovercraft is achieved.

CN119937570BActive Publication Date: 2025-07-04SHANGHAI ZHONGCHUAN SDT-NERC CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510421443.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-04
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

During the navigation process, due to the strong coupling of the heading and lateral control channels, the maneuverability is difficult and the navigation stability is affected by the wind speed and direction, which makes it prone to side drifting movement, increasing the risk of safety accidents.

Method used

The intelligent action planning and track control method of hovercraft based on reinforcement learning is adopted. By establishing a fully pad lift hovercraft model, the speed and heading-lateral decoupling control model is designed, the depth deterministic strategy gradient algorithm is used to plan the hovercraft navigation actions, and an intelligent adaptive linear self-immune controller is designed to realize the stable navigation of the hovercraft.

Benefits of technology

Effectively decoupling heading and lateral control improves the navigation stability and safety of hovercraft, reduces the dangerous navigation state of manual operation, and ensures the safe navigation of hovercraft.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937570B_ABST
    Figure CN119937570B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent motion planning and trajectory control method for an air-cushion vehicle based on reinforcement learning, which determines the motion environment information of the surface-effect air-cushion vehicle; establishes a model of the surface-effect air-cushion vehicle to obtain a decoupled control model of the speed, heading-lateral displacement; obtains the desired speed, desired heading and desired lateral position of the air-cushion vehicle; designs an air-cushion vehicle speed controller based on intelligent adaptive linear active disturbance rejection control, and designs an air-cushion vehicle heading-lateral decoupled controller based on the intelligent adaptive linear active disturbance rejection decoupling control algorithm to ensure the stability of the heading and lateral position of the air-cushion vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ship intelligent decision-making and motion control, and particularly to an intelligent action planning and track control method for an air-cushion vehicle based on reinforcement learning. Background Art

[0002] Due to the amphibious and high-speed characteristics of air-cushion vehicles, since air-cushion vehicles do not have underwater turning equipment and cannot turn freely, the maneuvering difficulty of air-cushion vehicles is greatly increased, and their navigation stability is greatly affected by wind speed and direction. During navigation, lateral drift motion is likely to occur, leading to safety accidents of air-cushion vehicles. Ship intelligence can effectively solve the problems faced by ships in terms of safety, energy efficiency, and cost, and is an inevitable trend in ship development. Implementing intelligent navigation according to the manual operation rules and corresponding motion control methods for special operation requirements will greatly reduce the dangerous navigation states that occur in manual operations and ensure the safe navigation of air-cushion vehicles. Summary of the Invention

[0003] In view of the above problems existing in the prior art, the present invention provides an intelligent action planning and track control method for an air-cushion vehicle based on reinforcement learning, including:

[0004] Step 1: Determine the motion environment information of a fully air-cushioned air-cushion vehicle; establish a model of the fully air-cushioned air-cushion vehicle to obtain the heading, position, speed, and roll of the air-cushion vehicle.

[0005] Step 2: Simplify the strong coupling problem between the heading control channel and the lateral control channel by using the heading, position, speed, and roll of the air-cushion vehicle obtained in Step 1 to obtain a speed control model and a heading-lateral displacement decoupling control model.

[0006] Step 3: Obtain the navigation information of the air-cushion vehicle according to the speed control model and the heading-lateral displacement decoupling control model in Step 2 and the fully air-cushioned air-cushion vehicle model in Step 1, and plan the intelligent actions during the navigation process of the air-cushion vehicle through the deep deterministic policy gradient in the reinforcement learning algorithm to obtain the desired speed, desired heading, and desired lateral position of the air-cushion vehicle.

[0007] Step 4: Design an air-cushion vehicle speed controller based on intelligent adaptive linear active disturbance rejection control according to the desired speed in Step 3 and the speed control model in Step 2, and rely on the thrust generated by the variable pitch air propeller to control the longitudinal speed to ensure the stable speed of the air-cushion vehicle.

[0008] Step 5: Design an air-cushion vehicle heading-lateral decoupling controller according to the desired heading and desired lateral position in Step 3 and the heading-lateral displacement decoupling control model in Step 2, and based on the intelligent adaptive linear active disturbance rejection decoupling control algorithm to ensure the stable heading and lateral position of the air-cushion vehicle.

[0009] In Step 1, the hovercraft model includes a kinematic model and a dynamic model;

[0010] The kinematic model is as follows:

[0011] ;

[0012] where, , , , respectively represent the northward position, eastward position, roll angle, and heading angle of the hovercraft in the north-east coordinate system, , , , respectively represent the longitudinal velocity, lateral velocity, roll angular velocity, and heading angular velocity of the hovercraft in the moving coordinate system.

[0013] The dynamic model described in Step 1 is as follows:

[0014] ;

[0015] where, is the mass of the hovercraft, , , , are successively the longitudinal resultant force, lateral resultant force, roll resultant moment, and yaw resultant moment acting on the hovercraft, , are successively the moments of inertia of the hovercraft about the x and z axes; , , , respectively represent the longitudinal velocity, lateral velocity, roll angular velocity, and heading angular velocity of the hovercraft in the moving coordinate system.

[0016] Step 2 specifically includes:

[0017] Step 2-1: Since the adaptive linear active disturbance rejection control is a model-free control method, the standard form of the hovercraft speed control model that can be directly obtained is:

[0018] ;

[0019] where, is the longitudinal velocity of the hovercraft in the moving coordinate system, is the total system disturbance, is the control gain, is the control quantity;

[0020] Step 2-2: From the hovercraft model established in Step 1, the heading control model and the lateral displacement control model are obtained as follows:

[0021] ;

[0022] ;

[0023] In the formula, is the yaw aerodynamic moment, is the yaw air momentum moment, is the yaw hydrodynamic moment, is the yaw moment of the air rudder, is the yaw moment of the air propeller, is the yaw moment of the bow nozzle, is the lateral aerodynamic force, is the lateral air momentum force, is the lateral hydrodynamic force, is the lateral force of the air rudder, is the lateral thrust of the air propeller, is the lateral thrust of the bow nozzle, is the moment of inertia of the hovercraft about the z axis, , represent the roll angle and yaw angle of the hovercraft in the north-east coordinate system, m is the mass of the hovercraft, , , , y respectively represent the longitudinal speed, lateral speed, yaw angular velocity, and lateral position of the hovercraft in the moving coordinate system.

[0024] Furthermore, in step 2, the yaw control actuator and the lateral displacement control actuator are the air rudder and the bow nozzle respectively. Define the yaw control input as the yaw moment generated by the air rudder, and define the lateral displacement control input as the lateral force generated by the bow nozzle. The mathematical form is expressed as follows:

[0025] ;

[0026] In the formula, is the lateral rudder force coefficient, is the air density, is the oncoming flow velocity received by the air rudder, is the area of the air rudder, is the longitudinal installation position, is the distance from the pressure center received by the air rudder to the rudder shaft, is the thrust of the bow nozzle, is the nozzle angle;

[0027] Define the lateral force generated by the air rudder as the interference of the yaw control channel on the lateral control channelτ uψ , define the turning moment generated by the bow nozzle as the interference of the lateral control channel on the heading control channel τ uy , which is expressed in mathematical form as follows:

[0028] ;

[0029] In the formula, is the installation position of the bow nozzle.

[0030] The heading decoupling control model in step 2 is:

[0031] ;

[0032] In the formula, , , , , , are the derived base values for calculating the non-dimensional hydrodynamic coefficients in the heading and lateral directions, is the longitudinal position where the hydrodynamic force or moment acts on the hull, , are the roll angle and heading angle of the hovercraft in the north-east coordinate system respectively, is the moment of inertia of the hovercraft about the z-axis, , are the lateral velocity and heading angular velocity of the hovercraft in the moving coordinate system respectively, is the turning moment generated by the air rudder, τ uy is the interference of the lateral control channel on the heading control channel, is the disturbance of the heading control channel;

[0033] The lateral displacement decoupling control model is:

[0034] ;

[0035] In the formula, is the disturbance of the lateral control channel, u y is the lateral displacement control input, y, respectively represent the lateral position, lateral velocity and heading angular velocity of the hovercraft in the moving coordinate system, m is the mass of the hovercraft, , , are the derived base values for calculating the non-dimensional hydrodynamic coefficients in the lateral direction, is the interference of the heading control channel on the lateral control channel.

[0036] The specific content in step 3 includes:

[0037] Step 3-1: Design the state space of all decision-making information affecting the hovercraft and the action space of all action sets of the hovercraft;

[0038] Step 3-2: Design the reward function: During the intelligent navigation of the hovercraft, the learning goal is for the hovercraft to approach the desired trajectory;

[0039] Step 3-3: Design the termination condition: Through experience and testing, design the termination step count to be 2000, and end this training when this step count is exceeded.

[0040] The hovercraft speed controller in the said Step 4 is specifically designed as:

[0041] ;

[0042] In the formula, is the control quantity, is the control gain, is the controller bandwidth, is the desired speed, 、 are the observation outputs of the speed observer.

[0043] The hovercraft heading-lateral decoupling controller in the said Step 5 is specifically designed as:

[0044] Step 5-1: Simplify the heading decoupling control model in Step 2:

[0045] ;

[0046] In the formula, is the virtual control quantity of the heading control channel, , where 、 、 are time-varying coefficients;

[0047] Step 5-2: Simplify the lateral displacement decoupling control model in Step 2:

[0048] ;

[0049] In the formula, is the virtual control quantity of the lateral control channel, , where 、 、 are time-varying coefficients;

[0050] Step 5-3: Combine the control models in Steps 5-1 and 5-2 into a matrix mode to obtain a bow-lateral decoupling controller for the bow control channel and the lateral displacement control channel of the hovercraft.

[0051] The bow-lateral decoupling controller for the bow control channel and the lateral displacement control channel of the hovercraft is:

[0052] ;

[0053] where is the desired bow error, is the desired bow direction, , , are the observation outputs of the bow observer, , , , , is the bow error feedback gain vector, is the virtual control quantity of the bow control channel, is the initial value of the virtual control quantity of the bow controller, is the bow control quantity;

[0054] ;

[0055] where is the desired lateral position error, , , are the observation outputs of the lateral position observer, , , , , is the lateral position error feedback gain vector, is the virtual control quantity of the lateral displacement control channel, is the initial value of the virtual control quantity of the lateral position controller, is the lateral position control quantity, is the desired lateral displacement.

[0056] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0057] 1) In the present invention, a bow-lateral displacement decoupling control model based on the strong coupling between the simplified bow control channel and the lateral control channel is proposed, clarifying the coupling relationship between the bow and the lateral displacement, providing a basis for the design of the subsequent decoupling controller;

[0058] 2) In the present invention, a method for intelligent motion planning design of an air-cushion vehicle based on the DDPG algorithm is proposed. The deep neural network used in this method greatly enhances the feature extraction ability. The gradient policy algorithm can effectively handle the continuous action space problem of the air-cushion vehicle. The experience replay technique and the added target network can improve the training stability and convergence speed, obtain the real-time expected navigation information of the air-cushion vehicle according to the operation requirements, and make motion planning;

[0059] 3) In the present invention, an intelligent adaptive linear active disturbance rejection method based on the BP neural network is proposed, which realizes the online adjustment of key parameters in active disturbance rejection control, greatly reduces the time for adjusting parameters, and also enables the controller to more accurately observe and compensate for disturbances, effectively avoiding the disadvantages of offline training complexity and non-real-time performance, improving the control accuracy and rapidity of the speed, heading-lateral decoupling controller, and enhancing the navigation safety of the air-cushion vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 Schematic diagram of the intelligent motion planning and trajectory control method for a surface-effect ship;

[0062] Figure 2 Network structure diagram of the deep deterministic policy gradient algorithm;

[0063] Figure 3 Schematic diagram of the intelligent adaptive linear active disturbance rejection speed controller;

[0064] Figure 4 Schematic diagram of the intelligent adaptive linear active disturbance rejection heading-lateral decoupling controller. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The present invention proposes an intelligent motion planning and trajectory control method for an air-cushion vehicle based on reinforcement learning, including:

[0066] Step 1: Determine the motion environment information of the surface-effect ship, mainly including the navigation information of the air-cushion vehicle. Establish a surface-effect ship model to obtain the heading, position, speed, and roll of the air-cushion vehicle; wherein, the air-cushion vehicle model includes a kinematic model and a dynamic model;

[0067] Step 1-1: The kinematic model of the air-cushion vehicle is as follows:

[0068] ;

[0069] In the formula, , , , respectively represent the northward position, eastward position, roll angle and heading angle of the hovercraft in the northeast coordinate system. , , , respectively represent the longitudinal velocity, lateral velocity, roll angular velocity and heading angular velocity of the hovercraft in the moving coordinate system.

[0070] Step 1-2: The dynamic model of the hovercraft is as follows:

[0071] ;

[0072] In the formula, is the mass of the hovercraft, , , , are successively the longitudinal resultant force, lateral resultant force, roll resultant moment and yaw resultant moment acting on the hovercraft, , are successively the moments of inertia of the hovercraft about x , z axes; , , , respectively represent the longitudinal velocity, lateral velocity, roll angular velocity and heading angular velocity of the hovercraft in the moving coordinate system.

[0073] Step 2: Simplify the strong coupling problem existing between the heading control channel and the lateral control channel by using the heading, position, speed and roll of the hovercraft obtained in Step 1 to obtain a speed control model and a heading-lateral displacement decoupling control model;

[0074] Specifically, Step 2 specifically includes:

[0075] Step 2-1: Since the adaptive linear active disturbance rejection control is a model-free control method, the standard form of the speed control model of the hovercraft can be directly obtained as:

[0076] ;

[0077] In the formula, is the longitudinal velocity of the hovercraft in the moving coordinate system, is the total system disturbance, is the control gain, is the control quantity.

[0078] Step 2-2: Based on the hovercraft model established in Step 1, the heading control model and the lateral displacement control model are obtained as follows:

[0079] ;

[0080] ;

[0081] In the formula, is the yaw aerodynamic moment, is the yaw air momentum moment, is the yaw hydrodynamic moment, is the yaw moment generated by the air rudder, is the yaw moment generated by the air propeller, is the yaw moment of the bow nozzle, is the lateral aerodynamic force, is the lateral air momentum force, is the lateral hydrodynamic force, is the lateral force of the air rudder, is the lateral thrust of the air propeller, is the lateral thrust of the bow nozzle, is the moment of inertia of the hovercraft about the z axis, , represent the roll angle and the heading angle of the hovercraft in the north-east coordinate system, m is the mass of the hovercraft, , , , y respectively represent the longitudinal velocity, the lateral velocity, the heading angular velocity, and the lateral position of the hovercraft in the moving coordinate system.

[0082] Since the heading control actuator and the lateral displacement control actuator are the air rudder and the bow nozzle respectively, the heading control input is defined as the yaw moment generated by the air rudder, and the lateral displacement control input is defined as the lateral force generated by the bow nozzle. The mathematical forms are as follows:

[0083] ;

[0084] In the formula, is the lateral rudder force coefficient, is the air density, is the oncoming flow velocity received by the air rudder, is the area of the air rudder, is the longitudinal installation position, is the distance from the pressure center received by the air rudder to the rudder shaft, is the thrust of the bow nozzle, is the nozzle angle.

[0085] The air rudder will also generate a lateral force, and the bow nozzle will also generate a yawing moment. Therefore, the mutual interference between the two control channels needs to be considered. The lateral force generated by the air rudder is defined as the interference of the heading control channel on the lateral control channel τ uψ , and the yawing moment generated by the bow nozzle is defined as the interference of the lateral control channel on the heading control channel τ uy , and the mathematical form is expressed as follows:

[0086] ;

[0087] In the formula, is the installation position of the bow nozzle.

[0088] Then the heading decoupling control model is:

[0089] ;

[0090] In the formula, , , , , , are the derived base values for calculating the non-dimensional hydrodynamic coefficients in the heading and lateral directions, is the longitudinal position where the hydrodynamic force or moment acts on the hull, , are the roll angle and heading angle of the hovercraft in the north-east coordinate system respectively, is the moment of inertia of the hovercraft about the z-axis, , are the lateral velocity and heading angular velocity of the hovercraft in the moving coordinate system respectively, is the yawing moment generated by the air rudder, τ uy is the interference of the lateral control channel on the heading control channel, is the disturbance of the heading control channel;

[0091] Then the lateral displacement decoupling control model is:

[0092] ;

[0093] In the formula, is the disturbance of the lateral control channel, u y is the lateral displacement control input, y, , respectively represent the lateral position, lateral velocity and heading angular velocity of the hovercraft in the moving coordinate system, m is the mass of the hovercraft, , , For calculating the derived base value of the lateral non-dimensional hydrodynamic coefficient, it is the interference of the heading control channel on the lateral control channel.

[0094] Step 3: Obtain the hovercraft navigation information according to the speed control model in Step 2, the heading-lateral decoupling control model, and the surface-effect ship model in Step 1. Design an intelligent action planning method for the hovercraft navigation process through the Deep Deterministic Policy Gradient (DDPG) in the reinforcement learning algorithm. The algorithm network structure is implemented using the Actor-Critic algorithm. The specific structure is as Figure 2 shown, mainly including the design or selection of the state space, action space, reward function, network structure, and termination condition, to obtain the desired speed, desired heading, and desired lateral position of the hovercraft;

[0095] The specific content of Step 3 includes:

[0096] Step 3-1: Design the state space that affects all the hovercraft decision-making information and the action space of all the hovercraft actions: In the scenario of the present invention, the required state information mainly includes the hovercraft attitude given for special operation requirements, which is represented by the vector S as:

[0097] ;

[0098] In the formula, the elements in the vector S x', u', y', v' , ', r' , ', p' are successively the longitudinal position, longitudinal speed, lateral position, lateral speed, heading angle, heading angular speed, roll angle, and roll angular speed of the hovercraft under special operation requirements, and are all continuous variables.

[0099] During the operation and navigation of the hovercraft, the actions should be designed as the desired speed, desired lateral displacement, and desired heading angle of the hovercraft. Therefore, 3 action information is selected to form the action space, which is represented by the vector A as:

[0100] ;

[0101] In the formula, 、 、 are respectively the desired speed, desired lateral displacement, and desired heading angle of the hovercraft.

[0102] Step 3-2: Design the reward function: During the intelligent navigation of the hovercraft, the goal of learning is for the hovercraft to approach the desired trajectory.

[0103] Apply a positive incentive for approaching the desired position, defined as follows:

[0104] ;

[0105] In the formula, is the relative longitudinal distance between the hovercraft and the desired position at the current moment, is a reward coefficient greater than 0, is the relative lateral distance between the hovercraft and the desired position at the current moment, is a reward coefficient greater than 0, r x is the longitudinal distance reward function, r y is the lateral distance reward function.

[0106] Apply a positive incentive for approaching the desired heading direction, defined as follows:

[0107] ;

[0108] In the formula, is the relative heading angle between the hovercraft and the desired heading at the current moment, is a reward coefficient greater than 0, is the heading reward function.

[0109] For the desired speed term, define its reward function:

[0110] ;

[0111] In the formula, , , are the relative longitudinal speed, relative lateral speed, and relative heading angular speed between the hovercraft and the desired speed at the current moment, , , are reward coefficients greater than 0, r u is the longitudinal speed reward function, r v is the lateral speed reward function, r r is the heading angular speed reward function.

[0112] During navigation, the hovercraft should also be kept stable. Therefore, apply a positive incentive for stable motion, and the higher the stability, the higher the reward value; the variables related to stability are the roll angle and roll angular speed of the hovercraft, and define the corresponding reward function as:

[0113] ;

[0114] ;

[0115] Wherein, , are respectively the roll angle and the roll angular velocity of the hovercraft at the current moment, , are respectively reward coefficients greater than 0, is the roll angle reward function, is the roll angular velocity reward function.

[0116] In summary, the total reward function r is:

[0117] ;

[0118] Step 3-3: Set the termination condition: Through experience and testing, set the termination step number to be 2000, and end this training when this step number is exceeded.

[0119] Step 4: According to the desired speed in Step 3 and the speed control model in Step 2, design a hovercraft speed controller based on intelligent adaptive linear active disturbance rejection control. The specific structure of the controller is as Figure 3 shown, and rely on the thrust generated by the variable pitch air propeller to control the longitudinal speed to ensure the stable speed of the hovercraft;

[0120] Linear active disturbance rejection control simplifies the nonlinear part in traditional active disturbance rejection control through linear transformation, omits the tracking differentiator, and introduces the concept of bandwidth, greatly reducing the number of parameters to be tuned, and focusing the configuration on the Linear Extended State Observer (ESO) and Linear States Error Feedback (LSEF). Rewrite the speed controller into the form of an extended system as follows:

[0121] ;

[0122] Wherein, x 1 , x 2 are extended states; u is the longitudinal speed of the hovercraft in the moving coordinate system, is the total system disturbance, is the control gain, is the control quantity;

[0123] Then, a linear extended observer is designed in the form of:

[0124] ;

[0125] wherein, , is the velocity error feedback gain vector, , is the observed output of the velocity observer, is the control gain, is the control quantity;

[0126] The corresponding hovercraft speed controller is:

[0127] ;

[0128] wherein, is the controller bandwidth, is the system input, i.e., the desired speed, , is the observed output of the velocity observer, is the control gain, is the control quantity.

[0129] Step 5: According to the desired heading and desired lateral position in Step 3 and the heading-lateral displacement decoupling control model in Step 2, and based on the intelligent adaptive linear active disturbance rejection decoupling control algorithm, design a hovercraft heading-lateral decoupling controller. The specific structure of the controller is as Figure 4 shown, to ensure the stability of the hovercraft heading and lateral positions and improve the navigation safety of the hovercraft.

[0130] The specific content of the said Step 5 includes:

[0131] Step 5-1: Simplify the heading decoupling control model in Step 2:

[0132] ;

[0133] wherein, is the virtual control quantity of the heading control channel, , where , , are time-varying coefficients;

[0134] Step 5-2: Simplify the lateral displacement decoupling control model in Step 2:

[0135] ;

[0136] wherein, is the virtual control quantity of the lateral control channel, , where , , are time-varying coefficients;

[0137] Step 5-3: Combine the control models in Steps 5-1 and 5-2 into a matrix form, and represent them using the bow acceleration and the lateral displacement acceleration as follows:

[0138] ;

[0139] In the formula, represents system disturbances and system state uncertainties, is the virtual control quantity matrix;

[0140] The specific expression form is:

[0141] , ;

[0142] Combine the bow control input and the lateral displacement control input in Step 2 to obtain the coupling system gain matrix as:

[0143] ;

[0144] Its inverse is the conversion matrix between the virtual control quantity and the actual control quantity. In actual situations, it may encounter the irreversible situation. At this time, an approximate invertible matrix can be found to replace it. is the longitudinal installation position, is the distance from the pressure center of the air rudder to the rudder shaft, is the installation position of the bow nozzle.

[0145] According to the virtual control quantity of the bow control channel, the bow-lateral decoupling controller of the hovercraft in the bow control channel is:

[0146] ;

[0147] In the formula, is the desired bow error, is the desired bow direction, , , are the observation outputs of the bow observer, , , , , is the bow error feedback gain vector, is the virtual control quantity of the bow control channel, is the initial value of the virtual control quantity of the heading controller, is the heading control quantity;

[0148] According to the virtual control quantity The decoupling controller for the heading-lateral direction of the hovercraft in the lateral control channel is obtained as:

[0149] ;

[0150] In the formula, is the expected lateral position error, , , are the observation outputs of the lateral position observer, , , , , are the feedback gain vectors of the lateral position error, is the virtual control quantity of the lateral displacement control channel, is the initial value of the virtual control quantity of the lateral position controller, is the lateral position control quantity, is the expected lateral displacement.

[0151] The intelligent motion planning method for hovercraft based on the DDPG algorithm proposed by the present invention can quickly provide accurate expected relative speed, expected relative heading, and expected relative lateral position for the hovercraft, providing a basis for the design of the lower-level controller; the intelligent adaptive linear active disturbance rejection control method proposed by the present invention designs the hovercraft speed controller and heading-lateral decoupling controller by combining the BP neural network and active disturbance rejection control. The designed controller can realize online adjustment of the controller parameters, greatly reducing the parameter adjustment time. At the same time, it can more accurately observe and compensate for disturbances, meeting the control requirements under special operation navigation; the intelligent navigation method proposed by the present invention realizes the autonomous trajectory tracking of the hovercraft through the mutual cooperation between the motion planning layer and the control layer, increasing the navigation safety of the hovercraft.

[0152] In this specification, the same and similar parts between various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments described later, the description is relatively simple, and the relevant parts can refer to the partial description of the foregoing embodiments.

[0153] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An intelligent action planning and trajectory control method for hovercraft based on reinforcement learning, characterized in that, Including: Step 1: Determine the motion environment information of the fully air-cushioned hovercraft; Establish a model of the fully air-cushioned hovercraft to obtain the heading, position, speed, and roll of the hovercraft; Step 2: Utilize the heading, position, speed, and roll of the hovercraft obtained in Step 1 to simplify the strong coupling problem between the heading control channel and the lateral control channel, and obtain a speed control model and a heading-lateral displacement decoupling control model; Step 3: According to the speed control model and the heading-lateral displacement decoupling control model in Step 2 and the fully air-cushioned hovercraft model in Step 1, obtain the navigation information of the hovercraft. Through the deep deterministic policy gradient in the reinforcement learning algorithm, plan the intelligent actions during the navigation process of the hovercraft to obtain the desired speed, desired heading, and desired lateral position of the hovercraft; Step 4: According to the desired speed in Step 3 and the speed control model in Step 2, design a hovercraft speed controller based on intelligent adaptive linear active disturbance rejection control, and rely on the thrust generated by the variable pitch airscrew to control the longitudinal speed to ensure the stability of the hovercraft speed; Step 5: According to the desired heading and desired lateral position in Step 3 and the heading-lateral displacement decoupling control model in Step 2, and based on the intelligent adaptive linear active disturbance rejection decoupling control algorithm, design a hovercraft heading-lateral decoupling controller to ensure the stability of the heading and lateral position of the hovercraft; The specific content of Step 2 includes: Step 2-1: Since the adaptive linear active disturbance rejection control is a model-free control method, the standard form of the speed control model of the hovercraft can be directly obtained as: where \(u\) is the longitudinal velocity of the hovercraft in the moving coordinate system, \(f\) is the total system disturbance, \(b\) u is the control gain, and \(u\) u is the control variable; Step 2-2: From the hovercraft model established in Step 1, the heading control model and the lateral displacement control model are respectively: where, M za is the bow aerodynamic moment, M zm is the bow air momentum moment, M zh is the bow hydrodynamic moment, M zR is the bow moment of the air rudder, M zp is the bow moment of the air propeller, M zn is the bow moment of the bow nozzle, F ya is the lateral aerodynamic force, F ym is the lateral air momentum force, F yh is the lateral hydrodynamic force, F yR is the lateral force of the air rudder, F yp is the lateral thrust of the air propeller, F yn is the lateral thrust of the bow nozzle, I z is the moment of inertia of the hovercraft about the z-axis, ψ represents the roll angle and the heading angle of the hovercraft in the north-east coordinate system, m is the mass of the hovercraft, and u, v, r, y respectively represent the longitudinal velocity, lateral velocity, heading angular velocity, and lateral position of the hovercraft in the motion coordinate system; the bow control actuator and the lateral displacement control actuator in step 2 are the air rudder and the bow nozzle respectively, and the bow control input u ψ is the bow moment generated by the air rudder, and the lateral displacement control input u y is the lateral force generated by the bow nozzle, and the mathematical form is expressed as follows: where C Ry is the lateral rudder force coefficient, ρ a is the air density, u aa is the oncoming flow velocity received by the air rudder, S R is the air rudder area, x R is the longitudinal installation position, X p is the distance from the pressure center received by the air rudder to the rudder shaft, T n is the bow nozzle thrust, is the nozzle angle; Define the lateral force generated by the air rudder as the interference τ of the heading control channel on the lateral control channel uψ , and define the yaw moment generated by the bow nozzle as the interference τ of the lateral control channel on the heading control channel uy , which is expressed in mathematical form as follows: where x n is the installation position of the bow nozzle; The heading decoupling control model in Step 2 is: Where, N r , Y r , are the derived base values for calculating the non-dimensional hydrodynamic coefficients in the heading and lateral directions, x h is the longitudinal position where the hydrodynamic force or moment acts on the hull, ψ are the roll angle and heading angle of the hovercraft in the north-east coordinate system, I z is the moment of inertia of the hovercraft about the z-axis, v and r are the lateral velocity and heading angular velocity of the hovercraft in the moving coordinate system, u ψ is the yawing moment generated by the air rudder, τuy is the interference of the lateral control channel on the heading control channel, τ ψ is the disturbance of the heading control channel; The lateral displacement decoupling control model is: where τ y is the lateral control channel perturbation, uy is the lateral displacement control input, y, v, and r respectively represent the lateral position, lateral velocity, and yaw angular velocity of the hovercraft in the moving coordinate system, m is the mass of the hovercraft, Y r , is the derived base value for calculating the lateral dimensionless hydrodynamic coefficient, τ uψ is the interference of the yaw control channel on the lateral control channel; The specific design of the intelligent adaptive active disturbance rejection heading-lateral decoupling controller in Step 5 is as follows: Step 5-1: Simplify the heading decoupling control model in Step 2-2: where U ψ is the virtual control quantity of the heading control channel, and U ψ = a3u ψ + a3τ uy , where a1, a2, and a3 are time-varying coefficients; Step 5-2: Simplify the lateral displacement decoupling control model in Step 2-2: where U y is the virtual control quantity of the lateral control channel, and U y = a6u ψ + a6τ uy , where a4, a5, and a6 are time-varying coefficients; Step 5-3: Combine the control models in Step 5-1 and Step 5-2 into a matrix form to obtain the intelligent adaptive active disturbance rejection controllers for the heading control channel and the lateral displacement control channel.

2. The intelligent action planning and trajectory control method for an air-cushion vehicle based on reinforcement learning according to claim 1, characterized in that The hovercraft model in Step 1 includes a kinematic model and a dynamic model; The kinematic model is: where ξ, η, ψ represent the northward position, eastward position, roll angle and heading angle of the hovercraft in the northeast coordinate system, respectively, and u, v, p, r represent the longitudinal velocity, lateral velocity, roll angular velocity and heading angular velocity of the hovercraft in the moving coordinate system, respectively.

3. The intelligent action planning and trajectory control method for an air-cushion vehicle based on reinforcement learning according to claim 2, wherein, The dynamic model in Step 1 is: where m is the mass of the hovercraft, F x , F y , M x , M z are respectively the longitudinal resultant force, the lateral resultant force, the transverse tilting resultant moment, and the yawing resultant moment acting on the hovercraft, and I x , I z are respectively the moments of inertia of the hovercraft about the x and z axes; u, v, p, and r respectively represent the longitudinal speed, lateral speed, roll angular velocity, and heading angular velocity of the hovercraft in the moving coordinate system.

4. A hovercraft intelligent motion planning and trajectory control method based on reinforcement learning according to claim 1, characterized in that, The specific content of Step 3 includes: Step 3-1: Design the state space of all decision-making information affecting the hovercraft and the action space of all action sets of the hovercraft; Step 3-2: Design the reward function: During the intelligent navigation process of the hovercraft, the goal of learning is that the hovercraft can approach the desired trajectory; Step 3-3: Design the termination condition: Through experience and testing, design the termination step T as 2000, and end this training when this number of steps is exceeded.

5. The intelligent action planning and trajectory control method for an air-cushion vehicle based on reinforcement learning according to claim 1, wherein The specific design of the intelligent adaptive active disturbance rejection speed controller in Step 4 is: where \(u\) u is the control quantity, \(b\) u is the control gain, \(\omega\) uc is the controller bandwidth, \(v\) u is the desired speed, \(z\) u1 and \(z\) u2 are the observed outputs of the speed observer.

6. The intelligent action planning and path control method for an air-cushion vehicle based on reinforcement learning according to claim 5, characterized in that, The intelligent adaptive active disturbance rejection controllers for the heading control channel and the lateral displacement control channel are: where, e ψ is the desired heading error, v ψ is the desired heading, z ψ1 , z ψ2 , z ψ3 are the observation outputs of the heading observer, β ψo1 , β ψo2 , β ψo3 , β ψc1 , β ψc2 is the heading error feedback gain vector, U ψ is the virtual control quantity of the heading control channel, U ψ0 is the initial value of the virtual control quantity of the heading controller, b ψ is the heading control quantity; where, e y is the expected lateral position error, z y1 , z y2 , z y3 are the observation outputs of the lateral position observer, β yo1 , β yo2 , β yo3 , β yc1 , β yc2 are the lateral position error feedback gain vectors, U y is the virtual control quantity of the lateral displacement control channel, U y0 is the initial value of the virtual control quantity of the lateral position controller, b y is the lateral position control quantity, v y is the expected lateral displacement.

Citation Information

Patent Citations

  • Full-hovering hovercraft path tracking method

    CN113867352A

  • Water surface target collaborative hunting method based on multi-agent reinforcement learning

    CN117806318A