An optimal heading control method for AUV with digital-analog collaborative driving in a three-dimensional unknown environment

By combining sliding mode control and reinforcement learning, the problem of area maintenance control for underwater autonomous vehicles in three-dimensional unknown environments was solved, achieving high-precision and robust optimal heading positioning in unknown ocean currents and improving the AUV's adaptive capabilities.

CN119536324BActive Publication Date: 2025-10-28HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411713508.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-10-28
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively maintain area control for autonomous underwater vehicles in unknown three-dimensional environments. Traditional control methods rely on prior knowledge and lack robustness and adaptability.

Method used

An optimal heading control method for AUVs driven by a combination of digital and analog models in a three-dimensional unknown environment is adopted. By combining sliding mode control and reinforcement learning, a sliding mode controller is designed by establishing kinematic and dynamic models, and the PPO reinforcement learning algorithm is used to compensate for unknown disturbance forces. An actor and evaluator network is constructed and trained to achieve optimal heading control.

Benefits of technology

It significantly improves the control accuracy and robustness of underwater autonomous vehicles in complex ocean current environments, and achieves optimal heading positioning control in three-dimensional space without relying on prior knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536324B_ABST
    Figure CN119536324B_ABST
Patent Text Reader

Abstract

This invention discloses an optimal heading control method for an AUV (Autonomous Underwater Vehicle) under a three-dimensional unknown environment, driven by a combined numerical and analog model. The method first establishes a three-dimensional six-degree-of-freedom kinematic and dynamic model of the underactuated AUV in an unknown time-varying ocean current environment, and designs a sliding mode controller. Secondly, to address the disturbance forces experienced by the sliding mode controller in the three-dimensional unknown ocean current environment, a compensation scheme using a PPO (Progressive Point of View) reinforcement learning algorithm is designed, and a state space, action space, reward function, actor network, and evaluator network are constructed to train the AUV. After training, the reinforcement learning algorithm outputs the disturbance force compensation in real time based on the position and velocity of the underactuated AUV under the influence of the current three-dimensional unknown time-varying ocean current, and the sliding mode controller outputs the final control quantity. This invention improves the accuracy and robustness of sliding mode control in complex ocean current environments, achieving optimal heading positioning control of an underwater autonomous vehicle in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of motion control technology for autonomous underwater vehicles, specifically referring to an optimal heading control method for AUVs driven by a combination of digital and analog models in a three-dimensional unknown environment. Background Art

[0002] The motion control of autonomous underwater vehicles (AUVs) can be used to perform various tasks and performs excellently with minimal human intervention, thus playing a vital role in ocean exploration and scientific missions. Some of these tasks, such as marine environmental monitoring and weather forecasting, require AUVs to remain within a certain area for extended periods, i.e., area-keeping capability. Currently, vessels with fixed-point positioning or area-keeping capabilities are primarily dynamic positioning (DP) vessels or large marine operational platforms. Fixed-point positioning or area-keeping control can be achieved through anchorage positioning or DP systems. The hydrodynamic coefficients of an AUV vary according to time-varying environmental disturbances and the AUV's motion information. These hydrodynamic coefficients are also affected by the motion of other degrees of freedom (DOFs). Furthermore, most AUVs are designed with an underactuated configuration, requiring the control of more DOFs than the number of independent control inputs to maintain actuator efficiency at high speeds. Therefore, AUVs are typically underactuated, highly nonlinear, and highly coupled. Simultaneously, AUVs are inevitably affected by unknown time-varying disturbances, complex hydrodynamics, and uncertainties. Due to the characteristics of AUVs, area-keeping control methods used for fully maneuverable powered vessels, such as anchoring and dynamic positioning (DP), cannot be directly applied to achieve area-keeping control of AUVs. Currently, research on area-keeping control for AUVs is still limited. Therefore, effective and reliable area-keeping control remains a very challenging task for AUVs.

[0003] The structure of autonomous underwater vehicles (AUVs) is highly complex, exhibiting highly coupled nonlinear characteristics and both structured and unstructured uncertainties. Changes in the hydrodynamic environment alter their motion characteristics, making accurate mathematical models difficult to obtain and posing significant challenges to control system design. Furthermore, in actual operation, AUVs are susceptible to disturbances from uncertainties in the marine environment, such as wind, waves, and currents, which significantly impact their control performance. Traditional control methods, such as PID control, sliding mode control, and fuzzy control, heavily rely on prior knowledge and struggle to cope with changes in complex environments. In contrast, reinforcement learning can comprehensively consider multiple objectives and constraints, learning nonlinear control strategies that adapt to the dynamic and disturbance characteristics of complex systems. Through continuous interaction with the environment, reinforcement learning can continuously update its strategies during operation, enabling the controller to respond promptly to system changes and disturbances, achieving superior control performance and demonstrating strong real-time learning and adaptive capabilities. However, current research combining reinforcement learning with sliding mode control for hovering and positioning of AUVs primarily focuses on tuning fixed parameters that require manual adjustment in sliding mode control; the system's robustness and adaptability still need further improvement. Summary of the Invention

[0004] To address the shortcomings of traditional model-based control methods in the three-dimensional spatial positioning of autonomous underwater vehicles (AUVs), this invention proposes an optimal heading position control method for AUVs driven by a combination of numerical and model data in unknown three-dimensional environments.

[0005] To address the aforementioned technical problems, this invention proposes an optimal bow control method for AUVs driven by a combination of digital and analog models in a three-dimensional unknown environment, comprising:

[0006] Step 1: Establish a three-dimensional six-degree-of-freedom kinematic and dynamic model of an underactuated AUV under unknown time-varying ocean current conditions;

[0007] Step 2: Based on the three-dimensional kinematic and dynamic models of the underactuated AUV, design a sliding mode controller based on position error observation information that can achieve three-dimensional optimal heading control;

[0008] Step 3: To address the difficult-to-observe disturbance forces experienced by the sliding mode controller in a three-dimensional unknown ocean current environment, a compensation scheme based on the data-driven proximal policy optimization (PPO) reinforcement learning algorithm is designed, and the state space, action space, and reward function of PPO are constructed.

[0009] Step 4: Based on the established state space, action space and reward function, design the actor network and evaluator network in the PPO algorithm, build a simulation scenario considering three-dimensional unknown time-varying ocean currents to train the AUV, and form a complete numerical-analog collaborative drive control method.

[0010] Step 5: After training is completed, reinforcement learning outputs the corresponding disturbance force compensation in sliding mode control in real time based on the position and velocity of the underactuated AUV under the influence of the current three-dimensional unknown time-varying ocean current. The sliding mode controller outputs the final control quantity to achieve the desired control effect.

[0011] Step one is as follows:

[0012] The kinematic model of a three-dimensional six-DOF AUV in an unknown time-varying ocean current environment is as follows:

[0013]

[0014] The dynamic model is as follows:

[0015]

[0016] in It is the transformation matrix from its own moving coordinate system to the fixed coordinate system, η=η=[ξ,η,ζ,φ,θ,ψ] T It is a vector in a fixed coordinate system composed of the AUV's current position (ξ,η,ζ) and the roll angle φ, pitch angle θ, and heading angle ψ. It is the velocity of the AUV in its own coordinate system. is the velocity of the ocean current in a fixed coordinate system, and M is the inertia matrix of the rigid body mass and the added mass; The Coriolis centripetal force is calculated from the rigid body weight and added mass matrix. This is the damping force matrix; τ is the thrust generated by the AUV propeller; τ c It refers to the interference force exerted by the AUV's environment on the AUV.

[0017] Step two is as follows:

[0018] Based on the three-dimensional six-degree-of-freedom kinematic and dynamic models, the sliding mode control of a three-dimensional underactuated AUV is divided into a forward controller and an attitude controller. The forward controller is defined according to the optimal heading principle. e and p e Let u be the error of the sliding surface of the forward controller. e It is the error between the current speed and the desired speed of the AUV, p e It is the error between the current position and the desired position of the AUV, which is controlled by adjusting the longitudinal thrust along the sliding surface s1=u e +λ1p e By controlling the AUV's position and velocity errors to approach zero, the AUV can remain on the virtual sphere for an extended period despite unknown ocean current interference, and select... The exponential approach law, where ε1, λ1, and k1 are design parameters greater than 0, can be used to obtain the forward control law by combining the sliding surface function and the dynamic model.

[0019] The attitude controller is divided into a pitch controller and a yaw controller. To ensure optimal yaw control, it is necessary to ensure that the underactuated AUV always points towards the virtual sphere center. Based on the optimal yaw principle, q is defined... e and θ e For the error of the sliding surface of the pitch controller, q e and θ e These are the errors between the current pitch rate and the desired pitch rate of the AUV, and the errors between the current pitch angle and the desired pitch angle of the AUV, respectively. The pitch moment is controlled along the sliding surface s2 = q. e +λ2θ e Control the AUV's pitch rate and pitch angle error to approach 0, and select... The exponential approach law, where ε2, λ2, and k2 are design parameters greater than 0, can be used to obtain the pitch control law by combining the sliding surface function and the dynamic model.

[0020] According to the optimal heading control principle, to ensure that the underactuated AUV always points towards the virtual sphere center, it is necessary to ensure that the current pitch angle reaches the desired pitch angle while the current heading angle also reaches the desired heading angle. Let r be defined as... e and ψ e For the error definition of the sliding surface of the heading controller, r e and ψ e These are the errors between the current and desired heading angular velocities of the AUV, and the errors between the current and desired heading angles of the AUV, respectively. The heading controller moves along the sliding surface s3 = r. e +λ3ψ e Control the AUV's heading angular velocity and heading angle error to approach 0, and select... The exponential approach law, where ε3, λ3, and k3 are design parameters greater than 0, can be used to obtain the heading control law by combining the sliding surface function and the dynamic model.

[0021] Step three specifically involves:

[0022] Since the control law contains unobservable disturbances, this invention designs a reinforcement learning network to compensate for these unknown disturbances. The designed reinforcement learning state space includes: the deviation e1 between the current AUV pose and the current desired pose, the deviation e2 between the current AUV position and the current desired position, and the current AUV's velocity and pose; the action space includes: three unknown disturbances f. u f q and f r The reward function includes: a reward / penalty item related to the deviation e1 between the current AUV position and the desired position, and a reward / penalty item related to the deviation e2 between the current AUV attitude and the desired attitude.

[0023] Step four is as follows:

[0024] Based on the state space, action space, and reward function designed in step three, a suitable actor network and evaluator network are constructed. The model required by the PPO algorithm includes four networks: an actor network, a critic network, a target actor network, and a target critic network. The input layer of the actor network consists of variables in the state space; the output layer outputs the disturbance forces at the positions of each sliding mode controller to compensate for the disturbances caused by the unknown environment to the AUV, ensuring that the AUV can work in the desired position and attitude. The input layer of the critic network consists of the actions and states of the reinforcement learning, and the output is the Q-value of the state-action evaluation. The structure of the actor network and the critic network is consistent with the corresponding target network. The AUV is trained using ocean current data under unknown ocean current environments in the constructed reinforcement learning network.

[0025] Step five is as follows:

[0026] The disturbance force of the output obtained from training is provided as compensation to the sliding mode controller to achieve optimal heading control of the underactuated AUV in a three-dimensional unknown ocean current environment.

[0027] Beneficial effects of this invention:

[0028] This invention utilizes reinforcement learning to design neural networks, enabling autonomous underwater vehicles (AUVs) to continuously learn from their environment in a three-dimensional space with unknown ocean currents, without relying on prior knowledge or precise mathematical models. Reinforcement learning is used to compensate for unknown components in sliding mode control, significantly improving the accuracy and robustness of traditional sliding mode control in complex ocean current environments, and achieving optimal heading positioning control for AUVs in three-dimensional space. Attached Figure Description

[0029] Figure 1 This is a control system framework diagram of the underwater autonomous vehicle according to the method of the present invention;

[0030] Figure 2 This is a schematic diagram of ocean current velocity in the simulated environment in the method of the present invention;

[0031] Figure 3 This is a schematic diagram of the three-dimensional coordinate system of the underwater autonomous vehicle in the method of the present invention;

[0032] Figure 4 This is a schematic diagram of the network structure of the reinforcement learning part in the method of the present invention;

[0033] Figure 5 This is a diagram showing the operational trajectory of the underwater autonomous vehicle according to the method of the present invention.

[0034] Figure 6 This is a schematic diagram illustrating the position error during the operation of the underwater autonomous vehicle according to the method of the present invention;

[0035] Figure 7 This is a schematic diagram illustrating the pitch angle error during the operation of the underwater autonomous vehicle according to the method of the present invention.

[0036] Figure 8 This is a schematic diagram of the heading angle error during the operation of the underwater autonomous vehicle according to the method of the present invention. Detailed Implementation

[0037] The implementation steps of the present invention will be further described below with reference to the accompanying drawings and specific examples:

[0038] This invention proposes an optimal heading control method for AUVs driven by a combination of digital and analog models in a three-dimensional unknown environment. The flowchart is as follows: Figure 1 As shown, the specific implementation steps are as follows:

[0039] Step 1: Based on the movement of the autonomous underwater vehicle in an unknown time-varying ocean current environment, such as... Figure 3 A three-dimensional, six-DOF underwater autonomous vehicle (AUV) kinematic and dynamic model is established as shown. Since roll is balanced by the AUV's own restoring torque, the simplified model is as follows:

[0040]

[0041] in It is the transformation matrix from its own moving coordinate system to the fixed coordinate system, η=[ξ,η,ζ,θ,ψ] T It is a vector in a fixed coordinate system consisting of the AUV's current position (ξ,η,ζ) and the pitch angle θ and heading angle ψ. It is the velocity of the AUV in its own coordinate system, such as Figure 2 As shown is the velocity of the ocean current in a fixed coordinate system, and M is the inertia matrix of the rigid body mass and the added mass; The Coriolis centripetal force is calculated from the rigid body weight and added mass matrix. This is the damping force matrix; τ is the thrust generated by the AUV propeller; τ c This represents the interference force exerted by the AUV's environment on the AUV; the specific details of each matrix are as follows:

[0042]

[0043] In the formula m x c xx d x All are scalars.

[0044] Step 2: Based on the established model, design a 3D underactuated AUV sliding mode controller, which consists of a forward controller and an attitude controller. The forward controller is defined according to the optimal heading principle. e and p eThe error definition for the sliding surface of the forward controller is given, where u e It is the error between the current speed and the desired speed of the AUV, p e It is the error between the current position and the desired position of the AUV, which is controlled by adjusting the longitudinal thrust along the sliding surface s1=u e +λ1p e By controlling the AUV's position and velocity errors to approach zero, the AUV can remain on the virtual sphere for an extended period despite unknown ocean current interference, and select... The exponential reaching law, combined with the sliding surface function and the dynamic model, yields the forward control law:

[0045]

[0046] Where f u It is the unknown disturbance force moving forward, and λ1, k1, and ε1 are the parameters to be designed.

[0047] The attitude controller is divided into a pitch controller and a yaw controller. To ensure optimal yaw control, it is necessary to ensure that the underactuated AUV always points towards the virtual sphere center. Based on the optimal yaw principle, q is defined... e and θ e For the error definition of the sliding surface of the pitch controller, q e and θ e These are the errors between the current pitch rate and the desired pitch rate of the AUV, and the errors between the current pitch angle and the desired pitch angle of the AUV, respectively. The pitch moment is controlled along the sliding surface s2 = q. e +λ2θ e Control the AUV's pitch rate and pitch angle error to approach 0, and select... The exponential reaching law, combined with the sliding surface function and the dynamic model, yields the pitch control law:

[0048]

[0049] Where f q The pitch is an unknown disturbance force, and λ2, k2, and ε2 are the parameters to be designed.

[0050] According to the optimal heading control principle, to ensure that the underactuated AUV always points towards the virtual sphere center, it is necessary to ensure that the current pitch angle reaches the desired pitch angle while the current heading angle also reaches the desired heading angle. Let r be defined as... e and ψ e For the error definition of the sliding surface of the heading controller, r e and ψ e These are the errors between the current and desired heading angular velocities of the AUV, and the errors between the current and desired heading angles of the AUV, respectively. The heading controller moves along the sliding surface s3 = r. e +λ3ψe Control the AUV's heading angular velocity and heading angle error to approach 0, and select... The exponential reaching law, combined with the sliding surface function and the dynamic model, yields the heading control law:

[0051]

[0052] Where f r It is the unknown disturbance force in the bow direction, and λ3, k3, and ε3 are the parameters to be designed.

[0053] Step 3: Since the control law contains unobservable disturbances, this invention designs a reinforcement learning compensation method based on the data-driven Proximity Policy Priority (PPO) algorithm to compensate for unknown disturbances. The constructed reinforcement learning state space is defined as S: [p, u t ,v t ,w t ,q t ,r t ,θ e ,ψ e Where p is the current distance between the AUV and the desired virtual sphere center, v t =[u t ,v t ,w t ,q t ,r t ] is the measured lag-average velocity, θ e ,ψ e These are the errors between the current pitch angle and the current desired pitch angle of the AUV, respectively; the motion space is: a:[f u ,f q ,f r There are three unknown disturbances; the reward function consists of three parts:

[0054]

[0055] r = r1 + r2 + r3;

[0056] Where p e It is the difference between the current distance between the AUV and the virtual sphere center and the expected working distance.

[0057] Step 4: Based on the state space, action space, and reward function established in Step 3, design the actor network and evaluator network in the PPO algorithm. The model required for the PPO algorithm includes four networks, such as... Figure 4As shown, the Actor and Critic networks, as well as the target Actor and Critic networks, are respectively. The Actor network has an input layer, two hidden layers, and an output layer. The input layer is the state space; the output layer outputs three unknown disturbance forces in the sliding mode controller to compensate for the disturbances caused by the unknown environment to the AUV, ensuring that the AUV can work in the desired position and attitude. The Critic network has an input layer, three hidden layers, and an output layer. The input layer is the action space and state space of reinforcement learning, and the output is the evaluation Q value of the pair. The structure of the Actor and Critic networks is consistent with the corresponding target networks. A simulation scenario considering three-dimensional unknown time-varying ocean currents is built, and the AUV is trained in the designed PPO network using time-varying ocean current data under unknown environment.

[0058] The trained neural network model is used as the policy in reinforcement learning, outputting estimates of action values ​​(i.e., disturbance forces), and these estimates are provided to the sliding mode controller. The sliding mode controller then generates a control law based on this information, which is ultimately applied to the underwater autonomous vehicle, thereby achieving three-dimensional optimal heading and positioning control of the underwater autonomous vehicle in an unknown time-varying ocean current environment.

[0059] Step 5: The disturbance force obtained from the training is used as compensation and provided to the sliding mode controller to achieve optimal heading control of the underactuated AUV in the three-dimensional environment. The flight trajectory is as follows: Figure 5 As shown. The control effect is as follows. Figure 6 , Figure 7 and Figure 8 As shown in the figure, the solid line represents the error of the autonomous underwater vehicle of the present invention, and the dotted line represents the error of the traditional sliding mode control. It can be seen from the figure that the traditional sliding mode does not process the disturbance force, and the error changes significantly. The method of combining reinforcement learning and sliding mode control can enable the autonomous underwater vehicle to basically maintain on the predetermined virtual sphere, achieving the expected control effect and accuracy.

[0060] The preferred embodiments of the present invention disclosed above are only used to illustrate the basic concepts of the present invention. The preferred embodiments do not describe all details exhaustively, nor do they limit the implementation of the present invention. Obviously, many modifications and changes can be made based on the content of this specification. The embodiments selected and specifically described herein are intended to better explain the principles of the present invention and its practical application, so that those skilled in the art can better understand and apply the present invention. The scope of protection of the present invention is limited to the claims and their full scope and equivalents.

Claims

1. A method for optimal heading control of an AUV driven by a combination of digital and analog models in a three-dimensional unknown environment, characterized in that, Includes the following steps: Step 1: Establish a three-dimensional six-degree-of-freedom kinematic and dynamic model of an underactuated AUV under unknown time-varying ocean current conditions; Step 2: Based on the three-dimensional kinematic and dynamic models of the underactuated AUV, design a sliding mode controller based on position error observation information that can achieve three-dimensional optimal heading control; Step 3: To address the disturbance forces experienced by the sliding mode controller in a three-dimensional unknown ocean current environment, a compensation scheme based on data-driven proximal policy optimization (PPO) reinforcement learning algorithm is designed, and the state space, action space, and reward function of PPO are constructed. Step 4: Based on the state space, action space, and reward function, design the actor network and evaluator network in the PPO algorithm, and build a simulation scenario that considers three-dimensional unknown time-varying ocean currents to train the AUV. The specific implementation process of step three is as follows: The design of the reinforcement learning network compensates for unknown disturbances in the control law. The state space of the reinforcement learning includes: the deviation e1 between the current AUV pose and the current desired pose, the deviation e2 between the current AUV position and the current desired position, and the current AUV velocity and pose; the action space includes: three unknown disturbance forces f. u f q and f r ; The reward function includes: a reward / penalty term related to the deviation e1 between the current AUV position and the desired position, and a reward / penalty term related to the deviation e2 between the current AUV pose and the desired pose. The reward function consists of three parts: r=r1+r2+r3; p e θ is the error between the current position and the desired position of the AUV. e ψ is the error between the current pitch angle and the desired pitch angle of the AUV. e It is the error between the current heading angle and the desired heading angle of the AUV; Step 5: After training is completed, reinforcement learning outputs the corresponding disturbance force compensation in sliding mode control in real time based on the position and velocity of the underactuated AUV under the influence of the current unknown time-varying ocean current in three dimensions. The sliding mode controller outputs the final control quantity to achieve optimal heading control of the underactuated AUV in the three-dimensional environment.

2. The optimal heading control method for AUVs driven by digital-analog co-drive in a three-dimensional unknown environment according to claim 1, characterized in that, The three-dimensional six-degree-of-freedom kinematic model is as follows: The dynamic model is as follows: in It is the transformation matrix from its own moving coordinate system to the fixed coordinate system, η=[ξ,η,ζ,φ,θ,ψ] T It is a vector in a fixed coordinate system composed of the AUV's current position (ξ,η,ζ) and the roll angle φ, pitch angle θ, and heading angle ψ. It is the velocity of the AUV in its own coordinate system. is the velocity of the ocean current in a fixed coordinate system, and M is the inertia matrix of the rigid body mass and the added mass; The Coriolis centripetal force is calculated from the rigid body weight and added mass matrix. This is the damping force matrix; τ is the thrust generated by the AUV propeller; τ c It refers to the interference force exerted by the AUV's environment on the AUV.

3. The optimal heading control method for AUVs driven by digital-analog co-drive in a three-dimensional unknown environment according to claim 2, characterized in that, The specific implementation process of step two is as follows: Based on the three-dimensional six-degree-of-freedom kinematic and dynamic models, the sliding mode control of a three-dimensional underactuated AUV is divided into a forward controller and an attitude controller. The forward controller is defined according to the optimal heading principle. e and p e Let u be the error of the sliding surface of the forward controller. e It is the error between the current speed and the desired speed of the AUV, p e It is the error between the current position and the desired position of the AUV, which is controlled by adjusting the longitudinal thrust along the sliding surface s1=u e +λ1p e By controlling the AUV's position and velocity errors to approach zero, the AUV can remain on the virtual sphere for an extended period despite unknown ocean current interference, and select... The exponential approach law, where ε1, λ1, and k1 are parameters greater than 0, is used to derive the forward control law by combining the sliding surface function and the dynamic model. The attitude controller is divided into a pitch controller and a yaw controller. q is defined based on the optimal yaw principle. e and θ e For the error of the sliding surface of the pitch controller, q e and θ e These are the errors between the current pitch rate and the desired pitch rate of the AUV, and the errors between the current pitch angle and the desired pitch angle of the AUV, respectively. The pitch moment is controlled along the sliding surface s2 = q. e +λ2θ e Control the AUV's pitch rate and pitch angle error to approach 0, and select... The exponential approach law, where ε2, λ2, and k2 are parameters greater than 0, is used to derive the pitch control law by combining the sliding surface function and the dynamic model. Based on the optimal heading control principle, r is defined as follows: e and ψ e For the error definition of the sliding surface of the heading controller, r e and ψ e These are the errors between the current and desired heading angular velocities of the AUV, and the errors between the current and desired heading angles of the AUV, respectively. The heading controller moves along the sliding surface s3 = r. e +λ3ψ e Control the AUV's heading angular velocity and heading angle error to approach 0, and select... The exponential approach law, where ε3, λ3, and k3 are design parameters greater than 0, is used to derive the heading control law by combining the sliding surface function and the dynamic model.

4. The optimal heading control method for AUVs driven by digital-analog co-drive in a three-dimensional unknown environment according to claim 3, characterized in that, The specific implementation process of step four is as follows: Based on the state space, action space, and reward function established in step three, an actor network and an evaluator network are constructed. The PPO algorithm model includes four networks: the actor network, the critic network, the target actor network, and the target critic network. The input layer of the actor network consists of variables in the state space; the output layer outputs the disturbance forces at the positions of each sliding mode controller to compensate for the disturbances caused by the unknown environment to the AUV, ensuring that the AUV can work in the desired position and pose. The input layer of the critic network consists of the actions and states learned in reinforcement learning, and the output is the Q-value of the state-action evaluation. The structure of the actor network and the critic network is consistent with the corresponding target network. The AUV was trained using ocean current data in an unknown ocean current environment within a reinforcement learning network.

Citation Information

Patent Citations

  • AUV environment optimal positioning control method based on sliding mode and reinforcement learning

    CN118011807A

  • AUV action plan and operation control method based on reinforcement learning

    JP2021034050A