Method for controlling attitude of foldable four-rotor cross-medium unmanned aerial vehicle based on dynamic programming
By establishing a dynamic model of a cross-medium UAV and introducing an adaptive dynamic programming controller, the problem of stable control of cross-medium aircraft in different media environments was solved, and high-precision and robust cross-medium flight control was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies lack a unified modeling method and a model-free controller with adaptive capabilities, making it difficult to achieve stable control of cross-medium aircraft in the air, underwater, and during medium transition phases, especially facing strong fluid disturbances and nonlinear interferences during medium transitions.
A dynamic model of a UAV, encompassing aerial, underwater, and medium-crossing phases, is established. A fixed-time convergent sliding mode observer and an adaptive dynamic programming controller are introduced. The control input is optimized through a Hamiltonian function, and online learning and adjustment are performed using a neural network to construct an attitude control method for a quadrotor UAV crossing medium.
It significantly improves the modeling accuracy and dynamic response prediction capability of aircraft cross-medium motion, overcomes strong disturbances and parameter mutations during medium conversion, achieves stable cross-domain operation and precise control, and enhances autonomous operation capability.
Smart Images

Figure CN121764153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) flight control technology, specifically to a dynamic programming-based attitude control method for a foldable quadrotor cross-medium UAV. Background Technology
[0002] In recent years, cross-medium aircraft, as a new type of intelligent equipment capable of both aerial and underwater operations, have attracted widespread attention in aerospace, underwater inspection, and emergency rescue fields. Compared with traditional single-medium aircraft, cross-medium aircraft face the challenges of dynamic modeling and control under two completely different flow field environments: air and water. Especially during the medium transition process, the aircraft needs to cross the water-air interface, accompanied by strong fluid abrupt changes and nonlinear disturbances, which places higher demands on control accuracy and system robustness.
[0003] Existing research mostly focuses on establishing separate motion models and control algorithms for aerial unmanned aerial vehicles (UAVs) and underwater vehicles, lacking a unified framework for aircraft dynamics modeling that integrates the entire process across air, underwater, and medium transitions. In the air phase, the aircraft is primarily influenced by gravity, lift, and aerodynamic forces, exhibiting typical six-degree-of-freedom rigid body flight characteristics. In the underwater phase, the aircraft faces complex nonlinear effects such as fluid drag, buoyancy, and viscous damping, resulting in sluggish dynamic response and significantly reduced propulsion efficiency. The medium transition phase is even more complex; surface disturbances, flow field disruptions, and surface tension combine to trigger instantaneous thrust shifts and attitude changes, making it difficult for dynamic models specific to a single medium to adapt. Therefore, constructing an integrated dynamics modeling method spanning all three phases is fundamental for achieving precise control and mission execution.
[0004] In controller design, existing control technologies are mainly divided into three categories: model-based controllers, nominal model-based controllers, and data-driven intelligent controllers. Traditional PID controllers, due to their simple structure and ease of implementation, are widely used in aerial drones. However, they inherently lack adaptive capabilities and struggle to handle parameter disturbances and sudden external interference in high-speed nonlinear systems. Model-based and nominal model-based control methods, such as linear quadratic regulators (LQR) or model predictive control (MPC), achieve optimal performance based on accurate system models. However, in cross-medium applications, due to drastic environmental changes and frequent abrupt changes in system model parameters, control accuracy and stability are significantly reduced. Especially when crossing water-air interfaces, the instantaneous switching between aerodynamics and fluid resistance leads to system dynamic mismatch, rendering model-based control methods ineffective.
[0005] With the development of intelligent algorithms, model-free control methods have gradually become an important research direction in cross-medium aircraft control. Model-free adaptive control adjusts the controller directly based on system input-output data, without requiring precise system modeling, making it suitable for complex system control in environments with high uncertainty and multiple disturbances. Currently, controllers based on reinforcement learning, neural networks, and adaptive dynamic programming (ADP) have shown initial success in traditional UAV control. However, these methods still face practical challenges when applied to cross-medium scenarios, including difficulties in underwater data acquisition and the non-stationary characteristics of the medium transition phase. Especially during mission execution, dynamic adjustment, performance constraints, and energy consumption balance of the aircraft at different stages are required, necessitating the exploration of a control strategy with online learning capabilities and full-condition adaptability.
[0006] In summary, for foldable quadcopter cross-medium aircraft, there is still a lack of a unified modeling method for the entire process from air to underwater to transition, as well as a model-free controller design framework that combines stability and adaptability. Therefore, there is an urgent need to conduct research on integrated dynamic modeling and robust intelligent control methods for cross-medium aircraft to achieve stable operation in highly maneuverable and multi-disturbance environments, supporting the technological development of future integrated air-sea combat, rescue, and inspection missions. Summary of the Invention
[0007] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a dynamic programming-based attitude control method for a foldable quadrotor cross-medium unmanned aerial vehicle to solve the aforementioned technical problems.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a dynamic programming-based attitude control method for a foldable quadrotor cross-medium UAV, comprising: S1: Establish a dynamic model of the UAV that includes the aerial, underwater and medium-crossing stages. The forces and moments caused by wave, wind and buoyancy effects are modeled as disturbance terms, and the control forces and control moments generated by the quadcopter are modeled as control terms. The relative immersion length is set to characterize the cross-medium state. S2: Define the tracking error vectors for altitude and attitude angle and construct the error dynamics with channel coupling perturbations. Design fixed-time convergent perturbation observers for the altitude channel and each attitude channel to obtain smooth error signals and perturbation estimates. S3: Based on the disturbance estimation, construct a performance index function, construct a Hamiltonian function and obtain an approximate optimal control input according to the HJB minimization condition, construct an evaluation neural network to approximate the cost function, construct an execution neural network to approximate the control strategy and update the network weights online, and output quadrotor control commands to achieve attitude stabilization tracking in the cross-medium process.
[0009] The present invention is further configured such that the multimodal dynamics model includes at least three sub-models: an air phase, an underwater phase, and a medium crossing phase; wherein the medium crossing phase sets the relative submersion length as the crossing state quantity, with a value range of 0 to 1, which is used to characterize the normalized submersion ratio from the intersection of the body and the water surface to the characteristic length of the body, and the gravity, buoyancy, fluid resistance and control torque terms are segmented or continuously interpolated and fused according to the relative submersion length.
[0010] The present invention is further configured to construct the altitude channel error and attitude channel error into a unified tracking error vector, wherein the attitude channel error includes at least roll angle error, pitch angle error and yaw angle error, and the tracking error vector serves as the common input of the fixed-time disturbance observer (FTDO) and the adaptive dynamic programming controller.
[0011] The present invention is further configured such that the FTDO sets up observer sub-channels according to the height channel and each attitude channel, and each sub-channel includes at least two levels of estimated state variables and corresponding observer gains. Fixed-time convergence is achieved through a composite sliding mode injection structure containing a sign function term and a power term. The power term adopts a first power exponent parameter and a second power exponent parameter. The value range of the first power exponent parameter is 0 to 1, and the value range of the second power exponent parameter is 1 to 2.
[0012] The present invention is further configured such that the disturbance estimate and smoothing error signal output by the FTDO are the input information of the adaptive dynamic programming controller, which are used to construct the performance index function and the Hamiltonian function to obtain the approximate optimal control input according to the HJB minimization condition.
[0013] The present invention is further configured such that the adaptive dynamic programming controller constructs a performance index function based on the observed tracking error vector, the performance index function including penalties for the error term and the control input term; wherein the error penalty matrix and the control penalty matrix are both positive definite symmetric matrices, and are used to generate the optimal control strategy related to the current state.
[0014] The present invention is further configured to construct a Critic network to approximate the cost function. The Critic network takes the tracking error vector as input and the estimated cost function as output. The cost function is represented by the inner product of the basis function vector and the weights of the Critic network. The error function constructed based on the Hamiltonian residual is used as the learning target. The weights of the Critic network are updated online through gradient descent to reduce the Hamiltonian residual and approximate the zero condition of the Hamiltonian equation.
[0015] The present invention is further configured to construct an Actor network to output a control strategy, wherein the Actor network generates control commands based on the tracking error vector, the system input channel mapping relationship, and the gradient information of the Critic network; the weights of the Actor network are set as adjustable parameters that have a mapping consistency constraint with the weights of the Critic network, so that the control strategy output by the Actor network and the cost function approximated by the Critic network converge collaboratively under the same performance index.
[0016] The present invention is further configured to apply stability constraints to the Actor network weight update through the Lyapunov function, wherein the Lyapunov function includes at least a tracking error related term and an Actor network weight related term.
[0017] The present invention is further configured such that, during the medium crossing phase, the cross-medium state is characterized by the relative submersion length, and the medium-related terms in the dynamic model are expressed based on the relative submersion length, so that the relative submersion length is 0 for the air state and 1 for the fully underwater state, and FTDO observation and adaptive dynamic programming control calculation are performed based on the model of the medium crossing phase.
[0018] This invention provides a dynamic programming-based attitude control method for a foldable quadrotor UAV that crosses media. It establishes a dynamic model of the UAV encompassing aerial, underwater, and media-crossing stages, modeling the forces and moments caused by wave, wind, and buoyancy effects as disturbance terms, and the control forces and moments generated by the quadrotor as control terms. A relative immersion length is set to characterize the cross-media state. Tracking error vectors for altitude and attitude angles are defined, and an error dynamic containing channel-coupled disturbances is constructed. Fixed-time convergent disturbance observers are designed for the altitude channel and each attitude channel to obtain smooth error signals and disturbance estimates. Based on the disturbance estimates, a performance index function is constructed, a Hamiltonian function is built, and an approximate optimal control input is obtained according to the HJB minimization condition. An evaluation neural network is constructed to approximate the cost function, and an execution neural network is constructed to approximate the control strategy and update the network weights online. The quadrotor control commands are output to achieve stable attitude tracking during the cross-media process. The beneficial effects include: 1. This invention establishes a complete mathematical model of a foldable quadcopter cross-medium aircraft, comprehensively characterizing the dynamic characteristics of the aircraft in the three stages of air, underwater and medium switching, significantly improving the modeling accuracy and dynamic response prediction capability of the aircraft's cross-medium motion.
[0019] 2. By introducing a fixed-time convergent sliding mode observer, the strong disturbances and parameter mutations in the medium conversion process can be effectively overcome, ensuring the real-time performance and robustness of the aircraft state estimation and improving the anti-interference performance of the control system.
[0020] 3. By combining an adaptive dynamic programming control strategy and utilizing neural networks to learn and optimize system performance indicators, the controller can adapt to multi-stage complex environmental changes and dynamically adjust control inputs, achieving stable cross-domain operation and precise control of the aircraft. The overall system possesses high-precision modeling, highly robust observation, and efficient intelligent control capabilities, significantly enhancing the autonomous operation capability of cross-medium aircraft in complex sea and air environments. This provides crucial support for its practical application in high-risk tasks such as submarine cable inspection, underwater search and rescue, and environmental monitoring, and has broad engineering application value.
[0021] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a structural framework diagram of adaptive dynamic programming in this invention; Figure 2 A schematic diagram of a foldable, cross-medium drone entering and exiting water; Figure 3 A schematic diagram of altitude command tracking for a foldable, cross-medium UAV; Figure 4 This is a diagram showing the pitch angle error of the cross-medium UAV in this invention; Figure 5 This is a diagram showing the roll angle error of the cross-medium UAV in this invention; Figure 6 This is a diagram showing the yaw angle error of the cross-medium UAV in this invention; Figure 7 This is a root mean square error diagram of the cross-medium UAV in this invention; Figure 8 This is a graph showing the variation of Hamiltonian values for the cross-medium UAV in this invention. Figure 9 A diagram showing the observation error of the position channel observer for a cross-medium UAV; Figure 10 A graph showing the weight changes of the evaluation neural network for the pitch channel of a cross-media UAV; Figure 11 This is a graph showing the weight changes of the execution neural network for the pitch channel of a cross-media UAV. Detailed Implementation
[0023] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.
[0024] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0025] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0026] A dynamic programming-based attitude control method for foldable quadrotor cross-medium UAVs, such as Figure 1 and Figure 2 As shown, it includes: S1: Establish a dynamic model of the UAV that includes the aerial, underwater and medium-crossing stages. The forces and moments caused by wave, wind and buoyancy effects are modeled as disturbance terms, and the control forces and control moments generated by the quadcopter are modeled as control terms. The relative immersion length is set to characterize the cross-medium state. S2: Define the tracking error vectors for altitude and attitude angle and construct the error dynamics with channel coupling perturbations. Design fixed-time convergent perturbation observers for the altitude channel and each attitude channel to obtain smooth error signals and perturbation estimates. S3: Based on the disturbance estimation, construct a performance index function, construct a Hamiltonian function and obtain an approximate optimal control input according to the HJB minimization condition, construct an evaluation neural network to approximate the cost function, construct an execution neural network to approximate the control strategy and update the network weights online, and output quadrotor control commands to achieve attitude stabilization tracking in the cross-medium process.
[0027] S1) Establish a dynamic model for a foldable quadrotor transmedium aircraft: The dynamic model of the UAV established is as follows:
[0028] The position vector and angular velocity vector of the UAV body coordinate system are expressed as follows: , Euler angles are represented as The velocity of the UAV relative to the inertial frame is expressed as... The transformation relationship between the inertial coordinate system and the volume coordinate system can be expressed as:
[0029]
[0030] The parameters are expressed as follows:
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038] in It is the stationary distance when the drone crosses the medium. and This indicates disturbances, primarily caused by forces and moments generated by waves, wind, and buoyancy. and It refers to the control force and torque of the drone. Defined as the relative submersion length, which is the distance from the point where the drone intersects the water surface to the point below the water surface on the drone. If the drone is in the air, The value is 0 if the drone is completely underwater. The value is 1.
[0039] The control forces and torques acting on the UAV during flight can be expressed as follows:
[0040] in, , , , The altitude of the aircraft rotor related: .
[0041] S2) Design based on a fixed-time convergent sliding mode observer: When designing the attitude controller, the stability of the height channel needs to be considered simultaneously. Therefore, the height and attitude angle tracking errors are defined as follows:
[0042] The target height is defined as The target attitude angle is expressed as , and Therefore, the attitude tracking control system can be written as: The coupled disturbances between the external environment and various altitudes and attitude channels are denoted as... .
[0043] To obtain smoother error signals and system disturbances, the fixed-time disturbance observer for each channel is designed as follows:
[0044] in , , … The parameters represent the estimated state of the observer. , and This is the observer gain of the FTDO, at which point the FTDO will remain stable for a fixed time.
[0045] Under the action of this observer It can be written in the following form: ,in, Here is the error matrix of FTDO. and It is a symmetric positive definite matrix. , and Defined as: , , Considering and ,definition , ; The state-space equation for the FTDO error can be expressed as: ; The candidate Lyapunov function is designed as follows: ; Differentiation The result obtained is: ; Symmetric positive definite matrix The following inequalities must be satisfied: ; in , We can obtain: ,in matrix The smallest eigenvalue. When When the size is sufficiently small, the observer converges within a fixed time: .
[0046] S3) Design a controller based on adaptive dynamic programming: In the design of the adaptive dynamic programming algorithm, two neural networks are designed to simulate the reinforcement learning structure, and an adaptive law is designed to adjust the weights of the neural networks online.
[0047] After FTDO observations, a cost function is defined based on the observed values to evaluate the current control strategy, which is related to the system state, error, and system input. ,in and It is a positive definite symmetric matrix.
[0048] Define the Hamiltonian function as: ,in The optimal cost function is obtained by solving the following Hamiltonian equation: ; This leads to the current optimal control input: ; The result of optimal control Substituting into the HJB equation, we get: ;in By solving the equation Substitute it into the equation To obtain the optimal control strategy. However, for a nonlinear system, solving the above nonlinear partial differential equations is extremely difficult.
[0049] Based on the properties of neural networks, the cost function and the optimal control strategy can each be represented by a single-layer neural network. Using a Critic neural network to approximate the cost function, it can be precisely expressed as: ,in Let the weights of the ideal critical neural network be represented as follows: The gradient of the cost function after approximation of the neural network is expressed as: ,in, and The optimal control strategy can be expressed as: ; The HJB equation can also be rewritten as: ,in It is the approximation error of the HJB equation, however, the ideal critic neural network weight vector Unknown. Therefore, the cost function for fitting the actual critic neural network can be expressed as: ,in These are the current estimated weights of the ideal critic neural network.
[0050] The optimal control strategy can be accurately represented by an actor neural network to ensure the smooth operation of the closed-loop system; ,in Ideal critic neural network weights The estimated value.
[0051] Furthermore, gradient descent can be used to define the weight vector of the critic neural network. and actor neural network weight vector Adaptive update rate: For the critic neural network, the adaptive update rate will be... Substituting into the HJB equation: Considering It should approach 0, therefore, by updating... To minimize the approximation error of the HJB equation : At the same time, the weight adjustment law of the critic neural network is obtained: ,in , Simultaneously, the adaptive update rate of the actor neural network is defined according to the Lyapunov function: Define the estimation error of the critic neural network and the weight estimation error of the actor neural network: Define Lyapunov functions: ; Its derivative is: Among them, Differentiation yields: ,in ,as well as .
[0052] So It can be represented as: ; right Differentiation yields:
[0053] right Differentiation yields:
[0054] Will , , Substitute In the middle, then It can be represented as:
[0055] Substituting the neural network estimation error into the above equation, we get:
[0056]
[0057] in , , .
[0058] therefore It can be transformed into:
[0059] in , , It is expressed as follows:
[0060]
[0061] During the controller design process, by selecting appropriate parameters, the following inequalities can be made to hold:
[0062] when When the derivative of the Lyapunov function is less than 0, the controller is stable.
[0063] In this embodiment, the control performance of the proposed control algorithm is verified through simulation. The initial state of the foldable quadrotor cross-medium UAV system is as follows: , The basis functions of a neural network are all defined as follows: The desired attitude angle is defined as: . , , , , , To demonstrate the superiority of the proposed method, a simulation comparison is conducted with the traditional ADRC algorithm.
[0064] Figures 3-11This is a comparison of the attitudes of a foldable, trans-medium quadrotor UAV under the control of FTDO-ADP and ADRC controllers during continuous underwater-to-aerial mode switching. The UAV follows... Figure 3 The way the altitude changes is shown, indicating that cross-media UAVs can track altitude commands well. Figures 4-6 It can be seen that in attitude control of continuous inlet and outlet water mode switching, FTDO-ADP has a shorter convergence time, smaller fluctuation amplitude, and significantly better dynamic response and stability than ADRC. Figure 7 The chart shows a comparison of the root mean square error (RMSE) of AAV. The RMSE of FTDO-ADP is consistently lower than that of ADRC after each switch and converges faster, demonstrating its superior overall control accuracy. Figure 8 The attitude channel Hamilton value shows that the Hamilton function value drops rapidly and tends to stabilize after switching, indicating that FTDO-ADP can continuously optimize control performance and balance input energy and system indicators in complex external disturbance environments. Figure 9 The attitude channel observation error shows that the observation error curve can still quickly approach zero after each switch, and the steady-state error is extremely small, which verifies the efficiency and accuracy of FTDO in the medium switching process. Figure 10 and Figure 11 The weights of the Critic and Actor networks in the pitch control ADP pitch channel demonstrate that the online learning mechanism of adaptive dynamic programming remains effective after mode switching. The Critic network accurately evaluates control actions, and the Actor network optimizes the strategy, ensuring that the attitude control quantity tends towards the optimum. In summary, in the joint simulation of continuous water entry and exit, the FTDO-ADP controller maintains the advantages of both airborne and underwater modes, significantly outperforming the ADRC controller in attitude response speed, control accuracy, and stability after each mode switch. It effectively copes with the complex operating conditions of continuous switching across media, further verifying the effectiveness and robustness of this strategy in the full-modal control of a dual-rotor cross-media aircraft.
[0065] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A dynamic programming-based attitude control method for a foldable quadrotor cross-medium UAV, characterized in that, include: S1: Establish a dynamic model of the UAV that includes the aerial, underwater and medium-crossing stages. The forces and moments caused by wave, wind and buoyancy effects are modeled as disturbance terms, and the control forces and control moments generated by the quadcopter are modeled as control terms. The relative immersion length is set to characterize the cross-medium state. S2: Define the tracking error vectors for altitude and attitude angle and construct the error dynamics with channel coupling perturbations. Design fixed-time convergent perturbation observers for the altitude channel and each attitude channel to obtain smooth error signals and perturbation estimates. S3: Based on the disturbance estimation, construct a performance index function, construct a Hamiltonian function and obtain an approximate optimal control input according to the HJB minimization condition, construct an evaluation neural network to approximate the cost function, construct an execution neural network to approximate the control strategy and update the network weights online, and output quadrotor control commands to achieve attitude stabilization tracking in the cross-medium process.
2. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 1, characterized in that, The multimodal dynamics model includes at least three sub-models: the air phase, the underwater phase, and the medium crossing phase. In the medium crossing stage, the relative immersion length is set as the crossing state quantity, with a value range of 0 to 1. It is used to characterize the normalized immersion ratio from the intersection of the body and the water surface to the characteristic length of the body. The gravity, buoyancy, fluid resistance and control torque terms are segmented or continuously interpolated and fused according to the relative immersion length.
3. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 2, characterized in that, The altitude channel error and attitude channel error are constructed into a unified tracking error vector. The attitude channel error includes at least roll angle error, pitch angle error and yaw angle error. The tracking error vector serves as the common input of the fixed-time disturbance observer (FTDO) and the adaptive dynamic programming controller.
4. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 3, characterized in that, FTDO sets up observer sub-channels for the height channel and each attitude channel. Each sub-channel includes at least two levels of estimated state variables and corresponding observer gains. Fixed-time convergence is achieved through a composite sliding mode injection structure containing a sign function term and a power term. The power term uses a first power exponent parameter and a second power exponent parameter. The value range of the first power exponent parameter is 0 to 1, and the value range of the second power exponent parameter is 1 to 2.
5. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 4, characterized in that, The disturbance estimate and smoothing error signal output by FTDO serve as input information for the adaptive dynamic programming controller, which is used to construct the performance index function and the Hamiltonian function to obtain the approximate optimal control input based on the HJB minimization condition.
6. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 1, characterized in that, The adaptive dynamic programming controller constructs a performance index function based on the observed tracking error vector, and the performance index function includes penalties for the error term and the control input term; The error penalty matrix and the control penalty matrix are both positive definite symmetric matrices, and are used to generate the optimal control policy related to the current state.
7. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 6, characterized in that, A Critic network is constructed to approximate the cost function. The Critic network takes the tracking error vector as input and the estimated cost function as output. The cost function is represented by the inner product of the basis function vector and the weights of the Critic network. The error function constructed based on the Hamiltonian residual is used as the learning objective. The weights of the Critic network are updated online through gradient descent to reduce the Hamiltonian residual and approximate the zero condition of the Hamiltonian equation.
8. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 7, characterized in that, An Actor network is constructed to output a control strategy. The Actor network generates control commands based on the tracking error vector, the system input channel mapping relationship, and the gradient information of the Critic network. The weights of the Actor network are set as adjustable parameters with a mapping consistency constraint with the weights of the Critic network, so that the control strategy output by the Actor network and the cost function approximated by the Critic network converge together under the same performance index.
9. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 8, characterized in that, The stability of the Actor network weight update is constrained by the Lyapunov function, which includes at least a tracking error term and an Actor network weight term.
10. The attitude control method for a foldable quadrotor cross-medium UAV based on dynamic programming according to claim 1, characterized in that, In the medium crossing phase, the cross-medium state is characterized by the relative submersion length, and the medium-related terms in the dynamic model are expressed based on the relative submersion length, so that the relative submersion length is 0 for the air state and 1 for the fully underwater state. FTDO observation and adaptive dynamic programming control calculations are performed based on the model of the medium crossing phase.