Inertia-like wheel hydrofoil and air wing integrated water-air cross-medium unmanned aerial vehicle and control algorithm thereof

Through the inertia wheel-like hydrofoil-air wing integrated structure and SAC-MPC control architecture, the structural complexity and control instability problems of water-air cross-medium UAVs are solved, efficient cross-medium flight stability and environmental adaptability are achieved, and the mission execution efficiency is improved.

CN120621740APending Publication Date: 2025-09-12BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510817249.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing water-air cross-medium UAVs have problems such as complex structure, low conversion efficiency, uncoordinated control, poor flight stability and delayed response, making it difficult to cope with the highly nonlinear and uncertain environment during the medium conversion process.

Method used

The system adopts an integrated structure of hydrofoils and aerofoils similar to an inertia wheel, and combines the Soft Actor-Critic (SAC) reinforcement learning algorithm with the multi-input multiple output (MIMO) control architecture of model predictive control (MPC) to achieve attitude control and propulsion in both water and air. By integrating the functions of hydrofoils and aerofoils, the system integration level is improved, and reinforcement learning is used to enhance the environmental adaptability of traditional control strategies.

Benefits of technology

It has achieved compact and integrated structure, strong control adaptability, good flight stability, and the ability to adapt to complex environments such as surges and airflow disturbances, thereby improving the stability of cross-media flight and mission execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120621740A_ABST
    Figure CN120621740A_ABST
Patent Text Reader

Abstract

The invention discloses an inertia-like wheel hydrofoil and air wing integrated water-air cross-medium unmanned aerial vehicle and a control algorithm thereof. A fixed wing air lift system and an adjustable hydrofoil propulsion assembly are integrated on the unmanned aerial vehicle structure, an underwater propulsion propeller is regarded as an inertia-like wheel, the attitude of the unmanned aerial vehicle is adjusted through rotation of the underwater propulsion propeller, and the high-dynamic control performance is achieved; the control method is based on a double-layer architecture combining Soft Actor-Critic (SAC) reinforcement learning and model predictive control (MPC), MPC output is corrected in real time by using a neural strategy network, and high-robustness attitude control in a complex environment is realized. The system is compact in structure, intelligent in control, high in adaptability and suitable for various cross-medium task scenes such as sea-air integrated inspection, search and rescue and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of water-air cross-medium UAVs, and specifically to an inertia wheel-hydrofoil-air wing integrated water-air cross-medium UAV and its control algorithm, belonging to the intersection of new water-air collaborative UAV design and intelligent control technology. Background Art

[0002] Water-to-air cross-medium UAVs are a rapidly developing cutting-edge technology in recent years, widely used in scenarios such as ocean exploration, rescue and search, and military reconnaissance. Currently, most water-to-air UAVs employ a split-body design, with the aerial drone and underwater propulsion systems independent of each other. This leads to complex structures, low conversion efficiency, and uncoordinated control. Furthermore, existing control methods often rely on single models or fixed controllers, making them difficult to handle the highly nonlinear and uncertain environment of the medium conversion process, resulting in poor flight stability and hysteresis.

[0003] Therefore, there is an urgent need for a new type of water-air cross-medium UAV with integrated structure, sensitive response and adaptive control capability and its supporting control algorithm. Summary of the Invention

[0004] The purpose of the present invention is to provide an inertia wheel-like hydrofoil-air wing integrated water-air cross-medium UAV and its control method based on the SAC and MPC fusion algorithm to solve the technical problems of the existing technology such as complex structure, unstable response, and difficult cross-medium transition control.

[0005] The present invention is achieved through the following technical solutions:

[0006] A water-air cross-medium UAV with an integrated inertia wheel, hydrofoil and airfoil, characterized in that it comprises an airfoil portion and a hydrofoil portion, wherein the hydrofoil portion and the airfoil portion are structurally integrated into one body to achieve attitude control and power propulsion in both water and air media;

[0007] The hydrofoil part includes: a wind resistance reduction column, a left hydrofoil, a left hydrofoil motor, a left propeller, a left connecting rod shaft, a left hydrofoil servo, a right hydrofoil servo, a right connecting rod shaft, a right propeller, a right hydrofoil motor, and a right hydrofoil;

[0008] A wind resistance reduction column is used to connect the airfoil and hydrofoil structure and optimize the fluid resistance during underwater propulsion;

[0009] The hydrofoil assembly is arranged symmetrically on both sides, including a left hydrofoil and a right hydrofoil, which are connected to the wind resistance reduction column through the left and right connecting rod shafts respectively, so that the hydrofoil can rotate around the shaft in a vertical plane;

[0010] The left and right hydrofoil motors are installed on the corresponding hydrofoils respectively, and generate axial rotation torque through the left and right propellers. The propellers and their motors are used together as an inertia wheel-like structure to assist in underwater attitude control;

[0011] The left and right hydrofoil servos are fixedly installed inside the wind resistance reduction column respectively. The two servos drive the hydrofoils on the corresponding sides through the rotating shaft to achieve adjustable pitch angle control.

[0012] The aerofoil portion is a fixed-wing structure, and its wing profile can be selected as a straight wing, a swept wing or a delta wing, which is used to provide the lift required for flight in the air. The wing profile selection can be flexibly configured according to mission requirements. The aerofoil and the hydrofoil work together to compensate each other during cross-media transition and attitude adjustment, thereby improving the dynamic stability and response performance of the UAV.

[0013] A control method for an inertia wheel-like hydrofoil-air wing integrated water-air cross-medium UAV is characterized by being based on a multiple-input multiple-output (MIMO) control architecture combining a soft actor-critic (SAC) reinforcement learning algorithm with model predictive control (MPC);

[0014] The control targets of the control architecture include the roll angle α, pitch angle β, and yaw angle of the UAV. γ and flight speed v;

[0015] The control inputs of the control architecture include the main power system speed u1=ω0, the horizontal tail pitch angle control value u2=η, the vertical tail yaw angle control value u3=μ, the left aileron folding angle u4=δ1, the right aileron folding angle u5=δ2, the left hydrofoil servo pitch angle u6=θ1, the right hydrofoil servo pitch angle u7=θ2, the left propeller speed u8=ω1, and the right propeller speed u9=ω2; wherein, u1=ω0 the main power system speed controls the flight speed v, u2=η the horizontal tail pitch angle controls the pitch angle β, u3=μ the vertical tail yaw angle controls the yaw angle γ, u4=δ1 the left aileron folding angle controls α, β, u5=δ2 the right aileron folding angle controls α, β; In the air, the left and right hydrofoil motors u8=ω1 and u9=ω2 drive the propellers to rotate, which can be used as two quasi-inertia wheels to provide reverse torsional torque to control the roll angle of the drone. Combined with the pitch angles of the left and right hydrofoil servos u6=θ1 and u7=θ2, they can be used to control the direction of the torsional torque to achieve pitch and yaw control of the drone. In water, the left and right hydrofoil motors u8=ω1 and u9=ω2 drive the propellers to rotate. The combination of the two can be used to control the drone's water gliding and underwater diving speed, and can also be used to control the drone's yaw angle. Combined with the pitch angles of the left and right hydrofoil servos u6=θ1 and u7=θ2, they can be used to control the drone's pitch and roll angles.

[0016] The control architecture consists of three parts: a UAV state-space system model under multi-media, a master controller based on model predictive control (MPC), and a disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm.

[0017] The state space system model of the UAV under the multi-media is:

[0018] y=C1x+D1u, in the air

[0019] y=C2x+D2u, in water

[0020] Among them, A1, B1, C1, D1 and A2, B2, C2, D2 are the state space matrices of the aerial and underwater drones respectively, and the state quantities can be taken as:

[0021]

[0022] The input is:

[0023]

[0024] The master controller based on model predictive control (MPC) uses the system dynamics model of the UAV to predict the future state, solves the control variable u to minimize the cost function (including minimizing the tracking error and minimizing the change of the control variable), and switches to different models for different media (air, water);

[0025] The cost function is designed as:

[0026]

[0027] Where Q is the state error weight matrix and R is the control input weight.

[0028] Solve for the optimal control input u under the following constraints k :

[0029] u min ≤u k ≤u max

[0030] The disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm learns the disturbance model or compensation term of the unknown dynamics; learns the control priority weight and adaptively adjusts the control weight; and serves as an auxiliary parameter adjustment or fault-tolerant compensation module of MPC;

[0031] Its state space is:

[0032]

[0033] The action space is a = Δu output by SAC:

[0034] a=[Δω0 Δη Δμ Δδ1 Δδ2 Δθ1 Δθ2 Δω1 Δω2] T

[0035] The reward function should prefer minimizing the posture error, smoothing the control amount, and minimizing the energy, and can be designed as:

[0036] r t =λ1||x t -x ref || 2 -λ2||Δu t || 2

[0037] Where: x ref =[α d ,β d ,γ d ,v d ] is the expected state, λ1,λ2 are weight coefficients;

[0038] The core objective function of SAC is:

[0039]

[0040] Where: r(·) is the reward function, is the strategy summary, α is the exploration-exploitation trade-off coefficient;

[0041] The final control quantity is

[0042] u t =u MPC +Δu SAC

[0043] Where: u MPC is the control quantity predicted by the model, Δu SAC It is the compensation or correction control amount obtained by SAC learning;

[0044] It is also characterized by using SAC to learn a policy network π φ (u|x); MPC is used as prior control knowledge (or initial strategy) to enable SAC to perform better in the early stages of learning; during training, MPC is used to generate expert trajectories as supervision signals for SAC; the controller can switch between different control models and strategy networks based on the current medium type (water or air), achieving continuous control and highly robust response of cross-medium flight attitude.

[0045] Compared with the existing technology, the present invention has the following advantages: compact and integrated structure: integrating the functions of hydrofoils and airfoils to improve system integration; strong control adaptability: using reinforcement learning to enhance the environmental adaptability of traditional control strategies; good flight stability: improving underwater stability through the combined action of inertia-like wheels and adjustable hydrofoils; adaptability to complex environments: having the ability to cope with sudden changes in fluid environments (such as surges and airflow disturbances). BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Overall schematic diagram of an inertia wheel-like hydrofoil-air wing integrated water-air cross-medium UAV

[0047] Figure 2 Schematic diagram of inertia wheel-like hydrofoil

[0048] Figure 3 Bottom view of inertia wheel hydrofoil

[0049] Figure 4 Schematic diagram of the front side of the inertia wheel-like hydrofoil

[0050] Figure 5 Schematic diagram of the working of inertia wheel-like hydrofoil

[0051] Figure 6 Control framework diagram of the inertia wheel hydrofoil-air wing integrated water-air cross-medium UAV

[0052] In the figure: a wind resistance reduction column (1), a left hydrofoil (2), a left hydrofoil motor (3), a left propeller (4), a left connecting rod shaft (5), a left hydrofoil steering gear (6), a right hydrofoil steering gear (7), a right connecting rod shaft (8), a right propeller (9), a right hydrofoil motor (10), a right hydrofoil (11), a hydrofoil portion (12), and an airfoil portion (13). DETAILED DESCRIPTION

[0053] The present invention will be further described with reference to the accompanying drawings:

[0054] A water-air cross-medium UAV with integrated inertia wheel, hydrofoil and air wing, characterized by: Figure 1 As shown, it includes an airfoil portion (13) and a hydrofoil portion (12), wherein the hydrofoil portion and the airfoil portion are structurally integrated to achieve attitude control and power propulsion in both water and air media;

[0055] The hydrofoil portion includes: Figure 2 ,3, wind resistance reduction column (1), left hydrofoil (2), left hydrofoil motor (3), left propeller (4), left connecting rod shaft (5), left hydrofoil servo (6), right hydrofoil servo (7), right connecting rod shaft (8), right propeller (9), right hydrofoil motor (10), right hydrofoil (11);

[0056] A wind resistance reduction column (1) is used to connect the airfoil and the hydrofoil structure and optimize the fluid resistance during the propulsion process in water;

[0057] A hydrofoil assembly arranged symmetrically on both sides includes a left hydrofoil (2) and a right hydrofoil (11), which are connected to the wind resistance reduction column (1) via left and right connecting rod shafts (5, 8) respectively, so that the hydrofoils can rotate around the shafts in a vertical plane;

[0058] The left hydrofoil motor (3) and the right hydrofoil motor (10) are respectively installed on the corresponding hydrofoils (2, 11), and generate axial rotation torque through the left and right propellers (4, 9). The propellers and their motors are used together as an inertia wheel-like structure to assist in underwater attitude control.

[0059] The left hydrofoil steering gear (6) and the right hydrofoil steering gear (7) are respectively fixedly installed inside the wind resistance reduction column (1), and the two steering gears respectively drive the hydrofoils on the corresponding sides through the rotating shaft to realize adjustable elevation angle control.

[0060] The aerofoil portion (13) is a fixed-wing structure, and its airfoil can be selected as a straight wing, a swept wing or a delta wing, and is used to provide the lift required for flight in the air. The airfoil selection can be flexibly configured according to mission requirements. The aerofoil and the hydrofoil work together to compensate each other during the cross-medium transition and attitude adjustment process, thereby improving the dynamic stability and response performance of the UAV.

[0061] A control method for a water-air cross-medium UAV with an integrated inertia wheel hydrofoil and air wing is characterized in that: Figure 6 As shown in the figure, this method is based on a multiple-input multiple-output (MIMO) control architecture that combines the Soft Actor-Critic (SAC) reinforcement learning algorithm with model predictive control (MPC);

[0062] The control targets of the control architecture include the roll angle α, pitch angle β, and yaw angle of the UAV. γ and flight speed v;

[0063] The control inputs of the control architecture include the main power system speed u1=ω0, the horizontal tail pitch angle control value u2=η, the vertical tail yaw angle control value u3=μ, the left aileron folding angle u4=δ1, the right aileron folding angle u5=δ2, the left hydrofoil servo pitch angle u6=θ1, the right hydrofoil servo pitch angle u7=θ2, the left propeller speed u8=ω1, and the right propeller speed u9=ω2; wherein, u1=ω0 the main power system speed controls the flight speed v, u2=η the horizontal tail pitch angle controls the pitch angle β, u3=μ the vertical tail yaw angle controls the yaw angle γ, u4=δ1 the left aileron folding angle controls α, β, u5=δ2 the right aileron folding angle controls α, β; In the air, the left and right hydrofoil motors u8=ω1 and u9=ω2 drive the propellers to rotate, which can be used as two quasi-inertia wheels to provide reverse torsional torque to control the roll angle of the drone. Combined with the pitch angles of the left and right hydrofoil servos u6=θ1 and u7=θ2, they can be used to control the direction of the torsional torque to achieve pitch and yaw control of the drone. In water, the left and right hydrofoil motors u8=ω1 and u9=ω2 drive the propellers to rotate. The combination of the two can be used to control the drone's water gliding and underwater diving speed, and can also be used to control the drone's yaw angle. Combined with the pitch angles of the left and right hydrofoil servos u6=θ1 and u7=θ2, they can be used to control the drone's pitch and roll angles.

[0064] The control architecture consists of three parts: a UAV state-space system model under multi-media, a master controller based on model predictive control (MPC), and a disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm.

[0065] The state space system model of the UAV under the multi-media is:

[0066] y=C1x+D1u, in the air

[0067] y=C2x+D2u, in water

[0068] Among them, A1, B1, C1, D1 and A2, B2, C2, D2 are the state space matrices of the aerial and underwater drones respectively, and the state quantities can be taken as:

[0069]

[0070] The input is:

[0071]

[0072] The master controller based on model predictive control (MPC) uses the system dynamics model of the UAV to predict the future state, solves the control variable u to minimize the cost function (including minimizing the tracking error and minimizing the change of the control variable), and switches to different models for different media (air, water);

[0073] The cost function is designed as:

[0074]

[0075] Where Q is the state error weight matrix and R is the control input weight.

[0076] Solve for the optimal control input u under the following constraints k :

[0077] u min ≤u k ≤u max

[0078] The disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm learns the disturbance model or compensation term of the unknown dynamics; learns the control priority weight and adaptively adjusts the control weight; and serves as an auxiliary parameter adjustment or fault-tolerant compensation module of MPC;

[0079] Its state space is:

[0080]

[0081] The action space is a = Δu output by SAC:

[0082] a=[Δω0 Δη Δμ Δδ1 Δδ2 Δθ1 Δθ2 Δω1 Δω2] T

[0083] The reward function should prefer minimizing the posture error, smoothing the control amount, and minimizing the energy, and can be designed as:

[0084] r t =λ1||x t -x ref || 2 -λ2||Δu t || 2

[0085] Where: x ref =[α d ,β d ,γ d ,v d ] is the expected state, λ1,λ2 are weight coefficients;

[0086] The core objective function of SAC is:

[0087]

[0088] Where: r(·) is the reward function, is the strategy summary, α is the exploration-exploitation trade-off coefficient;

[0089] The final control quantity is

[0090] u t =u MPC +Δu SAC

[0091] Where: u MPC is the control quantity predicted by the model, Δu SAC It is the compensation or correction control amount obtained by SAC learning;

[0092] It is also characterized by using SAC to learn a policy network π φ (u|x); MPC is used as prior control knowledge (or initial strategy) to enable SAC to perform better in the early stages of learning; during training, MPC is used to generate expert trajectories as supervision signals for SAC; the controller can switch between different control models and strategy networks based on the current medium type (water or air), achieving continuous control and highly robust response of cross-medium flight attitude.

[0093] The present invention highly integrates the hydrofoil, airfoil and propulsion system, proposes an inertia wheel control principle, and combines SAC reinforcement learning with MPC predictive control algorithm to construct a new water-air UAV with cross-media autonomous flight capability, compact structure, precise control and strong environmental adaptability. It effectively solves the problems of structural separation, conversion hysteresis and control instability in the existing technology, and significantly improves the flight stability and mission execution efficiency in complex environments.

Claims

1. A water-air cross-medium UAV with integrated inertia wheel, hydrofoil and airfoil, characterized by: It includes an airfoil part and a hydrofoil part, wherein the hydrofoil part and the airfoil part are structurally integrated to realize attitude control and power propulsion in both water and air media; The hydrofoil part includes: a wind resistance reduction column, a left hydrofoil, a left hydrofoil motor, a left propeller, a left connecting rod shaft, a left hydrofoil servo, a right hydrofoil servo, a right connecting rod shaft, a right propeller, a right hydrofoil motor, and a right hydrofoil; A wind resistance reduction column is used to connect the airfoil and hydrofoil structure and optimize the fluid resistance during underwater propulsion; The hydrofoil assembly is arranged symmetrically on both sides, including a left hydrofoil and a right hydrofoil, which are connected to the wind resistance reduction column through the left and right connecting rod shafts respectively, so that the hydrofoil can rotate around the shaft in a vertical plane; The left and right hydrofoil motors are installed on the corresponding hydrofoils respectively, and generate axial rotation torque through the left and right propellers. The propellers and their motors are used together as an inertia wheel-like structure to assist in underwater attitude control; The left and right hydrofoil servos are fixedly installed inside the wind resistance reduction column respectively. The two servos drive the corresponding hydrofoils through the rotating shaft to achieve adjustable pitch angle control; The aerofoil portion is a fixed-wing structure, and its wing profile can be selected as a straight wing, a swept wing or a delta wing, which is used to provide the lift required for flight in the air. The wing profile selection can be flexibly configured according to mission requirements. The aerofoil and the hydrofoil work together to compensate each other during cross-media transition and attitude adjustment, thereby improving the dynamic stability and response performance of the UAV.

2. An adaptive control method for the water-air cross-medium UAV according to claim 1, characterized in that: The method is based on a multiple-input multiple-output (MIMO) control architecture that combines the Soft Actor-Critic (SAC) reinforcement learning algorithm with model predictive control (MPC); The control targets of the control architecture include the roll angle α, pitch angle β, yaw angle γ and flight speed v of the UAV; The control inputs of the control architecture include the main power system speed u1=ω0, the horizontal tail pitch angle control value u2=η, the vertical tail yaw angle control value u3=μ, the left aileron folding angle u4=δ1, the right aileron folding angle u5=δ2, the left hydrofoil servo pitch angle u6=θ1, the right hydrofoil servo pitch angle u7=θ2, the left propeller speed u8=ω1, and the right propeller speed u9=ω2; Among them, u1=ω0 main power system speed controls flight speed v, u2=η horizontal tail pitch angle controls pitch angle β, u3=μ vertical tail yaw angle controls yaw angle γ, u4=δ1 left aileron folding angle controls α, β, u5=δ2 right aileron folding angle controls α, β; in the air, u8=ω1 and u9=ω2 left and right hydrofoil motors drive propellers to rotate and can be used as two inertia wheels to provide reverse torsional torque to control the roll angle of the drone, and then cooperate with u6 =θ1 and u7=θ2, the pitch angles of the left and right hydrofoil servos can be used to control the direction of the torsional moment, thereby controlling the pitch and yaw angles of the drone. In water, u8=ω1 and u9=ω2, the left and right hydrofoil motors drive the propellers to rotate. The combination of the two can be used to control the gliding and underwater diving speeds of the drone, and can also be used to control the yaw angle of the drone. Together with u6=θ1 and u7=θ2, the pitch angles of the left and right hydrofoil servos can be used to control the pitch and roll angles of the drone. The control architecture consists of three parts: a UAV state-space system model under multi-media, a master controller based on model predictive control (MPC), and a disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm. The state space system model of the UAV under the multi-media is: y=C1x+D1u, in the air y=C2x+D2u, in water Among them, A1, B1, C1, D1 and A2, B2, C2, D2 are the state space matrices of the aerial and underwater drones respectively, and the state quantities can be taken as: The input is: The master controller based on model predictive control (MPC) uses the system dynamics model of the UAV to predict the future state, solves the control variable u to minimize the cost function (including minimizing the tracking error and minimizing the change of the control variable), and switches to different models for different media (air, water); The cost function is designed as: Where Q is the state error weight matrix, R is the control input weight; Solve for the optimal control input u under the following constraints k : in min in k in max The disturbance compensator based on the reinforcement learning Soft Actor-Critic (SAC) algorithm learns the disturbance model or compensation term of the unknown dynamics; learns the control priority weight and adaptively adjusts the control weight; and serves as an auxiliary parameter adjustment or fault-tolerant compensation module of MPC; Its state space is: The action space is a = Δu output by SAC: a=[Δω0 Δη Δμ Δδ1 Δδ2 Δθ1 Δθ2 Δω1 Δω2] T The reward function should prefer minimizing the posture error, smoothing the control amount, and minimizing the energy, and can be designed as: r t =λ1||x t -x ref || 2 -λ2||Δu t || 2 Where: x ref =[α d ,β d ,γ d ,v d ] is the expected state, λ1,λ2 are weight coefficients; The core objective function of SAC is: Where: r(·) is the reward function, H is the strategy summary, and α is the exploration-exploitation trade-off coefficient; The final control quantity is u t =u MPC +Δu SAC Where: u MPC is the control quantity predicted by the model, Δu SAC It is the compensation or correction control amount obtained by SAC learning; It is also characterized by using SAC to learn a policy network π φ (u|x); MPC is used as prior control knowledge (or initial strategy) to enable SAC to perform better in the early stages of learning; during training, MPC is used to generate expert trajectories as supervision signals for SAC; the controller can switch between different control models and strategy networks based on the current medium type (water or air), achieving continuous control and highly robust response of cross-medium flight attitude.

3. A control method according to claim 2, characterized in that: The control structure is a two-layer control system integrating MPC and SAC, which specifically includes the following steps: (1) Establish a state space model. The state variables include the attitude angle (roll angle, pitch angle, yaw angle) and speed of the UAV, as well as the current control state of each actuator; (2) Based on the state model, an MPC controller is constructed to perform rolling optimization with input constraints and output a predictive control reference value; (3) Constructing a SAC neural network strategy structure based on the maximum entropy reinforcement learning framework. The input includes the current state, MPC output, and disturbance estimation value, and the output compensation is used to enhance the MPC control effect. (4) The fused control quantity is linearly superimposed and acts on each control actuator of the UAV, achieving high dynamic accuracy and continuous attitude stability control under cross-media switching; (5) In the reinforcement learning stage, the control strategy in the disturbance environment is pre-trained using the simulator, and the strategy network is continuously fine-tuned during actual flight to achieve model adaptive evolution and environmental robustness compensation.