Composite wing unmanned aerial vehicle transition process control method based on reinforcement learning

By constructing a transition process control model for a compound-wing UAV based on the PPO reinforcement learning algorithm, the dynamic challenges of the UAV during the transition between rotor and fixed-wing modes were solved, achieving stable and energy-efficient control.

CN121806995APending Publication Date: 2026-04-07SHANGHAI AEROSPACE CONTROL TECH INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the highly nonlinear, strongly coupled, and time-varying dynamic characteristics of compound-wing UAVs during the transition between rotor and fixed-wing modes, leading to decreased control capabilities. Furthermore, there is a lack of effective methods for applying reinforcement learning in this field.

Method used

A transition process control model for a compound-wing UAV is constructed using the PPO reinforcement learning algorithm. By designing a state space and action space that match physical meaning and dimension, and combining it with a reward function, a reinforcement learning agent is trained to achieve a weighted distribution of virtual control force and torque in rotor and fixed-wing modes. A classic controller is used to ensure the stability of the subsystem, and a stable, smooth, and energy-efficient control of the transition process is quickly achieved.

Benefits of technology

It achieves stable and efficient energy control during the transition process of compound-wing UAVs, avoiding the design difficulties of traditional methods under complex working conditions and the learning convergence difficulties of reinforcement learning, and ensuring a smooth transition in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121806995A_ABST
    Figure CN121806995A_ABST
Patent Text Reader

Abstract

The invention discloses a composite wing unmanned aerial vehicle transition process control method based on reinforcement learning, and the method comprises the steps: building a composite wing unmanned aerial vehicle transition process control model based on a PPO reinforcement learning algorithm; in the transition process of the composite wing unmanned aerial vehicle, rotor wing virtual control force and torque output by a rotor wing mode controller, fixed wing virtual control force and torque output by a fixed wing mode controller, and rotor wing virtual control weight and fixed wing virtual control weight output by a composite wing unmanned aerial vehicle transition process control model are obtained; and after weighting processing is carried out, distribution execution is carried out, composite wing unmanned aerial vehicle transition process control parameters are obtained, and control over the composite wing unmanned aerial vehicle transition process is completed. According to the method disclosed by the invention, the stability of the subsystem is guaranteed by a classical control algorithm, and the method is oriented to multi-performance index optimization, so that the network training convergence of the transition process can be quickly realized, and the stable and high-energy-efficiency control of the strong environment adaptability of the transition process of the composite wing unmanned aerial vehicle is comprehensively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of compound-wing unmanned aerial vehicle (UAV) control technology, and particularly relates to a method for controlling the transient process of a compound-wing UAV based on reinforcement learning. Background Technology

[0002] Compound-wing UAVs are hybrid-powered aircraft that combine the vertical takeoff and landing (VTOL) and hovering capabilities of multi-rotor aircraft with the long-endurance and efficient cruise capabilities of fixed-wing aircraft, and have broad application prospects. Compound-wing UAVs generally have three flight modes: rotor mode, fixed-wing mode, and a transition process between the two modes. Rotor mode primarily enables VTOL and hovering at low speeds, while fixed-wing mode primarily achieves energy-saving and long-endurance flight at low and high speeds. Both modes have mature controller design algorithms to achieve effective control under corresponding single-type actuator conditions. Although the transition process accounts for a very small percentage of the flight time, the highly nonlinear, strongly coupled, and time-varying characteristics of the aircraft's dynamics during this process are one of the biggest challenges to the stable flight of compound-wing UAVs. It is necessary to simultaneously consider the control capability degradation caused by low-speed stall of the fixed-wing wing and high-speed stall of the rotor blades, and achieve a safe and smooth transition between rotor and fixed-wing modes through reasonable mode-switching strategies or the design of redundant actuators for joint control.

[0003] Current control methods mainly include system switching and control allocation. System switching primarily involves dividing dynamic feature points based on actual task requirements, designing controllers for different feature subsystems accordingly, and then using appropriate switching laws to improve system performance, such as linear switching and fuzzy switching. However, this approach is difficult to design for complex conditions with too many changing feature points. Control allocation, on the other hand, fully utilizes redundant actuators and employs methods such as chain-incremental methods, pseudo-inverse methods, dynamic programming, and model prediction to optimize performance indicators under constraints, thereby improving system reliability. Reinforcement learning, as a popular optimization method in recent years, has been widely applied in aircraft control due to its ability to handle complex scenarios and adaptive optimization. However, research on its application in the transition process control of compound-wing UAVs is lacking. Summary of the Invention

[0004] The technical problem solved by this invention is to overcome the shortcomings of the prior art and provide a transition process control method for compound-wing UAVs based on reinforcement learning. This method aims to overcome the deficiencies of existing system switching and control allocation techniques. While ensuring the stability of subsystems through classical control algorithms, it optimizes multiple performance indicators and can quickly achieve network training convergence during the transition process. This comprehensively realizes stable, smooth, and energy-efficient control of compound-wing UAVs with strong environmental adaptability during the transition process.

[0005] To address the aforementioned technical problems, this invention discloses a transition process control method for a compound-wing unmanned aerial vehicle based on reinforcement learning, comprising: Based on the PPO reinforcement learning algorithm, a transition process control model for a compound-wing UAV was constructed. During the transition process of the compound wing UAV, the virtual control force and torque of the rotor mode controller, the virtual control force and torque of the fixed wing mode controller, and the virtual control weights of the rotor and the fixed wing output by the control model of the transition process of the compound wing UAV are obtained. The virtual control forces and torques of the rotor and the fixed wing are weighted with the virtual control weights of the rotor and the fixed wing, respectively. Based on the weighting results, the parameters are allocated and executed to obtain the control parameters for the transition process of the compound wing UAV, thus completing the control of the transition process of the compound wing UAV.

[0006] In the aforementioned reinforcement learning-based transition process control method for compound-wing UAVs, the transition process control parameters for the compound-wing UAV include: rotor motor throttle. Fixed-wing engine throttle 1. Fixed-wing rudder deflection angles in each channel ;in, Indicates the first One rotor motor throttle, , Indicates the number of rotor motors; This indicates the normalized rudder deflection angle of the fixed-wing pitch channel. This indicates the normalized rudder deflection angle for the fixed-wing roll path. This indicates the normalized rudder deflection angle of the yaw channel for a fixed-wing aircraft.

[0007] In the aforementioned reinforcement learning-based transition process control method for compound-wing UAVs, ,

[0008] in, , , and This represents the virtual control weights of the four rotors; , , and This represents the four virtual control weights for fixed-wing aircraft. , and The three-axis components represent the virtual control torque of the rotor. This represents the virtual control force of the rotor; , and The three-axis components represent the virtual control torque of a fixed-wing aircraft. This represents the virtual control force of a fixed-wing aircraft. This represents the rotor actuator allocation matrix that conforms to the dimensions. This represents the fixed-wing actuator allocation matrix that conforms to the dimensions.

[0009] In the aforementioned reinforcement learning-based transition process control method for compound-wing UAVs, and The percentage range for each is 0 to 1; , , , , , , and The value range is 0 to 1.

[0010] In the above reinforcement learning-based transition process control method for compound-wing UAVs, a transition process control model for the compound-wing UAV is constructed based on the PPO reinforcement learning algorithm, including: S1, based on the PPO reinforcement learning algorithm, constructs a reinforcement learning agent; where the reinforcement learning agent acts as the overall control unit, and understands the environment by perceiving the state space; the decision-making mechanism inside the reinforcement learning agent is jointly operated by the action network and the evaluation network, which calculates actions and outputs them to the action space based on the feedback of the current state and the reward function; after the action is applied to the environment, the environment generates new states and rewards, and the reinforcement learning agent continuously optimizes its strategy. S2, designing a reinforcement learning state space and action space that match physical meaning and dimension; S3, Design the reward function based on task metrics and optimization requirements; S4, Design the size and structure of the action network and evaluation network; S5, train the reinforcement learning agent and determine whether the convergence and control task requirements are met; if not, return to step S3 to adjust the design until the training converges and the control task requirements are met; if so, use the trained reinforcement learning agent as the transition process control model for the compound wing UAV.

[0011] In the above-mentioned reinforcement learning-based transition process control method for compound wing UAVs, the reinforcement learning state space includes: roll angle command, pitch angle command, yaw angle command, roll angle feedback, pitch angle feedback, yaw angle feedback, roll angular velocity, pitch angular velocity, yaw angular velocity, speed command, airspeed, angle of attack, and sideslip angle.

[0012] In the aforementioned reinforcement learning-based transient control method for compound-wing UAVs, the reinforcement learning action space includes eight independent, continuous control weights: four rotor virtual control weights. and 4 fixed-wing virtual control weights .

[0013] In the aforementioned reinforcement learning-based transition process control method for compound-wing UAVs, the reward function is expressed as follows:

[0014] in, Indicates the total reward. This represents the x-axis attitude channel reward. This represents the y-axis attitude channel reward. This represents the z-axis attitude channel reward. Indicates speed channel reward, Indicates executor reward; x-axis attitude channel reward It is expressed as follows:

[0015] in, This represents the non-negative steady-state reward parameter along the x-axis. This is represented as a non-negative penalty parameter for inconsistent x-axis deviation angle direction. This represents the non-negative x-axis steady-state angular velocity deviation penalty parameter. Indicates the roll angular velocity. This represents the steady-state angular velocity constraint parameter along the x-axis. Indicates roll angle deviation; Represents the steady-state region along the x-axis. , Indicates the initial roll angle deviation. This represents the standard constraint parameter for x-axis overshoot. This represents the steady-state error constraint parameter along the x-axis; y-axis attitude channel reward It is expressed as follows:

[0016] in, This represents the non-negative steady-state reward parameter along the y-axis. This is represented as a non-negative penalty parameter for inconsistent y-axis deviation angle direction. This represents the penalty parameter for non-negative steady-state angular velocity along the y-axis. Indicates pitch angular velocity, This represents the steady-state angular velocity constraint parameter along the y-axis. Indicates pitch angle deviation; Indicates the steady-state region along the y-axis. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for y-axis overshoot. This represents the steady-state error constraint parameter along the y-axis; z-axis attitude channel reward It is expressed as follows:

[0017] in, This represents the non-negative steady-state reward parameter along the z-axis. This is represented as a non-negative penalty parameter for inconsistent z-axis deviation angle direction. This represents the penalty parameter for non-negative steady-state angular velocity deviation along the z-axis. Indicates yaw rate, This represents the z-axis steady-state angular velocity constraint parameter. Indicates yaw angle deviation; Represents the steady-state region along the z-axis. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for z-axis overshoot. This represents the z-axis steady-state error constraint parameter; Speed ​​Channel Rewards It is expressed as follows:

[0018] in, This represents the non-negative steady-state reward parameter. This indicates the deviation between the speed command and the feedback. Executor reward It is expressed as follows:

[0019] in, , , , , , and These represent the non-negative penalty parameters used for performance optimization of different performance metrics; express Time of the first One rotor motor throttle, express Constantly maintain the throttle position of the fixed-wing engine. express Constantly adjust the pitch channel rudder deflection angle of the fixed-wing aircraft. express Constantly fix the wing roll channel rudder deflection angle. express Constantly fix the yaw channel rudder deflection angle of the wing. This indicates the time interval for the controller to process data.

[0020] In the aforementioned reinforcement learning-based transition process control method for compound-wing UAVs, the size and structure of the action network and evaluation network are as follows: Action network: The input layer dimension is the same as the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output layer dimension is the same as the action space dimension. Network evaluation: The input layer dimension is consistent with the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output is a scalar value.

[0021] The aforementioned reinforcement learning-based transition process control method for compound-wing UAVs also includes: independently designing the rotor mode controller and fixed-wing mode controller for the compound-wing UAV based on classical controller design methods.

[0022] The present invention has the following advantages: This invention discloses a reinforcement learning-based transient control method for compound-wing unmanned aerial vehicles (UAVs). This method, while retaining the classic rotor and fixed-wing control systems, utilizes reinforcement learning to distribute virtual control forces (torques) during the transient process. This avoids the design difficulties of traditional transient system switching in complex operating conditions, and also avoids the problems of learning convergence difficulties and control stability assurance issues that arise from directly using reinforcement learning for end-to-end control. Furthermore, the reward function design incorporates convergence direction guidance in addition to the classic error convergence reward, enabling rapid training convergence. Simultaneously, it considers actuator energy saving and variation penalties, comprehensively achieving stable, smooth, and energy-efficient control with strong environmental adaptability during the transient process of the compound-wing UAV. Attached Figure Description

[0023] Figure 1 This is a flowchart of a transition process control method for a compound-wing unmanned aerial vehicle based on reinforcement learning, according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments disclosed herein will be described in further detail below with reference to the accompanying drawings.

[0025] Reference Figure 1 In this embodiment, the reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles includes: S0, based on the classic controller design method, independently designs the rotor mode controller and fixed-wing mode controller for a compound wing UAV.

[0026] In this embodiment, both the rotor mode controller and the fixed-wing mode controller are classic controllers designed using classical controller design methods, including but not limited to non-optimized real-time control methods such as PID control and total energy control. Various input desired pose commands can be easily converted into speed and three-axis attitude angle control. Both the rotor mode controller and the fixed-wing mode controller need to independently ensure stable control of attitude, altitude, speed, and other commands in their respective modes. Furthermore, the outputs of both the rotor mode controller and the fixed-wing mode controller do not directly affect actuator actions, facilitating subsequent reinforcement learning design: the rotor mode controller output does not directly affect the throttle of each motor, and the fixed-wing mode controller output does not directly affect the engine throttle and the deflection angle of each channel. The outputs of both the rotor mode controller and the fixed-wing mode controller are represented as virtual control forces and torques, and can be converted into actuator actions through a fixed allocation strategy.

[0027] S1. Based on the PPO reinforcement learning algorithm, a transition process control model for a compound-wing UAV is constructed.

[0028] In this embodiment, the transition process control model for the compound-wing UAV can be constructed and trained in the following manner: S11, based on the PPO reinforcement learning algorithm, constructs a reinforcement learning agent.

[0029] As the central control unit, the reinforcement learning agent understands the environment by perceiving the state space. The decision-making mechanism within the reinforcement learning agent consists of the collaborative work of the action network and the evaluation network. Based on the feedback from the current state and the reward function, the agent calculates actions and outputs them to the action space. After the actions are applied to the environment, the environment generates new states and rewards, and the reinforcement learning agent continuously optimizes its strategy.

[0030] S12, design reinforcement learning state space and action space that match physical meaning and dimension.

[0031] The reinforcement learning state space mainly includes 13 state variables: roll angle command, pitch angle command, yaw angle command, roll angle feedback, pitch angle feedback, yaw angle feedback, roll rate, pitch rate, yaw rate, speed command, airspeed, angle of attack, and sideslip angle.

[0032] The reinforcement learning action space includes eight independent, consecutive control weights: four rotor virtual control weights. and 4 fixed-wing virtual control weights .in, , , , , , , and The value range is 0 to 1.

[0033] S13. Design the reward function based on the task indicators and optimization requirements.

[0034] The reward function is expressed as follows:

[0035] in, Indicates the total reward. This represents the x-axis attitude channel reward. This represents the y-axis attitude channel reward. This represents the z-axis attitude channel reward. Indicates speed channel reward, This indicates the executor reward.

[0036] x-axis attitude channel reward It is expressed as follows:

[0037] in, This represents the non-negative steady-state reward parameter along the x-axis. It is represented as a non-negative x-axis deviation angle direction inconsistency penalty parameter, used to guide the roll angular velocity integral to develop in the direction of reducing angle deviation during the process; This represents the non-negative x-axis steady-state angular velocity deviation penalty parameter. Indicates the roll angular velocity. This represents the steady-state angular velocity constraint parameter along the x-axis. Indicates roll angle deviation; Representing the steady-state region along the x-axis, considering the relative changes in the steady-state error and overshoot constraint under different step spans, the minimum of the two is taken, i.e. , Indicates the initial roll angle deviation. This represents the standard constraint parameter for x-axis overshoot. This represents the steady-state error constraint parameter along the x-axis.

[0038] y-axis attitude channel reward It is expressed as follows:

[0039] in, This represents the non-negative steady-state reward parameter along the y-axis. It is represented as a non-negative penalty parameter for inconsistent y-axis deviation angle direction, used to guide the mid-process to ensure that the integral of pitch angular velocity develops in the direction of reducing angular deviation; This represents the penalty parameter for non-negative steady-state angular velocity along the y-axis. Indicates pitch angular velocity, This represents the steady-state angular velocity constraint parameter along the y-axis. Indicates pitch angle deviation; Representing the steady-state region along the y-axis, considering the relative changes in the steady-state error and overshoot constraint under different step spans, the minimum of the two is taken, i.e. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for y-axis overshoot. This represents the steady-state error constraint parameter along the y-axis.

[0040] z-axis attitude channel reward It is expressed as follows:

[0041] in, This represents the non-negative steady-state reward parameter along the z-axis. It is represented as a non-negative z-axis deviation angle direction inconsistency penalty parameter, used to guide the mid-process to ensure that the yaw rate integral develops in the direction of reducing angle deviation; This represents the penalty parameter for non-negative steady-state angular velocity deviation along the z-axis. Indicates yaw rate, This represents the z-axis steady-state angular velocity constraint parameter. Indicates yaw angle deviation; Representing the steady-state region along the z-axis, considering the relative changes in the steady-state error and overshoot constraint under different step spans, the minimum of the two is taken, i.e. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for z-axis overshoot. This represents the z-axis steady-state error constraint parameter.

[0042] Speed ​​Channel Rewards It is expressed as follows:

[0043] in, This represents the non-negative steady-state reward parameter. This indicates the deviation between the speed command and the feedback.

[0044] Executor reward It is expressed as follows:

[0045] in, , , , , , and These represent the non-negative penalty parameters used for performance optimization of different performance metrics; express Time of the first One rotor motor throttle, express Constantly maintain the throttle position of the fixed-wing engine. express Constantly adjust the pitch channel rudder deflection angle of the fixed-wing aircraft. express Constantly fix the wing roll channel rudder deflection angle. express Constantly fix the yaw channel rudder deflection angle of the wing. This indicates the time interval for the controller to process data.

[0046] S14, Design the size and structure of the action network and evaluation network.

[0047] Action network: The input layer dimension is the same as the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output layer dimension is the same as the action space dimension. Network evaluation: The input layer dimension is consistent with the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output is a scalar value.

[0048] Based on design experience, network training generally starts with fewer hidden layers and neurons. If it is found that the learning speed of the reinforcement learning agent is too slow or the performance is difficult to improve, it is advisable to appropriately increase the network depth (i.e., the number of hidden layers) or increase the network width (i.e., the number of neurons).

[0049] S15, train the reinforcement learning agent and determine whether the convergence and control task requirements are met; if not, return to step S13 to adjust the design until the training converges and the control task requirements are met; if so, use the trained reinforcement learning agent as the transition process control model for the compound wing UAV.

[0050] If the training curve shows oscillations in the reward curve (i.e., no sustained growth trend and drastic fluctuations), or if the reward curve diverges (i.e., the reward continues to decline or suddenly collapses), it indicates that the training is unstable and it is necessary to appropriately increase the depth or width of the action network and evaluation network, or adjust hyperparameters such as reducing the learning rate. If the training curve has stabilized, but the task performance indicators do not meet the requirements, then the penalty parameters corresponding to the reward function should be adjusted.

[0051] S2, during the transition process of the compound wing UAV, acquires the virtual control force and torque of the rotor mode controller, the virtual control force and torque of the fixed wing mode controller, and the virtual control weights of the rotor and the fixed wing output by the control model of the transition process of the compound wing UAV.

[0052] S3 weights the virtual control force and torque of the rotor and the virtual control force and torque of the fixed wing with the virtual control weights of the rotor and the fixed wing, respectively, and allocates the control parameters for the transition process of the compound wing UAV based on the weighting result, thus completing the control of the transition process of the compound wing UAV.

[0053] In this embodiment, the control parameters for the transition process of the compound-wing UAV mainly include: rotor motor throttle. Fixed-wing engine throttle 1. Fixed-wing rudder deflection angles in each channel ;in, Indicates the first One rotor motor throttle, , Indicates the number of rotor motors; This indicates the normalized rudder deflection angle of the fixed-wing pitch channel. This indicates the normalized rudder deflection angle for the fixed-wing roll path. This indicates the normalized rudder deflection angle of the yaw channel for a fixed-wing aircraft.

[0054] The relationship between the virtual control force and torque of the rotor, the virtual control force and torque of the fixed wing, the virtual control weight of the rotor, the virtual control weight of the fixed wing, and the transient control parameters of the compound wing UAV is expressed as follows: ,

[0055] in, , and The three-axis components represent the virtual control torque of the rotor. This represents the virtual control force of the rotor; , and The three-axis components represent the virtual control torque of a fixed-wing aircraft. This represents the virtual control force of a fixed-wing aircraft. This represents the rotor actuator allocation matrix that conforms to the dimensions. Represents the fixed-wing actuator allocation matrix conforming to the dimensions; and The percentage ranges from 0 to 1.

[0056] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.

[0057] The contents not described in detail in this specification are common knowledge to those skilled in the art.

Claims

1. A transient control method for a compound-wing unmanned aerial vehicle based on reinforcement learning, characterized in that, include: Based on the PPO reinforcement learning algorithm, a transition process control model for a compound-wing UAV was constructed. During the transition process of the compound wing UAV, the virtual control force and torque of the rotor mode controller, the virtual control force and torque of the fixed wing mode controller, and the virtual control weights of the rotor and the fixed wing output by the control model of the transition process of the compound wing UAV are obtained. The virtual control forces and torques of the rotor and the fixed wing are weighted with the virtual control weights of the rotor and the fixed wing, respectively. Based on the weighting results, the parameters are allocated and executed to obtain the control parameters for the transition process of the compound wing UAV, thus completing the control of the transition process of the compound wing UAV.

2. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 1, characterized in that, The control parameters for the transition process of the compound-wing UAV include: rotor motor throttle. Fixed-wing engine throttle 1. Fixed-wing rudder deflection angles in each channel ;in, Indicates the first One rotor motor throttle, , Indicates the number of rotor motors; This indicates the normalized rudder deflection angle of the fixed-wing pitch channel. This indicates the normalized rudder deflection angle for the fixed-wing roll path. This indicates the normalized rudder deflection angle of the yaw channel for a fixed-wing aircraft.

3. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 2, characterized in that, , in, , , and This represents the virtual control weights of the four rotors; , , and This represents the virtual control weights for four fixed-wing aircraft. , and The three-axis components represent the virtual control torque of the rotor. This represents the virtual control force of the rotor; , and The three-axis components represent the virtual control torque of a fixed-wing aircraft. This represents the virtual control force of a fixed-wing aircraft. This represents the rotor actuator allocation matrix that conforms to the dimensions. This represents the fixed-wing actuator allocation matrix that conforms to the dimensions.

4. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 2, characterized in that, and The percentage range for each is 0 to 1; , , , , , , and The value range is 0 to 1.

5. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 1, characterized in that, Based on the PPO reinforcement learning algorithm, a transition process control model for a compound-wing UAV is constructed, including: S1, based on the PPO reinforcement learning algorithm, constructs a reinforcement learning agent; where the reinforcement learning agent acts as the overall control unit, and understands the environment by perceiving the state space; the decision-making mechanism inside the reinforcement learning agent is jointly operated by the action network and the evaluation network, which calculates actions and outputs them to the action space based on the feedback of the current state and the reward function; after the action is applied to the environment, the environment generates new states and rewards, and the reinforcement learning agent continuously optimizes its strategy. S2, designing a reinforcement learning state space and action space that match physical meaning and dimension; S3, Design the reward function based on task metrics and optimization requirements; S4, Design the size and structure of the action network and evaluation network; S5, train the reinforcement learning agent and determine whether the convergence and control task requirements are met; if not, return to step S3 to adjust the design until the training converges and the control task requirements are met; if so, use the trained reinforcement learning agent as the transition process control model for the compound wing UAV.

6. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 5, characterized in that, The reinforcement learning state space includes: roll angle command, pitch angle command, yaw angle command, roll angle feedback, pitch angle feedback, yaw angle feedback, roll rate, pitch rate, yaw rate, speed command, airspeed, angle of attack, and sideslip angle.

7. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 5, characterized in that, The reinforcement learning action space includes eight independent, sequential control weights: four rotor virtual control weights. and 4 fixed-wing virtual control weights .

8. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 5, characterized in that, The reward function is expressed as follows: in, Indicates the total reward. This represents the x-axis attitude channel reward. This represents the y-axis attitude channel reward. This represents the z-axis attitude channel reward. Indicates speed channel reward, Indicates executor reward; x-axis attitude channel reward It is expressed as follows: in, This represents the non-negative steady-state reward parameter along the x-axis. This is represented as a non-negative penalty parameter for inconsistent x-axis deviation angle direction. This represents the non-negative x-axis steady-state angular velocity deviation penalty parameter. Indicates the roll angular velocity. This represents the steady-state angular velocity constraint parameter along the x-axis. Indicates roll angle deviation; Represents the steady-state region along the x-axis. , Indicates the initial roll angle deviation. This represents the standard constraint parameter for x-axis overshoot. This represents the steady-state error constraint parameter along the x-axis; y-axis attitude channel reward It is expressed as follows: in, This represents the non-negative steady-state reward parameter along the y-axis. This is represented as a non-negative penalty parameter for inconsistent y-axis deviation angle direction. This represents the penalty parameter for non-negative steady-state angular velocity along the y-axis. Indicates pitch angular velocity, This represents the steady-state angular velocity constraint parameter along the y-axis. Indicates pitch angle deviation; Indicates the steady-state region along the y-axis. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for y-axis overshoot. This represents the steady-state error constraint parameter along the y-axis; z-axis attitude channel reward It is expressed as follows: in, This represents the non-negative steady-state reward parameter along the z-axis. This is represented as a non-negative penalty parameter for inconsistent z-axis deviation angle direction. This represents the penalty parameter for non-negative steady-state angular velocity deviation along the z-axis. Indicates yaw rate, This represents the z-axis steady-state angular velocity constraint parameter. Indicates yaw angle deviation; Represents the steady-state region along the z-axis. , Indicates the initial pitch angle deviation. This represents the standard constraint parameter for z-axis overshoot. This represents the z-axis steady-state error constraint parameter; Speed ​​Channel Rewards It is expressed as follows: in, This represents the non-negative steady-state reward parameter. This indicates the deviation between the speed command and the feedback. Executor reward It is expressed as follows: in, , , , , , and These represent the non-negative penalty parameters used for performance optimization of different performance metrics; express Time of the first One rotor motor throttle, express Constantly maintain the throttle position of the fixed-wing engine. express Constantly adjust the pitch channel rudder deflection angle of the fixed-wing aircraft. express Constantly fix the wing roll channel rudder deflection angle. express Constantly fix the yaw channel rudder deflection angle of the wing. This indicates the time interval for the controller to process data.

9. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 5, characterized in that, The size and structure of the action network and evaluation network are as follows: Action network: The input layer dimension is the same as the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output layer dimension is the same as the action space dimension. Network evaluation: The input layer dimension is consistent with the state space dimension; there are 2 to 5 fully connected hidden layers, each with 128 to 512 neurons; the output is a scalar value.

10. The reinforcement learning-based transition process control method for compound-wing unmanned aerial vehicles according to claim 1, characterized in that, Also includes: Based on classic controller design methods, we independently designed the rotor mode controller and fixed-wing mode controller for a compound wing UAV.