Flight control method, flight control system, and storage medium
Patent Information
- Application Number
- PCT/CN2026/080427
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-10
- Filing Date
- 2026-02-27
- Publication Date
- 2026-09-17
Smart Images

Figure CN2026080427_17092026_PF_FP_ABST
Abstract
Description
Flight control methods, systems and storage media
[0001] This application claims priority to Chinese patent application No. 2025102807437, filed on March 10, 2025, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This specification relates to the field of flight control technology, and in particular to a flight control method, system and storage medium. Background Technology
[0003] Compound wing flight platforms (such as compound wing flying cars) are a new type of aircraft that combines the characteristics of fixed wings and rotors, possessing vertical takeoff and landing and level flight cruise capabilities. Due to their complex airfoils and diverse flight modes (requiring different flight modes at different stages of flight), traditional flight control methods cannot be applied to compound wing flight platforms.
[0004] Among the existing solutions, the flight parameters of the compound wing flight platform can be controlled by sliding mode control, thereby enabling flight control of the compound wing flight platform in different flight modes.
[0005] However, in the existing sliding mode control process, the sliding surface parameters are preset based on the developer's experience, which requires a high level of experience from the developer. Moreover, the preset sliding surface parameters cannot be adjusted according to the actual flight mode, which may affect the stability and maneuverability of the compound wing flight platform.
[0006] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the date of this application, nor does it imply that it can be considered prior art in this disclosure. Summary of the Invention
[0007] This manual provides a flight control method, system, and storage medium that can match corresponding flight control parameters in different flight modes to ensure the stability and maneuverability of the flight platform.
[0008] Firstly, this specification provides a flight control method applied to a flight control system of a flight platform. The method includes: acquiring the flight state of the flight platform, the flight state including flight state parameters and a target flight mode; obtaining target flight control parameters of the flight platform based on the flight state of the flight platform through a pre-trained reinforcement learning agent, the target flight control parameters of the flight platform including at least sliding mode control parameters, the sliding mode control parameters being used to characterize the target flight state parameters corresponding to the flight platform being in the target flight mode, the reinforcement learning agent being trained through reinforcement learning with the flight state of the flight platform as state input and the control parameters of the flight platform as action output, with the reward value of a preset reward function meeting preset conditions as the objective; and controlling the flight platform according to the target flight control parameters, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters is less than or equal to a first threshold.
[0009] In some embodiments, the target flight control parameters further include a control gain parameter corresponding to the target flight mode; controlling the flight platform according to the target flight control parameters includes: obtaining a target sliding surface and a target control law of the flight platform according to the sliding mode control parameters, wherein the target sliding surface is used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode, and the target control law is used to characterize the control logic when controlling the flight platform; and controlling the flight platform based on the target control law and the control gain parameter, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to a first threshold, wherein the control gain parameter is used to characterize the target response rate of the flight platform when responding to the target control law in the target flight mode, and the control gain parameter is positively correlated with the target response rate.
[0010] In some embodiments, controlling the flight platform based on the target control law and the control gain parameter, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold, includes: determining a second threshold corresponding to the target sliding surface; and when it is detected that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the second threshold, continuously controlling the flight platform based on the target control law and the control gain parameter, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
[0011] In some embodiments, the step of continuously controlling the flight platform based on the target control law and the control gain parameter when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is detected to be less than or equal to a second threshold includes: gradually increasing the control frequency of the target control law until the flight platform is continuously controlled when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is detected to be less than or equal to a second threshold.
[0012] In some embodiments, the target control law includes a first control law and a second control law, wherein the control function corresponding to the second control law is a continuous function; when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is detected to be greater than a second threshold, the flight platform is controlled based on the first control law and the control gain parameter;
[0013] The step of continuously controlling the flight platform based on the target control law and the control gain parameter when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the second threshold includes: continuously controlling the flight platform based on the second control law and the control gain parameter when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the second threshold.
[0014] The target control law includes a first control law and a second control law, wherein the control function corresponding to the second control law is a continuous function; when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is detected to be less than or equal to a second threshold, the flight platform is continuously controlled based on the target control law and the control gain parameter, including: controlling the flight platform based on the first control law and the control gain parameter; and when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is detected to be less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0015] In some embodiments, determining the second threshold corresponding to the target sliding surface includes: determining the second threshold based on the target flight mode corresponding to the target sliding surface.
[0016] In some embodiments, determining the second threshold based on the target flight mode corresponding to the target sliding surface includes: obtaining the target response rate corresponding to the target flight mode, determining the second threshold based on the target response rate, wherein the second threshold is negatively correlated with the target response rate; or the target flight control parameters include the second threshold, wherein the second threshold is determined by the reinforcement learning agent based on the flight state parameters and the target flight mode.
[0017] In some embodiments, the flight mode of the flight platform includes at least a takeoff mode, a cruise mode, a transition flight mode, and a landing mode; obtaining the target flight mode of the flight platform includes: determining the target flight mode of the flight platform based on the flight route corresponding to the flight platform, the geographical location information of the flight platform, and the flight altitude of the flight platform, wherein the target flight mode is one of the flight modes of the flight platform.
[0018] In some embodiments, the flight state parameters include at least: attitude parameters, acceleration parameters, velocity parameters, and energy consumption parameters; the preset reward function includes at least: attitude error reward function, acceleration error reward function, velocity error reward function, and energy consumption reward function, wherein the reward values of the attitude error reward function, the acceleration error reward function, and the velocity error reward function are negatively correlated with the error, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
[0019] In some embodiments, the reward value of the preset reward function meets preset conditions, including: maximizing the reward value of the preset reward function.
[0020] Secondly, this specification provides a flight control system, comprising: at least one storage medium storing at least one instruction set for flight control; and at least one processor communicatively connected to the at least one storage medium, wherein, when the flight platform is running, the at least one processor reads the at least one instruction set and executes the flight control method provided in the first aspect according to the instructions of the at least one instruction set.
[0021] Thirdly, this specification also provides a computer-readable non-volatile storage medium, wherein the computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the flight control method provided in the first aspect.
[0022] Other functions of the flight control methods, systems, and storage media provided in this specification will be partially listed in the following description. The inventive aspects of the flight control methods, systems, and storage media provided in this specification can be fully understood through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 illustrates a schematic diagram of an application scenario of a flight control method provided according to an embodiment of this specification;
[0025] Figure 2 shows a schematic diagram of the structure of a flight platform provided according to an embodiment of this specification;
[0026] Figure 3 shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0027] Figure 4 shows a flowchart of a flight control method provided according to an embodiment of this specification; and
[0028] Figure 5 shows a schematic diagram of a reinforcement learning agent training according to an embodiment of this specification. Detailed Implementation
[0029] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.
[0030] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.
[0031] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "multiple" means two or more, unless otherwise explicitly specified. Additionally, the use of "and / or" or "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution that simultaneously satisfies A and B. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of a person skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this disclosure.
[0032] If the embodiments of this disclosure involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0033] In view of the following description, these and other features of this disclosure, as well as the operation and function of the related elements of the structure, and the economy of assembly and manufacture of the components, can be significantly improved. All of these form part of this disclosure with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this disclosure. It should also be understood that the drawings are not drawn to scale.
[0034] The flowcharts used in this disclosure illustrate operations implemented according to some embodiments of this disclosure. It should be clearly understood that the operations in the flowcharts may not be implemented sequentially. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0035] It should also be further understood that the term "and / or" as used in this disclosure and claims refers to any combination of one or more of the associated listed items, as well as all possible combinations, and includes such combinations. In this disclosure, "X includes at least one of A, B, or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only any one of A, B, and C, or any combination of A, B, and C, as well as other possible content / elements. The arbitrary combination of A, B, and C can be A, B, C, AB, AC, BC, or ABC.
[0036] In this disclosure, unless explicitly stated otherwise, the relationships between structures can be direct or indirect, complete or partial. For example, when describing "A is connected to B," unless explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is above B," unless explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). Furthermore, when describing "A is inside B," unless explicitly stated that A is entirely inside B, it should be understood that A can be entirely inside B or partially inside B. And so on.
[0037] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0038] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0039] For ease of description, the terms that will appear later in this manual will be explained first.
[0040] Term 1: Sliding Mode Control. Sliding Mode Control (SMC) is a nonlinear control strategy. The core idea of sliding mode control is to use a control law to make the state trajectory of the controlled object reach and slide on a sliding surface.
[0041] Term 2: Sliding surface. A sliding surface is a specific hyperplane in the state space of a controlled object. Once the dynamic behavior of the controlled object reaches this hyperplane, it will slide along it to reach the desired state or output.
[0042] Term 3: Control Law. A control law is a set of rules or equations that define how to calculate the control input based on the current state or output of the controlled object and the desired target state or output. In other words, a control law is the logical or mathematical expression used by a sliding mode controller to determine its control actions.
[0043] Term 4: Sliding Mode Controller. A sliding mode controller is a virtual component used to implement sliding mode control. It includes a sliding surface and a control law. The input parameters of the sliding mode controller are the target flight control parameters, and the output parameters are the control commands of the flight platform's power system. In this specification, when describing a flight control system implementing a sliding mode control method, it means that the flight control system implements the sliding mode control method through a sliding mode controller.
[0044] Term 5: Buffeting. Buffeting refers to the high-frequency vibration caused by the high-frequency switching control of the flight platform when the flight state parameters reach the target sliding mode surface in the sliding mode control method.
[0045] The following section introduces the application scenarios of this manual.
[0046] The technical solutions provided in this specification are applicable to scenarios involving sliding mode control of compound wing flight platforms. In this scenario, existing technical solutions rely on pre-configured sliding surfaces based on experience when performing sliding mode control on compound wing flight platforms. The flight control effect of these pre-configured sliding surfaces depends on the developer's experience; when the developer lacks experience, the pre-configured sliding surfaces cannot provide satisfactory flight control. Furthermore, the pre-configured sliding surfaces cannot be dynamically adjusted according to the flight mode of the compound wing flight platform, which may affect the stability and maneuverability of the compound wing flight platform.
[0047] This specification provides a flight control method applicable to sliding mode control of a compound wing flight platform. In this scenario, the flight control system, based on the flight platform's flight state parameters and the target flight mode, obtains the target flight control parameters (which may include sliding mode control parameters, such as sliding surface parameters and control laws) through a pre-trained reinforcement learning agent. Then, the flight platform is controlled according to these target flight control parameters, causing the platform's flight state parameters to approximate the target flight state parameters indicated by the target flight control parameters. The control system can obtain the target flight control parameters of the flight platform through the reinforcement learning agent based on the flight state parameters and the target flight mode, eliminating the need for manual setting of the flight control parameters. Moreover, the target flight control parameters obtained through the reinforcement learning agent represent a control strategy that performs well under the current flight state and mode. These target flight control parameters can be dynamically updated as the flight state or mode changes, ensuring the stability and maneuverability of the flight platform.
[0048] Figure 1 illustrates an application scenario of a flight control method provided according to an embodiment of this specification. As shown in Figure 1, scenario 100 includes a flight platform 11. In scenario 100, the flight platform 11 is shown in the form of a compound-wing flying car. It should be noted that, in this specification, the flight platform 11 can be any type of aircraft; for example, the flight platform 11 can also be a manned drone, a cargo drone, or a drone for a specific industry.
[0049] Figure 2 shows a schematic diagram of the structure of a flight platform provided according to an embodiment of this specification.
[0050] In some embodiments, referring to FIG2, the flight platform 11 may include a flight control system, a power system, and a power battery. The flight control system may include a sensor array, a reinforcement learning agent, and a sliding mode controller. The sensor array includes multiple sensors for collecting flight state parameters of the flight platform 11. The flight control system can determine the target flight mode of the flight platform 11 based on the flight state parameters, thereby obtaining the flight state of the flight platform. The reinforcement learning agent is used to determine target flight control parameters based on the flight state, and the sliding mode controller is used to output control commands for the power system based on the target flight control parameters. The power system can be powered by the power battery and executes the control commands output by the sliding mode controller.
[0051] In some embodiments, when the flight platform 11 is a compound-wing flying car, its power system includes a power system (such as a turbofan engine) required for fixed-wing flight and a power system (such as a motor equipped with a propeller) required for multi-rotor flight. Alternatively, in some embodiments, a single power system (such as a motor equipped with a propeller) can be shared for both fixed-wing and multi-rotor flight. The power system can adjust its angle according to the flight mode to adapt to the flight. In this case, when the power system executes the control commands output by the sliding mode controller, it also includes the control of the power system angle.
[0052] In some embodiments, the flight platform 11 may also be configured with different carrying systems depending on the type of aircraft. For example, when the flight platform 11 is a compound wing flying car, the carrying system may be a passenger cabin; when the flight platform 11 is a cargo drone, the carrying system may be a cargo hold.
[0053] Referring to Figure 1, the flight process of flight platform 11 from takeoff at address A to landing at address B may include a takeoff phase, a takeoff transition phase, a cruise phase, a landing transition phase, and a landing phase, with each phase corresponding to a flight mode.
[0054] Throughout the flight, the flight platform 11 can perform flight control at each stage based on its built-in flight control system and the flight control methods provided in this manual. The flight control methods provided in this manual can determine the corresponding target flight control parameters at different stages based on the flight status of the flight platform 11 (including flight status parameters and target flight mode), and perform flight control based on the target flight control parameters.
[0055] In some embodiments, the flight control method provided herein can be executed by the flight control system of the flight platform 11. In this case, the flight control system of the flight platform 11 may store data or instructions for executing the flight control method described herein, and may execute or be used to execute said data or instructions. In some embodiments, the flight control system of the flight platform 11 may include hardware devices with data processing capabilities and the necessary programs required to drive the hardware devices.
[0056] The flight control system of flight platform 11 can correspond to a single computing device or a computing cluster composed of multiple computing devices. For example, the flight control system can be deployed in the computing device on flight platform 11; or, the flight control system can be distributed across multiple computing devices on flight platform 11; or, part of it can be deployed in the computing device on flight platform 11, and another part can be deployed on a cloud server connected to the computing device network.
[0057] Figure 3 shows a hardware structure diagram of a computing system provided according to an embodiment of this specification. The computing system can serve as the flight control system of the flight platform 11 in Figure 1, executing the flight control methods described in this specification.
[0058] As shown in Figure 3, the computing system 200 may include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 may also include a communication port 250 and an internal communication bus 210. The computing system 200 may also include I / O components 260.
[0059] The internal communication bus 210 can connect to different system components. For example, the internal communication bus 210 can connect to storage medium 230, processor 220, communication port 250, and I / O component 260, etc.
[0060] I / O component 260 supports input / output between computing system 200 and other components.
[0061] Communication port 250 is used for data communication between computing system 200 and the outside world. For example, communication port 250 can be used for data communication between computing system 200 and a network. Communication port 250 can be a wired communication port or a wireless communication port.
[0062] Storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. Storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc.
[0063] At least one processor 220 may be communicatively connected to at least one storage medium 230. When the computing system 200 is running, at least one processor 220 reads the at least one instruction set and executes the flight control method provided in this specification according to the instructions of the at least one instruction set. The processor 220 may perform the steps included in the flight control method. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, microprocessor, reduced instruction set computer (RISC), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof.
[0064] For illustrative purposes only, the accompanying drawings show only one processor 220 for the computing system 200. However, it should be noted that the computing system 200 may also include multiple processors; therefore, the operations and / or method steps disclosed herein may be executed by one processor or by multiple processors in combination. For example, if the processor 220 of the computing system 200 described in this specification executes steps A and B, it should be understood that steps A and B may also be executed jointly or separately by two different processors 220 (e.g., a first processor executes step A, a second processor executes step B, or the first and second processors jointly execute steps A and B).
[0065] Figure 4 shows a flowchart of a flight control method provided according to an embodiment of this specification. As previously described, the flight control system deployed in flight platform 11 can execute the flight control method of this specification.
[0066] As shown in Figure 4, flight control methods may include:
[0067] S410: Acquire the flight status of the flight platform, including flight status parameters and target flight mode.
[0068] In some embodiments, the flight state parameters include at least: attitude parameters, acceleration parameters, velocity parameters, and energy consumption parameters.
[0069] Optionally, the flight control system can obtain attitude and acceleration parameters by reading the measurement results from sensors installed on the flight platform. For example, the flight control system can read the measurement results from attitude angle sensors installed on the flight platform to obtain attitude parameters such as pitch angle, roll angle, and yaw angle. The flight control system can also read the parameters from accelerometers installed on the flight platform to obtain acceleration parameters.
[0070] Optionally, the flight control system can obtain energy consumption parameters such as the remaining charge and discharge power of the power battery through the battery status sensor installed on the power battery in the flight platform.
[0071] Optionally, the flight control system can obtain speed parameters based on the measurement results of speed sensors installed on the flight platform. Preferably, the speed sensor can be a pitot tube. At least one pitot tube can be installed on the flight platform, and the flight control system can obtain the flight speed of the flight platform based on the measurement results of the pitot tube. Then, the flight control system can calculate speed parameters such as the longitudinal speed, lateral speed, and vertical speed of the flight platform based on the attitude parameters and the flight speed.
[0072] Optionally, the flight mode of the flight platform includes at least a takeoff mode, a cruise mode, a transition flight mode, and a landing mode. During the flight of the flight platform, different phases can each correspond to a flight mode. Preferably, referring to Figure 1, the takeoff phase of the flight platform can correspond to a takeoff mode, the takeoff transition phase and the landing transition phase of the flight platform can correspond to a transition mode, the cruise phase of the flight platform can correspond to a cruise mode, and the landing phase of the flight platform can correspond to a landing mode.
[0073] In takeoff mode, the flight platform needs to provide a large thrust in the vertical direction to support its vertical ascent. The flight control system can adjust the attitude, speed and acceleration of the flying car to quickly respond to the instantaneous needs of the flying car and avoid unstable dynamic behavior of the flying car during takeoff.
[0074] Optionally, the flight control system acquires the target flight mode of the flight platform, including: determining the target flight mode of the flight platform based on the flight route corresponding to the flight platform, the geographical location information of the flight platform, and the flight altitude of the flight platform, wherein the target flight mode is one of the flight modes of the flight platform.
[0075] Before takeoff, the flight platform needs to set the flight path. Preferably, the flight path can include the start point, end point, cruising altitude, and multiple navigation points along the way.
[0076] Optionally, when the flight platform flies along the route, its geographical location information can be compared with the geographical location information of the start point, end point, and navigation points along the route, as well as the flight altitude of the flight platform with the cruising altitude along the route, to determine the target flight mode of the flight platform. Preferably, the target flight mode is one of the following: takeoff mode, transition mode, cruise mode, and landing mode.
[0077] Preferably, if the geographical location of the flight platform obtained by the flight control system is the same as or similar to the geographical location of the starting point, and the flight altitude of the flight platform is less than or equal to the preset takeoff altitude, then it can be determined that the flight platform is in the takeoff phase, that is, the target flight mode is the takeoff mode.
[0078] Preferably, if the flight control system obtains that the geographical location of the flight platform is between the starting point and the first waypoint, and the flight altitude of the flight platform is greater than the preset takeoff altitude but less than the cruise altitude, then it can be determined that the flight platform is in the takeoff transition phase, that is, the target flight mode is the transition mode.
[0079] Preferably, if the geographical location of the flight platform obtained by the flight control system is between the first waypoint and the last waypoint, and the flight altitude of the flight platform is the same as or similar to the cruise altitude, then it can be determined that the flight platform is in the cruise phase, that is, the target flight mode is the cruise mode.
[0080] Preferably, if the geographical location of the flight platform obtained by the flight control system is between the last waypoint and the destination, and the flight altitude of the flight platform is greater than the preset takeoff altitude but less than the cruise altitude, then it can be determined that the flight platform is in the landing transition phase, that is, the target flight mode is the transition mode.
[0081] Preferably, if the geographical location of the flight platform obtained by the flight control system is the same as or similar to the geographical location of the destination, and the flight altitude of the flight platform is less than or equal to the preset takeoff altitude, then it can be determined that the flight platform is in the landing phase, that is, the target flight mode is the landing mode.
[0082] It should be noted that two geographical locations being the same or similar can mean that the coordinates of the two geographical locations are the same, or that the distance between the coordinates of the two geographical locations is less than a preset distance threshold.
[0083] S420: Based on the flight state of the flight platform, the target flight control parameters of the flight platform are obtained through a pre-trained reinforcement learning agent. The target flight control parameters of the flight platform include at least sliding mode control parameters. The sliding mode control parameters are used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode. The reinforcement learning agent is trained by taking the flight state of the flight platform as the state input and the control parameters of the flight platform as the action output, with the goal of the reward value of the preset reward function meeting the preset conditions.
[0084] Optionally, when creating a reinforcement learning agent, it is necessary to first construct the agent's state space, action space, and reward function. The state space serves as the input to the reinforcement learning agent. Preferably, in this specification, the state space may include all types of parameters in the flight state. The action space serves as the output of the reinforcement learning agent. Preferably, in this specification, the action space may include target flight control parameters.
[0085] After the flight control system inputs the acquired flight state into the reinforcement learning agent, the agent can input the target flight control parameters corresponding to the flight state. The flight control system then controls the flight platform based on these parameters. The reward function characterizes the control effect of the target flight control parameters on the flight platform; a higher reward value indicates a better control effect.
[0086] Optionally, the flight control system can train the constructed reinforcement learning agent. That is, the flight control system can perform multiple iterations on the reinforcement learning agent, with the goal of each iteration being to ensure that the reward value of a preset reward function meets preset conditions. As an example, the preset conditions could be maximizing the reward value of the preset reward function.
[0087] In other words, after multiple iterations, the reward values obtained in two consecutive iterations are the same or similar. The two iterations being the same or similar can mean that the two reward values are identical, or that the difference between the two reward values is less than a preset difference threshold.
[0088] Figure 5 shows a schematic diagram of a reinforcement learning agent training according to an embodiment of this specification.
[0089] Referring to Figure 5, in the t-th iteration, the flight control system can adjust the configuration parameters of the reinforcement learning agent based on the reward value of the (t-1)-th iteration, and input the flight state obtained in the (t-1)-th iteration into the reinforcement learning agent to generate the target flight control parameters for the t-th iteration. Next, based on the target flight control parameters for the t-th iteration, the flight control system generates the flight state for the t-th iteration through interaction with the environment, and obtains the reward value for the t-th iteration based on the flight state and the preset reward function. Finally, the flight state and the reward value for the t-th iteration are used for the (t+1)-th iteration.
[0090] The environment shown in Figure 5 can simulate the flight state of the flight platform and the environmental state of the target space where the flight platform is located.
[0091] Optionally, the flight control system can generate the flight state for the t-th iteration by interacting with the environment based on the target flight control parameters for the t-th iteration. The flight control system can first input the target flight control parameters for the t-th iteration into the environment, and then use the environment to simulate the flight state of the flight platform flying in the target space according to the target flight control parameters, thereby generating the flight state for the t-th iteration.
[0092] Optionally, the flight state parameters include at least: attitude parameters, acceleration parameters, velocity parameters, and energy consumption parameters. The preset reward function can then include at least: an attitude error reward function, an acceleration error reward function, a velocity error reward function, and an energy consumption reward function. The reward values of the attitude error reward function, acceleration error reward function, and velocity error reward function are negatively correlated with the error, while the reward value of the energy consumption reward function is negatively correlated with energy consumption.
[0093] Preferably, the reward value obtained by the flight control system according to the preset reward function can be the sum of the reward value corresponding to the attitude error reward function, the reward value corresponding to the acceleration error reward function, the reward value corresponding to the velocity error reward function, and the reward value corresponding to the energy consumption reward function.
[0094] Optionally, the flight control system controls the flight platform based on sliding mode control. In this case, the target flight control parameters may include sliding mode control parameters, control gain parameters corresponding to the target flight mode, and the rate of change of the control input. The sliding mode control parameters may include parameters of the target sliding surface and parameters of the target control law. The control gain parameter is a weighting coefficient in the target control law, used to characterize the target response rate of the flight platform in the target flight mode when responding to the target control law. The control gain parameter is positively correlated with the target response rate; that is, the larger the control gain parameter, the faster the flight platform responds to the target control law. The rate of change of the control input is used to characterize the rate of change of the control input (such as controlling the steer or throttle).
[0095] Optionally, when applying a pre-trained reinforcement learning agent to the flight control system of a flight platform, the method of application is similar to the iterative method used during training. The difference lies in that, when applying a pre-trained reinforcement learning agent to the flight control system of a flight platform, the flight control system is based on the flight status acquired by sensors installed on the flight platform.
[0096] In this embodiment, the flight control system trains a reinforcement learning agent through a reinforcement learning scheme. Based on the flight state of the flight platform, the system obtains the target flight control parameters of the flight platform through the pre-trained reinforcement learning agent, thereby achieving adaptive adjustment of the target flight control parameters and optimizing the robustness of the flight platform under different flight modes.
[0097] S430: Control the flight platform according to the target flight control parameters, so that the absolute value of the difference between the flight platform's flight state parameters and the target flight state parameters is less than or equal to a first threshold.
[0098] In some embodiments, the flight control system can obtain the target sliding surface and target control law of the flight platform based on sliding mode control parameters. The target sliding surface is used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode, and the target control law is used to characterize the control logic when controlling the flight platform. Furthermore, the flight control system controls the flight platform based on the target control law and control gain parameters, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to a first threshold.
[0099] Optionally, the sliding mode controller in the flight control system can determine the target sliding surface and target control law of the flight platform based on the sliding mode control parameters. The target sliding surface can be a linear or nonlinear combination of target flight state parameters. For example, the target sliding surface can include a multi-dimensional function composed of control errors (such as attitude and velocity errors) and their derivatives (such as angular velocity and acceleration errors). The target control law can be a function based on the flight state parameters, used to make the flight state parameters approach the target sliding surface from any initial value within a finite time. Preferably, the target control law can include a sign function based on the target sliding surface and a linear function of the derivative of the target sliding surface, wherein the linear coefficients of the linear function are adjustment gain parameters.
[0100] Optionally, when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold, it indicates that the flight state parameters of the flight platform are approaching the target sliding surface. The first threshold is determined by the reinforcement learning agent based on the flight state.
[0101] In some embodiments, chattering occurs when the flight state parameters of the flight platform approach the target sliding surface. Here, the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is defined as the sliding variable, that is, the closer the flight state parameters are to the target sliding surface, the closer the absolute value of the sliding variable is to the first threshold.
[0102] To mitigate chattering, the flight control system can introduce a boundary layer near the target sliding surface. When the sliding variable corresponding to the flight platform is outside the boundary layer, the target control law switches the control of the flight platform. When the sliding variable corresponding to the flight platform enters the boundary layer, the flight platform is continuously controlled based on the target control law.
[0103] Optionally, to achieve the above scheme, the flight control system can first determine a second threshold corresponding to the target sliding surface. Then, when the flight control system detects that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the second threshold, it continuously controls the flight platform based on the target control law and control gain parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
[0104] The second threshold represents the thickness of the boundary layer. When the flight control system detects that the absolute value of the sliding mode variable is less than or equal to the second threshold, it determines that the flight platform's flight state parameters have entered the boundary layer. In this case, the flight control system can continuously control the flight platform based on the target control law and control gain parameters, so that the flight platform's flight state parameters smoothly approach the target sliding mode surface, significantly reducing chattering.
[0105] Optionally, the flight control system can determine a second threshold based on the target flight mode corresponding to the target sliding surface. Preferably, the flight control system can acquire the target response rate corresponding to the target flight mode and determine the second threshold based on the target response rate, wherein the second threshold is negatively correlated with the target response rate.
[0106] Specifically, the thicker the boundary layer (i.e., the larger the second threshold), the longer it takes for the flight control system to smoothly approach the target sliding mode surface through the target control law. In other words, the thickness of the boundary layer affects the response rate of the flight platform to sliding mode control. Different flight modes also have different requirements for response rate.
[0107] Optionally, the flight control system can configure a corresponding second threshold for different flight modes. For example, when a flight mode requires a faster response rate, the flight control system can configure a smaller second threshold for that flight mode. When the flight control system determines the target flight mode corresponding to the flight platform, it can determine that the second threshold corresponding to the target flight mode is the same as the second threshold corresponding to the target sliding surface.
[0108] Optionally, the second threshold corresponding to the target sliding surface can also be directly determined by the reinforcement learning agent. Preferably, when iteratively training the reinforcement learning agent, the second threshold corresponding to the target sliding surface can be used as a parameter in the action space, that is, the target flight control parameters output by the reinforcement learning agent also include the second threshold.
[0109] The flight control system can output target flight control parameters, including a second threshold, based on flight state parameters and target flight mode, through a reinforcement learning agent, thereby determining the second threshold corresponding to the target sliding surface.
[0110] In this embodiment, the flight control system can dynamically adjust the thickness of the boundary layer according to the target flight mode corresponding to the flight platform, so that when the flight platform is subjected to sliding mode control, the chattering phenomenon can be significantly reduced, the response rate of the flight platform to sliding mode control can be guaranteed, and the stability of the flight platform can be improved.
[0111] Optionally, when the flight control system detects that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the second threshold, the control frequency of the target control law can be gradually increased until the flight platform is continuously controlled.
[0112] As the flight control system gradually increases the control frequency of the target control law, the sliding mode controller will respond to changes in flight state parameters more quickly. Furthermore, with the increased control frequency, the sliding mode controller can adjust the intervals between output control commands to be shorter, resulting in finer granularity and higher control precision. When the control frequency is sufficiently high, the control of the flight platform by the sliding mode controller can be considered continuous control.
[0113] Optionally, the target control law includes a first control law and a second control law, where the control function corresponding to the second control law is a continuous function. When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is detected to be greater than a second threshold, the flight platform is controlled based on the first control law and the control gain parameter. When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is detected to be less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0114] Flight control systems can use different control laws outside and inside the boundary layer. For example, a flight control system can use a first control law (e.g., a switching control law) when the sliding mode variable is outside the boundary layer, and a second control law (e.g., a continuous control law) when the sliding mode variable is inside the boundary layer.
[0115] Preferably, the target control law u(t) can be expressed by the following formula:
[0116] (First Control Law)
[0117] u(t) = -K·sat(φS(x)) (Second control law)
[0118] Here, `sat()` is a saturation function, `K` is the control gain parameter, `φ` represents the second threshold, and `S(x)` is the sliding mode variable. In this specification, `S(x)` represents the sliding mode variable. When `|S(x)|` is greater than `φ` (i.e., the sliding mode variable is outside the boundary layer), the flight control system controls the flight platform based on the first control law and the control gain parameter. When `|S(x)|` is less than or equal to `φ` (i.e., the sliding mode variable is inside the boundary layer), the flight platform is continuously controlled based on the second control law and the control gain parameter. In this case, the control input will saturate, thereby avoiding the occurrence of chattering.
[0119] In summary, this specification provides a flight control method system and storage medium. The flight control system can obtain target flight control parameters (which may include sliding mode control parameters, such as sliding surface parameters and control laws) of the flight platform based on its flight state parameters and target flight mode through a pre-trained reinforcement learning agent. Then, the flight platform is controlled according to the target flight control parameters, so that the flight state parameters of the flight platform approach the target flight state parameters indicated by the target flight control parameters. The control system can obtain the target flight control parameters of the flight platform through the reinforcement learning agent based on the flight state parameters and target flight mode, without requiring manual setting of the flight control parameters. Moreover, the target flight control parameters obtained through the reinforcement learning agent represent a control strategy that performs well under the current flight state and flight mode. The target flight control parameters can be dynamically updated as the flight state or flight mode changes, ensuring the stability and maneuverability of the flight platform.
[0120] This specification, in another aspect, provides a computer-readable non-transitory storage medium storing at least one instruction set for flight control. When the at least one instruction set is executed by a processor, it instructs the processor to perform the steps of the flight control method described herein. In some possible embodiments, various aspects of this specification may also be implemented as a program product comprising program code. When the program product is run on a computing system 200, the program code causes the computing system 200 to perform the steps of the flight control method described herein. The program product for implementing the above method may employ a portable compact disk read-only memory (CD-ROM) containing program code and may run on the computing system 200. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on computing system 200, partially on computing system 200, as a standalone software package, partially on computing system 200 and partially on a remote computing device, or entirely on a remote computing device.
[0121] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0122] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.
[0123] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.
[0124] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.
[0125] Every patent, patent application, publication of a patent application, and other material such as articles, books, specifications, publications, documents, articles, etc., cited herein, except for those inconsistent with or conflicting with this document, or those having a restrictive effect on the widest scope of the claims, may be incorporated herein by reference for all purposes now or hereafter associated with this document. Furthermore, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0126] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.
Claims
1. A flight control method applied to the flight control system of a flight platform, the method comprising: The flight status of the flight platform is obtained, including flight status parameters and target flight mode; Based on the flight state of the flight platform, target flight control parameters of the flight platform are obtained through a pre-trained reinforcement learning agent. These target flight control parameters include at least sliding mode control parameters, which characterize the target flight state parameters corresponding to when the flight platform is in the target flight mode. The reinforcement learning agent is trained using the flight state of the flight platform as the state input and the control parameters of the flight platform as the action output, with the goal of achieving a reward value that meets preset conditions. The flight platform is controlled according to the target flight control parameters, such that the absolute value of the difference between the flight platform's flight state parameters and the target flight state parameters is less than or equal to a first threshold.
2. The method according to claim 1, wherein, The target flight control parameters also include the control gain parameters corresponding to the target flight mode; The step of controlling the flight platform according to the target flight control parameters includes: The target sliding surface and target control law of the flight platform are obtained based on the sliding mode control parameters. The target sliding surface is used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode, and the target control law is used to characterize the control logic when controlling the flight platform. The flight platform is controlled based on the target control law and the control gain parameter, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold. The control gain parameter is used to characterize the target response rate of the flight platform in the target flight mode when responding to the target control law. The control gain parameter is positively correlated with the target response rate.
3. The method according to any one of claims 1-2, wherein, The control of the flight platform based on the target control law and the control gain parameter, such that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold, includes: Determine the second threshold corresponding to the target sliding surface; and When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a second threshold, the flight platform is continuously controlled based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
4. The method according to any one of claims 1-3, wherein, When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a second threshold, continuous control of the flight platform is performed based on the target control law and the control gain parameter, including: When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the second threshold, the control frequency of the target control law is gradually increased until the flight platform is continuously controlled.
5. The method according to any one of claims 1-4, wherein, The target control law includes a first control law and a second control law, and the control function corresponding to the second control law is a continuous function. When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is greater than the second threshold, the flight platform is controlled based on the first control law and the control gain parameter. When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a second threshold, the flight platform is continuously controlled based on the target control law continuity and the control gain parameter, including: When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
6. The method according to any one of claims 1-5, wherein, Determining the second threshold corresponding to the target sliding surface includes: The second threshold is determined based on the target flight mode corresponding to the target sliding surface.
7. The method according to any one of claims 1-6, wherein, Determining the second threshold based on the target flight mode corresponding to the target sliding surface includes: Obtain the target response rate corresponding to the target flight mode, and determine the second threshold based on the target response rate, wherein the second threshold is negatively correlated with the target response rate; or The target flight control parameters include the second threshold, which is determined by the reinforcement learning agent based on the flight state parameters and the target flight mode.
8. The method according to any one of claims 1-7, wherein, The flight modes of the flight platform include at least takeoff mode, cruise mode, transition flight mode and landing mode; The acquisition of the target flight mode of the flight platform includes: Based on the flight path corresponding to the flight platform, the geographical location information of the flight platform, and the flight altitude of the flight platform, the target flight mode of the flight platform is determined, and the target flight mode is one of the flight modes of the flight platform.
9. The method according to any one of claims 1-8, wherein, The flight state parameters include at least: attitude parameters, acceleration parameters, velocity parameters, and energy consumption parameters; The preset reward function includes at least: an attitude error reward function, an acceleration error reward function, a velocity error reward function, and an energy consumption reward function, wherein the reward values of the attitude error reward function, the acceleration error reward function, and the velocity error reward function are negatively correlated with the error, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
10. The method according to any one of claims 1-9, wherein, The reward value of the preset reward function meets preset conditions, including: The preset reward function maximizes the reward value.
11. A flight control system, comprising: At least one storage medium storing at least one instruction set for flight control; as well as At least one processor is communicatively connected to the at least one storage medium. When the flight control system is running, the at least one processor reads the at least one instruction set. And execute the flight control method according to any one of claims 1-10 according to the instructions of the at least one set of instructions.
12. A computer-readable non-volatile storage medium, wherein, The computer-readable non-volatile storage medium stores at least one instruction set, which, when executed by at least one processor, implements the flight control method as described in any one of claims 1-10.