Flight control method and system and storage medium
By using pre-trained reinforcement learning agents to dynamically obtain target flight control parameters on the composite wing flight platform, the problem of difficult to dynamically adjust the sliding mode control parameters is solved, and the stability and handling of the flight platform are improved.
Patent Information
- Application Number
- CN202510280743.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing sliding mode control method is difficult to dynamically adjust the sliding mode surface parameters on the composite wing flight platform, resulting in limited flight stability and handling.
Through pre-trained reinforcement learning agents, the target flight control parameters, including skid mode control parameters, are dynamically obtained according to the flight status and target flight mode of the flight platform, and then adjust the flight control strategy.
It realizes matching appropriate flight control parameters in different flight modes, improves the stability and handling of the composite wing flight platform, and reduces the dependence on developer experience.
Smart Images

Figure CN120143680A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of flight control, and particularly to a flight control method, system, and storage medium. Background Art
[0002] A compound-wing flight platform (such as a compound-wing flying car) is a new type of aircraft that combines the characteristics of fixed wings and rotors and has the capabilities of vertical takeoff and landing and level flight cruising. Due to the complex airfoil and variable flight modes (different flight modes are required in different flight phases) of the compound-wing flight platform, traditional flight control methods cannot be applied to the compound-wing flight platform.
[0003] In the existing solutions, the flight parameters of the compound-wing flight platform can be controlled by means of sliding mode control, and thus the flight control of the compound-wing flight platform in different flight modes can be realized.
[0004] However, in the existing sliding mode control process, the sliding mode surface parameters are preset based on the experience of developers, which requires high experience of developers. Moreover, the preset sliding mode surface parameters cannot be adjusted according to the actual flight mode, which may affect the stability and maneuverability of the compound-wing flight platform.
[0005] The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention
[0006] This specification provides a flight control method, system, and storage medium that can match corresponding flight control parameters in different flight modes to ensure the stability and maneuverability of the flight platform.
[0007] In a first aspect, this specification provides a flight control method, which is applied to the flight control system of a flight platform. The method includes: obtaining the flight state of the flight platform, where the flight state includes flight state parameters and a target flight mode; according to the flight state of the flight platform, obtaining the target flight control parameters of the flight platform through a pre-trained reinforcement learning agent. The target flight control parameters of the flight platform at least include sliding mode control parameters, and the sliding mode control parameters are at least used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode. The reinforcement learning agent takes the flight state of the flight platform as the state input and the control parameters of the flight platform as the action output, and is trained by means of reinforcement learning with the goal that the reward value of a preset reward function meets a preset condition; controlling the flight platform according to the target flight control parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters is less than or equal to a first threshold.
[0008] In some embodiments, the target flight control parameters further include control gain parameters corresponding to the target flight mode; the controlling the flight platform according to the target flight control parameters includes: obtaining a target sliding mode surface and a target control law of the flight platform according to the sliding mode control parameters, where the target sliding mode surface is used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode, and the target control law is used to characterize the control logic when controlling the flight platform; and controlling the flight platform based on the target control law and the control gain parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to a first threshold. The control gain parameters are used to characterize the target response rate of the flight platform when responding to the target control law in the target flight mode, and the control gain parameters are positively correlated with the target response rate.
[0009] In some embodiments, the controlling the flight platform based on the target control law and the control gain parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to a first threshold includes: determining a second threshold corresponding to the target sliding mode surface; and when it is detected that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to the second threshold, continuously controlling the flight platform based on the target control law and the control gain parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to a first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
[0010] In some embodiments, when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to a second threshold, continuous control of the flight platform based on the target control law and the control gain parameter includes: when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to the second threshold, gradually increasing the control frequency of the target control law until continuous control of the flight platform is performed.
[0011] In some embodiments, the target control law includes a first control law and a second control law, and the control function corresponding to the second control law is a continuous function; when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is greater than the second threshold, the flight platform is controlled based on the first control law and the control gain parameter.
[0012] When the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to the second threshold, continuous control of the flight platform based on the continuous target control law and the control gain parameter includes: when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0013] The target control law includes a first control law and a second control law, and the control function corresponding to the second control law is a continuous function; when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to the second threshold, continuous control of the flight platform based on the target control law and the control gain parameter includes: controlling the flight platform based on the first control law and the control gain parameter; when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding surface is less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0014] In some embodiments, determining the second threshold corresponding to the target sliding surface includes: determining the second threshold according to the target flight mode corresponding to the target sliding surface.
[0015] In some embodiments, determining the second threshold according to the target flight mode corresponding to the target sliding mode surface includes: obtaining the target response rate corresponding to the target flight mode, and determining the second threshold according to the target response rate, where the second threshold is negatively correlated with the target response rate; or the target flight control parameter includes the second threshold, and the second threshold is determined by the reinforcement learning agent according to the flight state parameters and the target flight mode.
[0016] In some embodiments, the flight modes of the flight platform at least include a takeoff mode, a cruise mode, a transition flight mode, and a landing mode; obtaining the target flight mode of the flight platform includes: determining the target flight mode of the flight platform according to the corresponding route of the flight platform, the geographical location information of the flight platform, and the flight altitude of the flight platform, where the target flight mode is one of the flight modes of the flight platform.
[0017] In some embodiments, the flight state parameters at least include: attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters; the preset reward function at least includes: an attitude error reward function, an acceleration error reward function, a speed error reward function, and an energy consumption reward function, where the reward values of the attitude error reward function, the acceleration error reward function, and the speed error reward function are negatively correlated with the errors, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
[0018] In some embodiments, the reward value of the preset reward function meets a preset condition, including: maximizing the reward value of the preset reward function.
[0019] In a second aspect, this specification provides a flight control system, including: at least one storage medium storing at least one instruction set for performing flight control; and at least one processor communicatively connected to the at least one storage medium, where when the flight platform is running, the at least one processor reads the at least one instruction set and executes the flight control method provided in the first aspect according to the instructions of the at least one instruction set.
[0020] In a third aspect, this specification further provides a computer-readable non-volatile storage medium, where at least one instruction set is stored in the computer-readable non-volatile storage medium, and when the at least one instruction set is executed by at least one processor, the flight control method provided in the first aspect is implemented.
[0021] Other functions of the flight control method, system, and storage medium provided in this specification will be partially listed in the following description. The creative aspects of the flight control method, system, and storage medium provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the detailed examples below. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 The figure shows a schematic diagram of an application scenario of a flight control method provided according to an embodiment of this specification;
[0024] Figure 2 The figure shows a schematic diagram of the structure of a flight platform provided according to an embodiment of this specification;
[0025] Figure 3 The figure shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0026] Figure 4 The figure shows a flowchart of a flight control method provided according to an embodiment of this specification; and
[0027] Figure 5 The figure shows a schematic diagram of the training of a reinforcement learning agent provided according to an embodiment of this specification. Detailed Implementation Modes
[0028] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the shown embodiments, but has the broadest scope consistent with the claims.
[0029] The terms used herein are for the purpose of describing particular example embodiments only and are not limiting. For example, unless the context clearly dictates otherwise, as used herein, the singular forms "a", "an" and "the" may also include the plural forms. When used in this specification, the terms "comprises", "comprising" and / or "having" mean that the associated integers, steps, operations, elements and / or components are present, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups in the system / method.
[0030] In view of the following description, these features of the present specification and other features, as well as the operations and functions of the related elements of the structure, and the combination and manufacturing economy of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0031] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments of this specification. It should be clearly understood that the operations of the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0032] For ease of description, the terms that will appear later in this specification are first explained.
[0033] Term 1: Sliding mode control method. Sliding Mode Control (SMC) is a non - linear control strategy. The core idea of sliding mode control is to make the state trajectory of the controlled object reach and slide on a sliding mode surface through the control law.
[0034] Term 2: Sliding mode surface. The sliding mode surface is a specific hyperplane in the state space of the controlled object. Once the dynamic behavior of the controlled object reaches this hyperplane, it will slide along it to reach the desired state or output.
[0035] Term 3: Control law. The control law is a set of rules or equations that define how to calculate the control input based on the current state or output of the controlled object and the desired target state or output. That is, the control law is the logical or mathematical expression used by the sliding mode controller to determine its control actions.
[0036] Term 4: Sliding mode controller. A sliding mode controller is a virtual component used to implement sliding mode control, which includes a sliding mode surface and a control law. The input parameter of the sliding mode controller is the target flight control parameter, and the output parameter is the control instruction for the flight platform power system. In this specification, when describing that the flight control system implements the sliding mode control method, it means that the flight control system realizes the sliding mode control method through the sliding mode controller.
[0037] Term 5: Chattering. Chattering refers to the high-frequency vibration caused by the high-frequency switching control of the flight platform when the target control law reaches the target sliding mode surface in the sliding mode control method.
[0038] The application scenarios of this specification are introduced below.
[0039] The technical solution provided in this specification is applicable to the scenario of sliding mode control of a compound-wing flight platform. In this scenario, there is currently a technical solution where when performing sliding mode control on a compound-wing flight platform, the sliding mode surface is pre-configured based on experience. The flight control effect of the pre-set sliding mode surface depends on the experience of the developers. When the developers have insufficient experience, the pre-set sliding mode surface cannot provide a good flight control effect. Moreover, the pre-set sliding mode surface cannot be dynamically adjusted according to the flight mode of the compound-wing flight platform, which may affect the stability and maneuverability of the compound-wing flight platform.
[0040] This specification provides a flight control method that can be applied to the scenario of sliding mode control of a compound-wing flight platform. In this scenario, the flight control system can obtain the target flight control parameters of the flight platform (which may include sliding mode control parameters, such as sliding mode surface parameters, control laws, etc.) through a pre-trained reinforcement learning agent based on the flight state parameters and the target flight mode of the flight platform. Then, the flight platform is controlled according to the target flight control parameters, so that the flight state parameters of the flight platform approach the target flight state parameters indicated by the target flight control parameters. The control system can obtain the target flight control parameters of the flight platform through the reinforcement learning agent according to the flight state parameters and the target flight mode of the flight platform, and the flight control parameters do not need to be set manually. Moreover, the target flight control parameters obtained through the reinforcement learning agent are control strategies with better effects in the current flight state and current flight mode, and the target flight control parameters can be dynamically updated as the flight state or flight mode changes, ensuring the stability and maneuverability of the flight platform.
[0041] Figure 1 Shows a schematic diagram of the application scenario of a flight control method provided according to an embodiment of this specification. As Figure 1As shown, the scenario 100 includes a flight platform 11. In the scenario 100, the flight platform 11 is shown in the form of a compound-wing flying car. It should be noted that in this specification, the flight platform 11 can be any type of aircraft. For example, the flight platform 11 can also be a manned drone, a cargo drone, or a specific-industry drone, etc.
[0042] Figure 2 Shows a schematic structural diagram of a flight platform provided according to an embodiment of this specification.
[0043] In some embodiments, referring to Figure 2 , the flight platform 11 may include a flight control system, a power system, and a power battery. The flight control system may include a sensor group, a reinforcement learning agent, and a sliding mode controller. Among them, the sensor group includes multiple sensors for collecting flight state parameters of the flight platform 11. The flight control system can determine the target flight mode of the flight platform 11 according to the flight state parameters, and then obtain the flight state of the flight platform. The reinforcement learning agent is used to determine the target flight control parameters according to the flight state, and the sliding mode controller is used to output a control instruction for the power system according to the target flight control parameters. The power system can be powered by the power battery and execute the control instruction output by the sliding mode controller.
[0044] In some embodiments, when the flight platform 11 is a compound-wing flying car, its power system includes the power system required for fixed-wing flight (such as a turbofan engine) and the power system required for multi-rotor flight (such as a motor equipped with a propeller). Or, in some embodiments, the same power system (such as a motor equipped with a propeller) can be shared for fixed-wing flight and multi-rotor flight. The power system can adjust the angle according to the flight mode to adapt to flight. In this case, when the power system executes the control instruction output by the sliding mode controller, it also includes the control of the angle of the power system.
[0045] In some embodiments, the flight platform 11 can also be configured with different carrying systems according to the type of aircraft. For example, when the flight platform 11 is a compound-wing flying car, the carrying system can be a passenger cabin. When the flight platform 11 is a cargo drone, the carrying system can be a cargo hold.
[0046] See Figure 1 , the flight process of the flight platform 11 taking off from address A and landing at address B can include a takeoff stage, a takeoff transition stage, a cruise stage, a landing transition stage, and a landing stage, and each stage corresponds to a flight mode.
[0047] During the entire flight process, the flight platform 11 can perform flight control for each stage based on the built-in flight control system and adopt the flight control method provided in this specification. The flight control method provided in this specification can, at different stages, determine corresponding target flight control parameters according to the flight state of the flight platform 11 (including flight state parameters and target flight modes), and perform flight control on the flight platform based on the target flight control parameters.
[0048] In some embodiments, the flight control method provided in this specification can be executed by the flight control system of the flight platform 11. At this time, the flight control system of the flight platform 11 can store data or instructions for executing the flight control method described in this specification, and can execute or be used to execute the data or instructions. In some embodiments, the flight control system of the flight platform 11 can include a hardware device with data information processing functions and necessary programs for driving the hardware device to work.
[0049] The flight control system of the flight platform 11 can correspond to a single computing device or a computing cluster composed of multiple computing devices. For example, the flight control system can be deployed in the computing device on the flight platform 11; or, the flight control system can also be distributedly deployed in multiple computing devices on the flight platform 11; or, part of it can be deployed in the computing device on the flight platform 11, and the other part can be deployed on the cloud server network-connected to the computing device.
[0050] Figure 3 Shows a hardware structure diagram of a computing system provided according to an embodiment of this specification. The computing system can be used as Figure 1 the flight control system of the flight platform 11 in and execute the flight control method described in this specification.
[0051] As Figure 3 shown, the computing system 200 can include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 can also include a communication port 250 and an internal communication bus 210. The computing system 200 can also include I / O components 260.
[0052] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250, and the I / O components 260, etc.
[0053] The I / O components 260 support input / output between the computing system 200 and other components.
[0054] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and a network. The communication port 250 can be a wired communication port or a wireless communication port.
[0055] The storage medium 230 can include a data storage device. The data storage device can be a non-transitory storage medium or a transitory storage medium. For example, the data storage device can include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. The storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set can include computer program code, and the computer program code can include programs, routines, objects, components, data structures, procedures, modules, and so on.
[0056] At least one processor 220 can be communicatively connected to at least one storage medium 230. When the computing system 200 is running, at least one processor 220 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the flight control method provided in this specification. The processor 220 can execute the steps included in the flight control method. The processor 220 can be in the form of one or more processors. In some embodiments, the processor 220 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof.
[0057] For illustrative purposes only, only one processor 220 is shown for the computing system 200 in the drawings. However, it should be noted that the computing system 200 in this specification can also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification can be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 220 of the computing system 200 executes step A and step B, it should be understood that step A and step B can also be jointly or separately executed by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0058] Figure 4The flowchart of a flight control method provided according to an embodiment of this specification is shown. As mentioned before, the flight control system deployed in the flight platform 11 can execute the flight control method of this specification.
[0059] As Figure 4 shown, the flight control method may include:
[0060] S410: Obtain the flight state of the flight platform, where the flight state includes flight state parameters and a target flight mode.
[0061] In some embodiments, the flight state parameters at least include: attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters.
[0062] In some embodiments, the flight control system can obtain the attitude parameters and acceleration parameters by reading the measurement results of sensors set on the flight platform. For example, the flight control system can read the measurement results of the attitude angle sensor set on the flight platform to obtain attitude parameters such as pitch angle, roll angle, and heading angle. The flight control system can also read the parameters of the accelerometer set on the flight platform to obtain acceleration parameters.
[0063] In some embodiments, the flight control system can obtain energy consumption parameters such as the remaining power and discharge power corresponding to the power battery through the battery state sensor set on the power battery in the flight platform.
[0064] In some embodiments, the flight control system can obtain the speed parameter based on the measurement results of the speed sensor set on the flight platform.
[0065] As an example, the speed sensor can be a pitot tube. At least one pitot tube can be set on the flight platform, and the flight control system can obtain the flight speed of the flight platform based on the measurement results of the pitot tube. Then, the flight control system can calculate speed parameters such as the longitudinal speed, lateral speed, and vertical speed of the flight platform according to the attitude parameters and flight speed.
[0066] In some embodiments, the flight modes of the flight platform at least include a takeoff mode, a cruise mode, a transition flight mode, and a landing mode. During the flight of the flight platform, different stages can each correspond to a flight mode. As an example, referring to Figure 1 , the takeoff stage of the flight platform can correspond to the takeoff mode, the takeoff transition stage and the landing transition stage of the flight platform can correspond to the transition mode, the cruise stage of the flight platform can correspond to the cruise mode, and the landing stage of the flight platform can correspond to the landing mode.
[0067] Among them, in the takeoff mode, the flying platform needs to provide a large thrust in the vertical direction to support its vertical lift-off. The flight control system can adjust the attitude, speed, and acceleration of the flying car, quickly respond to the instantaneous needs of the flying car, and avoid unstable dynamic behaviors during the takeoff process of the flying car.
[0068] In some embodiments, the flight control system obtains the target flight mode of the flying platform, including: determining the target flight mode of the flying platform according to the corresponding flight route of the flying platform, the geographical location information of the flying platform, and the flight altitude of the flying platform. The target flight mode is one of the flight modes of the flying platform.
[0069] Among them, before takeoff, the flying platform needs to set the corresponding flight route for the flight process. The flight route corresponding to the flying platform may include the starting point, the ending point, the cruising altitude, and multiple waypoints of the flight.
[0070] In some embodiments, when the flying platform flies along the flight route, the geographical location information of the flying platform can be compared with the geographical location information of the starting point, the ending point, and the waypoints in the flight route, and the flight altitude of the flying platform can be compared with the cruising altitude in the flight route, so as to determine the target flight mode of the flying platform. As an example, the target flight mode is one of the takeoff mode, the transition mode, the cruising mode, and the landing mode.
[0071] As an example, if the geographical location of the flying platform obtained by the flight control system is the same as or close to the geographical location of the starting point, and the flight altitude of the flying platform is less than or equal to the preset takeoff altitude, it can be determined that the flying platform is in the takeoff stage, that is, the target flight mode is the takeoff mode.
[0072] As an example, if the geographical location of the flying platform obtained by the flight control system is between the starting point and the first waypoint, and the flight altitude of the flying platform is greater than the preset takeoff altitude and less than the cruising altitude, it can be determined that the flying platform is in the takeoff transition stage, that is, the target flight mode is the transition mode.
[0073] As an example, if the geographical location of the flying platform obtained by the flight control system is between the first waypoint and the last waypoint, and the flight of the flying platform is the same as or close to the cruising altitude, it can be determined that the flying platform is in the cruising stage, that is, the target flight mode is the cruising mode.
[0074] As an example, if the geographical location of the flying platform obtained by the flight control system is between the last waypoint and the ending point, and the flight altitude of the flying platform is greater than the preset takeoff altitude and less than the cruising altitude, it can be determined that the flying platform is in the landing transition stage, that is, the target flight mode is the transition mode.
[0075] As an example, if the geographical location of the flight platform obtained by the flight control system is the same as or close to the geographical location of the destination, and the flight altitude of the flight platform is less than or equal to the preset takeoff altitude, it can be determined that the flight platform is in the landing phase, that is, the target flight mode is the landing mode.
[0076] It should be noted that two geographical locations being the same or close can mean that the coordinates corresponding to the two geographical locations are the same, or the distance between the coordinates corresponding to the two geographical locations is less than the preset distance threshold.
[0077] S420: According to the flight state of the flight platform, through a pre-trained reinforcement learning agent, obtain the target flight control parameters of the flight platform. The target flight control parameters of the flight platform at least include sliding mode control parameters. The sliding mode control parameters are at least used to characterize the target flight state parameters corresponding to the flight platform when it is in the target flight mode. The reinforcement learning agent takes the flight state of the flight platform as the state input and the control parameters of the flight platform as the action output, and aims to make the reward value of the preset reward function meet the preset conditions, and is trained through the reinforcement learning method.
[0078] In some embodiments, when creating a reinforcement learning agent, it is necessary to first construct the state space, action space, and reward function of the reinforcement learning agent. Among them, the state space is the input of the reinforcement learning agent. In this specification, the state space can include all types of parameters in the flight state. The action space is the output of the reinforcement learning agent. In this specification, the action space can include the target flight control parameters.
[0079] In some embodiments, after the flight control system inputs the obtained flight state into the reinforcement learning agent, the reinforcement learning agent can input the target flight control parameters of the flight platform corresponding to the flight state. Then, the flight control system can control the flight platform according to the target flight control parameters. Among them, the reward function can be used to characterize the control effect of the target flight control parameters on the flight platform. The higher the reward value of the reward function, the better the control effect of the target flight control parameters on the flight platform.
[0080] In some embodiments, the flight control system can train the constructed reinforcement learning agent, and can perform multiple rounds of iteration on the reinforcement learning agent. The goal of multiple rounds of iteration is that the reward value of the preset reward function meets the preset conditions. As an example, the preset condition can be to maximize the reward value of the preset reward function.
[0081] That is to say, after multiple rounds of iteration, the reward values obtained in two consecutive rounds of iteration are the same or close. Among them, the reward values obtained in two rounds of iteration being the same or close can mean that the numerical values of the two reward values are the same, or the difference between the numerical values of the two reward values is less than the preset difference threshold.
[0082] Figure 5 Shows a schematic diagram of training an enhanced learning agent provided according to an embodiment of this specification.
[0083] Referring to Figure 5 , in the t-th round of iteration, the flight control system can adjust the configuration parameters of the enhanced learning agent according to the reward value of the (t - 1)-th round of iteration, and input the flight state obtained in the (t - 1)-th round of iteration into the enhanced learning agent to generate the target flight control parameters for the t-th round of iteration. Then, based on the target flight control parameters of the t-th round of iteration, the flight control system generates the flight state of the t-th round of iteration by interacting with the environment, and obtains the reward value of the t-th round of iteration based on the flight state of the t-th round of iteration and a preset reward function. Finally, the flight state of the t-th round of iteration and the reward value of the t-th round of iteration are used for the (t + 1)-th round of iteration.
[0084] In some embodiments, Figure 5 the environment shown in can implement the flight state of a simulated flight platform and the environmental state of the target space where the simulated flight platform is located.
[0085] For example, for the flight control system to generate the flight state of the t-th round of iteration by interacting with the environment based on the target flight control parameters of the t-th round of iteration, the target flight control parameters of the t-th round of iteration can be first input into the environment, and then the environment simulates the flight state when the flight platform flies in the target space according to the target flight control parameters, so as to generate the flight state of the t-th round of iteration.
[0086] In some embodiments, the flight state parameters at least include: attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters. Then the preset reward function can at least include: attitude error reward function, acceleration error reward function, speed error reward function, and energy consumption reward function. Among them, the reward values of the attitude error reward function, acceleration error reward function, and speed error reward function are negatively correlated with the errors, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
[0087] As an example, the reward value obtained by the flight control system according to the preset reward function can be the sum of the reward values corresponding to the attitude error reward function, acceleration error reward function, speed error reward function, and energy consumption reward function.
[0088] In some embodiments, the flight control system controls the flight platform based on the sliding mode control method. In this case, the target flight control parameters may include the sliding mode control parameters, the control gain parameters corresponding to the target flight mode, and the change rate of the control input. Among them, the sliding mode control parameters may include the parameters of the target sliding mode surface and the parameters of the target control law. The control gain parameter is the weight coefficient in the target control law, which is used to characterize the target response rate of the flight platform when responding to the target control law in the target flight mode. The control gain parameter is positively correlated with the target response rate. That is, the larger the control gain parameter, the faster the response rate of the flight platform when responding to the target control law. The change rate of the control input is used to characterize the change rate when controlling the control input (such as the control steering gear, the control throttle / throttle).
[0089] In some embodiments, when applying the pre-trained reinforcement learning agent to the flight control system of the flight platform, its usage method is similar to the iterative method during training. The difference is that when applying the pre-trained reinforcement learning agent to the flight control system of the flight platform, the flight control system is based on the flight state obtained by the sensors set on the flight platform.
[0090] In this embodiment, the flight control system trains to obtain a reinforcement learning agent through a reinforcement learning scheme, and based on the flight state of the flight platform, obtains the target flight control parameters of the flight platform through the pre-trained reinforcement learning agent, realizes the adaptive adjustment of the target flight control parameters, and optimizes the robustness of the flight platform in different flight modes.
[0091] S430: Control the flight platform according to the target flight control parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters is less than or equal to the first threshold.
[0092] In some embodiments, the flight control system can obtain the target sliding mode surface and the target control law of the flight platform according to the sliding mode control parameters. The target sliding mode surface is used to characterize the target flight state parameters corresponding to the flight platform in the target flight mode, and the target control law is used to characterize the control logic when controlling the flight platform. Moreover, the flight control system controls the flight platform based on the target control law and the control gain parameters, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to the first threshold.
[0093] In some embodiments, the sliding mode controller in the flight control system may determine the target sliding mode surface and the target control law of the flight platform according to the sliding mode control parameters. Among them, the target sliding mode surface may be a linear or non-linear combination of target flight state parameters. For example, the target sliding mode surface may include a multi-dimensional function composed of control errors (such as attitude error and velocity error) and derivatives of control errors (such as angular velocity and acceleration error). The target control law may be a function based on flight state parameters, and is used to make the flight state parameters approach the target sliding mode surface from any initial value within a finite time. For example, the target control law may include a sign function based on the target sliding mode surface and a linear function of the derivative of the target sliding mode surface, where the linear coefficient of the linear function is an adjustment gain parameter.
[0094] In some embodiments, when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is less than or equal to the first threshold, it means that the flight state parameters of the flight platform approach the target sliding mode surface. Among them, the first threshold is determined by the reinforcement learning agent according to the flight state.
[0095] In some embodiments, when the flight state parameters of the flight platform approach the target sliding mode surface, a chattering phenomenon will occur. Here, the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface is defined as the sliding mode variable, that is, the absolute value of the sliding mode variable is closer to the first threshold when the flight state parameters are closer to the target sliding mode surface.
[0096] To improve the chattering phenomenon, the flight control system may introduce a boundary layer near the target sliding mode surface. When the sliding mode variable corresponding to the flight platform is outside the boundary layer, the target control law performs switching control on the flight platform. When the sliding mode variable corresponding to the flight platform enters the boundary layer, continuous control is performed on the flight platform based on the target control law.
[0097] To implement the above solution, in some embodiments, the flight control system may first determine the second threshold corresponding to the target sliding mode surface. Then, when the flight control system detects that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface (that is, the absolute value of the sliding mode variable) is less than or equal to the second threshold, continuous control is performed on the flight platform based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters characterized by the target sliding mode surface (that is, the absolute value of the sliding mode variable) is less than or equal to the first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
[0098] Among them, the second threshold is the thickness of the boundary layer. When the flight control system detects that the absolute value of the sliding mode variable is less than or equal to the second threshold, it can be determined that the flight state parameters of the flight platform enter the boundary layer. In this case, the flight control system can continuously control the flight platform based on the target control law and control gain parameters, so that the flight state parameters of the flight platform smoothly approach the target sliding mode surface, significantly reducing the chattering phenomenon.
[0099] In some embodiments, the flight control system can determine the second threshold according to the target flight mode corresponding to the target sliding mode surface.
[0100] As an example, the flight control system can obtain the target response rate corresponding to the target flight mode, determine the second threshold according to the target response rate, and the second threshold is negatively correlated with the target response rate.
[0101] Among them, when the boundary layer is thicker (i.e., the second threshold is larger), the longer it takes for the flight control system to control the flight state parameters of the flight platform to smoothly approach the target sliding mode surface through the target control law. That is to say, the thickness of the boundary layer will affect the response rate of the flight platform to sliding mode control. And different flight modes have different requirements for the response rate.
[0102] In some embodiments, the flight control system can configure corresponding second thresholds for different flight modes respectively. For example, when the flight mode requires a faster response rate, the flight control system can configure a smaller second threshold for this flight mode. When the flight control system determines the target flight mode corresponding to the flight platform, it can determine that the second threshold corresponding to the target flight mode is the second threshold corresponding to the target sliding mode surface.
[0103] In some embodiments, the second threshold corresponding to the target sliding mode surface can also be directly determined by the reinforcement learning agent. For example, when iteratively training the reinforcement learning agent, the second threshold corresponding to the target sliding mode surface can be used as a parameter in the action space, that is, the target flight control parameters output by the reinforcement learning agent also include the second threshold.
[0104] The flight control system can output the target flight control parameters including the second threshold according to the flight state parameters and the target flight mode through the reinforcement learning agent, and then determine the second threshold corresponding to the target sliding mode surface.
[0105] In this embodiment, the flight control system can dynamically adjust the thickness of the boundary layer according to the target flight mode corresponding to the flight platform, so that when performing sliding mode control on the flight platform, it can not only significantly reduce the chattering phenomenon, but also ensure the response rate of the flight platform to sliding mode control, improving the stability of the flight platform.
[0106] In some embodiments, when the absolute value of the difference between the flight state parameters of the flight platform detected by the flight control system and the target flight state parameters characterized by the target sliding mode surface (i.e., the absolute value of the sliding mode variable) is less than or equal to a second threshold, the control frequency of the target control law can be gradually increased until continuous control of the flight platform is performed.
[0107] The flight control system gradually increasing the control frequency of the target control law indicates that the sliding mode controller will respond faster to changes in the flight state parameters. Moreover, as the control frequency increases, the interval of the control commands output by the sliding mode controller can be shorter, the control granularity of the control action commands that can be output can be finer, and the control accuracy can be higher. When the control frequency is high enough, the control performed by the sliding mode controller on the flight platform can be regarded as continuous control.
[0108] In some embodiments, the target control law includes a first control law and a second control law, and the control function corresponding to the second control law is a continuous function. When the absolute value of the difference between the flight state parameters of the flight platform detected and the target flight state parameters characterized by the target sliding mode surface (i.e., the absolute value of the sliding mode variable) is greater than the second threshold, the flight platform is controlled based on the first control law and the control gain parameter. When the absolute value of the difference between the flight state parameters of the flight platform detected and the target flight state parameters characterized by the target sliding mode surface (i.e., the absolute value of the sliding mode variable) is less than or equal to the second threshold, continuous control of the flight platform is performed based on the second control law and the control gain parameter.
[0109] The flight control system can use different control laws outside and inside the boundary layer respectively. For example, the flight control system can use the first control law (e.g., the switching control law) when the sliding mode variable is outside the boundary layer; and use the second control law (e.g., the continuous control law) when the sliding mode variable is inside the boundary layer.
[0110] As an example, the target control law u(t) can be expressed by the following formula:
[0111]
[0112] u(t) = -K·sat(φS(x)) (second control law)
[0113] Among them, sat() is a saturation function, K is a control gain parameter, φ represents the second threshold, and S(x) is a sliding mode variable. In this specification, S(x) represents the sliding mode variable. When |S(x)| is greater than φ (i.e., the sliding mode variable is outside the boundary layer), the flight control system controls the flight platform based on the first control law and the control gain parameter. When |S(x)| is less than or equal to φ (i.e., the sliding mode variable is inside the boundary layer), continuous control is performed on the flight platform based on the second control law and the control gain parameter. In this case, the control input will saturate, thus avoiding the occurrence of chattering phenomena.
[0114] In summary, this specification provides a flight control method system and a storage medium. The flight control system can obtain the target flight control parameters of the flight platform (which may include sliding mode control parameters such as sliding mode surface parameters, control laws, etc.) through a pre-trained reinforcement learning agent based on the flight state parameters and the target flight mode of the flight platform. Then, the flight platform is controlled according to the target flight control parameters, so that the flight state parameters of the flight platform approach the target flight state parameters indicated by the target flight control parameters. The control system can obtain the target flight control parameters of the flight platform through the reinforcement learning agent according to the flight state parameters and the target flight mode of the flight platform, and the flight control parameters do not need to be set manually. Moreover, the target flight control parameters obtained through the reinforcement learning agent are control strategies with better effects under the current flight state and the current flight mode, and the target flight control parameters can be dynamically updated as the flight state or the flight mode changes, ensuring the stability and maneuverability of the flight platform.
[0115] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for performing flight control. When the at least one instruction set is executed by a processor, the at least one instruction set directs the processor to implement the steps of the flight control method described in this specification. In some possible implementation manners, various aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on the computing system 200, the program code is used to cause the computing system 200 to execute the steps of the flight control method described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) including program code and can run on the computing system 200. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the computing system 200, partially on the computing system 200, executed as an independent software package, partially on the computing system 200 and partially on a remote computing device, or entirely on a remote computing device.
[0116] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0117] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0118] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0119] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature, and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well mark out some of the devices as separate embodiments when reading this specification. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0120] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for those that are inconsistent with or conflict with this document, or those that have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. Further, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0121] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Accordingly, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A flight control method, applied to a flight control system of a flight platform, the method comprising: Acquiring a flight state of the flight platform, wherein the flight state includes flight state parameters and a target flight mode; According to the flight state of the flight platform, target flight control parameters of the flight platform are obtained through a pre-trained reinforcement learning agent, wherein the target flight control parameters of the flight platform at least include sliding mode control parameters, and the sliding mode control parameters are at least used to characterize the target flight state parameters corresponding to when the flight platform is in the target flight mode, and the reinforcement learning agent is trained by reinforcement learning with the flight state of the flight platform as state input and the control parameters of the flight platform as action output, and with the goal of the reward value of a preset reward function meeting a preset condition; and The flight platform is controlled according to the target flight control parameter so that an absolute value of a difference between a flight state parameter of the flight platform and the target flight state parameter is less than or equal to a first threshold.
2. The method according to claim 1, wherein: The target flight control parameters also include control gain parameters corresponding to the target flight mode; The step of controlling the flight platform according to the target flight control parameter comprises: Obtaining a target sliding mode surface and a target control law of the flight platform according to the sliding mode control parameters, wherein the target sliding mode surface is used to characterize a target flight state parameter corresponding to the flight platform when the flight platform is in the target flight mode, and the target control law is used to characterize a control logic when controlling the flight platform; and The flight platform is controlled based on the target control law and the control gain parameter so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to a first threshold value. The control gain parameter is used to characterize the target response rate of the flight platform when responding to the target control law in the target flight mode, and the control gain parameter is positively correlated with the target response rate.
3. The method according to claim 2, wherein: The controlling the flight platform based on the target control law and the control gain parameter so that an absolute value of a difference between a flight state parameter of the flight platform and a target flight state parameter represented by the target sliding surface is less than or equal to a first threshold value includes: Determining a second threshold corresponding to the target sliding surface; and When it is detected that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a second threshold, the flight platform is continuously controlled based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold, and the absolute value of the second threshold is greater than the absolute value of the first threshold.
4. The method according to claim 3, wherein: When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to a second threshold, continuously controlling the flight platform based on the target control law and the control gain parameter includes: When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to a second threshold, the control frequency of the target control law is gradually increased until the flight platform is continuously controlled.
5. The method according to claim 3, wherein: The target control law includes a first control law and a second control law, and the control function corresponding to the second control law is a continuous function; When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is greater than a second threshold, controlling the flight platform based on the first control law and the control gain parameter; When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to a second threshold, continuously controlling the flight platform based on the target control law and the control gain parameter includes: When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to a second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
6. The method according to claim 3, wherein: The determining a second threshold value corresponding to the target sliding surface includes: The second threshold is determined according to a target flight mode corresponding to the target sliding surface.
7. The method according to claim 6, wherein: The determining the second threshold according to the target flight mode corresponding to the target sliding surface includes: acquiring the target response rate corresponding to the target flight mode, and determining the second threshold according to the target response rate, wherein the second threshold is negatively correlated with the target response rate; or The target flight control parameters include the second threshold, which is determined by the reinforcement learning agent according to the flight state parameters and the target flight mode.
8. The method according to claim 1, wherein: The flight modes of the flight platform include at least a take-off mode, a cruise mode, a transition flight mode and a landing mode; The obtaining of the target flight mode of the flight platform includes: A target flight mode of the flight platform is determined according to the flight route corresponding to the flight platform, the geographical location information of the flight platform, and the flight altitude of the flight platform. The target flight mode is one of the flight modes of the flight platform.
9. The method according to claim 1, wherein: The flight state parameters include at least: attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters; The preset reward function includes at least: a posture error reward function, an acceleration error reward function, a speed error reward function, and an energy consumption reward function, wherein the reward values of the posture error reward function, the acceleration error reward function, and the speed error reward function are negatively correlated with the errors, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
10. The method according to claim 9, wherein: The reward value of the preset reward function meets the preset conditions, including: The reward value of the preset reward function is maximized.
11. A flight control system comprising: at least one storage medium storing at least one instruction set for performing flight control; as well as at least one processor, in communication with the at least one storage medium, wherein, when the flight control system is running, the at least one processor reads the at least one instruction set, And execute the flight control method described in any one of claims 1-10 according to the instructions of the at least one instruction set.
12. A computer-readable non-volatile storage medium, wherein: The computer-readable non-volatile storage medium stores at least one instruction set, and when the at least one instruction set is executed by at least one processor, the flight control method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Aircraft intelligent sliding mode formation control method based on reinforcement learning
CN114545979A
Training method of aircraft six-degree-of-freedom control model and aircraft control method
CN116382326A
Parameter setting method and system for sliding mode controller
CN117111476A
Self-healing control method for multi-mode vertical take-off and landing aircraft
CN117250867A
A method for training a reinforcement learning agent to provide inputs to an automated system on an aircraft
GB2631381A