Flight control method, system, and storage medium
By using a pre-trained reinforcement learning agent and a dynamic adjustment method, the problem of sliding surface parameters relying on empirical settings, which was not effectively addressed in the prior art, was solved, thereby optimizing the stability and maneuverability of the compound wing flight platform under different flight modes.
Patent Information
- Application Number
- CN202510280743.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In existing sliding mode control methods, the sliding surface parameters rely on the developer's experience to set, and cannot be adjusted according to the actual flight mode of the compound wing flight platform, which affects stability and maneuverability.
A pre-trained reinforcement learning agent is used to dynamically adjust sliding mode control parameters, including sliding surfaces and control laws, based on the flight state of the flight platform and the target flight mode, thereby optimizing the flight control strategy through reinforcement learning.
It achieves dynamic optimization of the stability and maneuverability of the flight platform under different flight modes, improving the adaptability and effectiveness of flight control.
Smart Images

Figure CN120143680B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of flight control, and particularly relates to a flight control method, system and storage medium. BACKGROUND
[0002] A compound wing flight platform (such as a compound wing flying car) is a new type of aircraft combining the characteristics of fixed wings and rotors, and has the ability of vertical take-off and landing and flat flight cruising. The traditional flight control method cannot be applied to the compound wing flight platform due to the complex wing type and the changeable flight mode (different flight modes need to be used in different flight stages).
[0003] In the existing scheme, the flight parameters of the compound wing flight platform can be controlled by the method of sliding mode control, and then the flight control of the compound wing flight platform in different flight modes is realized.
[0004] However, in the existing sliding mode control process, the sliding mode surface parameters are pre-set based on the experience of the developer, which requires higher experience of the developer, and the pre-set sliding mode surface parameters cannot be adjusted according to the actual flight mode, which may affect the stability and maneuverability of the compound wing flight platform.
[0005] The content of the background section only represents the inventor's own knowledge and does not mean that the above information has entered the public domain before the filing date of the present disclosure, nor does it mean that it can be prior art of the present disclosure. SUMMARY
[0006] The present specification provides a flight control method, system and storage medium, which can match the corresponding flight control parameters in different flight modes, and ensure the stability and maneuverability of the flight platform.
[0007] In a first aspect, the present specification provides a flight control method applied to a flight control system of a flight platform, the method comprising: obtaining a flight state of the flight platform, the flight state comprising a flight state parameter and a target flight mode; obtaining a target flight control parameter of the flight platform according to the flight state of the flight platform by a pre-trained reinforcement learning agent, the target flight control parameter of the flight platform at least comprising a sliding mode control parameter, the sliding mode control parameter at least being used to represent a target flight state parameter corresponding to the target flight mode of the flight platform, the reinforcement learning agent being trained by reinforcement learning in a manner that a reward value of a preset reward function meets a preset condition, with the flight state of the flight platform as a state input and the control parameter of the flight platform as an action output; and controlling the flight platform according to the target flight control parameter, so that an absolute value of a difference between the flight state parameter of the flight platform and the target flight state parameter is less than or equal to a first threshold value.
[0008] In some embodiments, the target flight control parameter further comprises a control gain parameter corresponding to the target flight mode; the controlling the flight platform according to the target flight control parameter comprises: obtaining a target sliding mode surface and a target control law of the flight platform according to the sliding mode control parameter, the target sliding mode surface being used to represent the target flight state parameter corresponding to the target flight mode of the flight platform, and the target control law being used to represent a control logic when the flight platform is controlled; and controlling the flight platform based on the target control law and the control gain parameter, so that an absolute value of a difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the first threshold value, the control gain parameter being used to represent a target response rate of the flight platform in the target flight mode in response to the target control law, and the control gain parameter being positively correlated with the target response rate.
[0009] In some embodiments, the controlling the flight platform based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the first threshold value, comprises: determining a second threshold value corresponding to the target sliding mode surface; and when detecting that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value, continuously controlling the flight platform based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the first threshold value, and an absolute value of the second threshold value is greater than an absolute value of the first threshold value.
[0010] In some embodiments, when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold, the continuous control of the flight platform based on the target control law and the control gain parameter comprises: gradually increasing the control frequency of the target control law until the continuous control of the flight platform is performed when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold.
[0011] In some embodiments, the target control law comprises a first control law and a second control law, and the second control law corresponds to a continuous function; and when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is greater than the second threshold, the flight platform is controlled based on the first control law and the control gain parameter.
[0012] When the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold, the continuous control of the flight platform based on the target control law and the control gain parameter comprises: when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0013] The target control law comprises a first control law and a second control law, and the second control law corresponds to a continuous function; and when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold, the continuous control of the flight platform based on the target control law and the control gain parameter comprises: the flight platform is controlled based on the first control law and the control gain parameter; and when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0014] In some embodiments, the determination of the second threshold corresponding to the target sliding mode surface comprises: determining the second threshold according to the target flight mode corresponding to the target sliding mode surface.
[0015] In some embodiments, the determining the second threshold value according to the target flight mode corresponding to the target sliding mode surface comprises: obtaining the target response rate corresponding to the target flight mode, determining the second threshold value according to the target response rate, and the second threshold value is negatively correlated with the target response rate; or the target flight control parameter comprises the second threshold value, and the second threshold value is determined by the reinforcement learning agent according to the flight state parameter and the target flight mode.
[0016] In some embodiments, the flight mode of the flight platform comprises at least a take-off mode, a cruising mode, a transition flight mode, and a landing mode; and the obtaining the target flight mode of the flight platform comprises: determining the target flight mode of the flight platform according to a corresponding route of the flight platform, geographical position information of the flight platform, and a flight height of the flight platform, and the target flight mode is one of the flight modes of the flight platform.
[0017] In some embodiments, the flight state parameter comprises at least an attitude parameter, an acceleration parameter, a speed parameter, and an energy consumption parameter; and the preset reward function comprises at least an attitude error reward function, an acceleration error reward function, a speed error reward function, and an energy consumption reward function, wherein the reward values of the attitude error reward function, the acceleration error reward function, and the speed error reward function are negatively correlated with errors, and the reward value of the energy consumption reward function is negatively correlated with energy consumption.
[0018] In some embodiments, the reward value of the preset reward function meets a preset condition, which comprises: the reward value of the preset reward function is maximized.
[0019] In a second aspect, the specification provides a flight control system, comprising: at least one storage medium storing at least one instruction set for performing flight control; and at least one processor in communication connection with the at least one storage medium, wherein when the flight platform is running, the at least one processor reads the at least one instruction set, and executes the flight control method provided in the first aspect according to the indication of the at least one instruction set.
[0020] In a third aspect, the specification further provides a computer-readable non-volatile storage medium, wherein the computer-readable non-volatile storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the flight control method provided in the first aspect.
[0021] Other functions of the flight control method, system and storage medium provided by the present specification will be partially listed in the following description. The creative aspects of the flight control method, system and storage medium provided by the present specification can be fully explained by practicing or using the methods, devices and combinations described in the following detailed examples. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present specification, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0023] Figure 1 An application scenario schematic diagram of a flight control method provided by an embodiment of the present specification is shown;
[0024] Figure 2 A structural schematic diagram of a flight platform provided by an embodiment of the present specification is shown;
[0025] Figure 3 A hardware structure diagram of a computing system provided by an embodiment of the present specification is shown;
[0026] Figure 4 A flowchart of a flight control method provided by an embodiment of the present specification is shown; and
[0027] Figure 5 A schematic diagram of reinforcement learning agent training provided by an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0028] The following description provides specific application scenarios and requirements of the present specification, which is to enable those skilled in the art to manufacture and use the contents in the present specification. Various partial modifications of the disclosed embodiments are obvious to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present specification. Therefore, the present specification is not limited to the shown embodiments, but is consistent with the widest scope of the claims.
[0029] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. For example, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein the terms "includes", "including", "has", "having" and / or "contains", "containing" means that the listed element(s) is(are) present and not excluding the presence of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0030] These and other features of the present specification, as well as the structure and operation of the various elements of the system and the combination of parts and economies of manufacture will become more apparent upon consideration of the following description and the accompanying drawings. Aspects of the description will appear from the following description and the associated drawings, which are intended to illustrate and not to limit the scope of the present specification. It should be understood that the drawings are not necessarily drawn to scale.
[0031] The flow diagrams herein illustrate the operations of system implementations in accordance with some embodiments of the present specification. It should be clearly understood that the operations of the flow diagrams can not be implemented in order. Rather, the operations can be implemented in reverse order or simultaneously. Furthermore, one or more other operations can be added to the flow diagrams. One or more operations can be removed from the flow diagrams.
[0032] For the convenience of description, the terms to be used in the latter part of the present specification are first explained.
[0033] Term 1: Sliding Mode Control Method. Sliding Mode Control (SMC) is a nonlinear control strategy, the core idea of which is to make the state trajectory of the controlled object reach and slide on a sliding surface through a control law.
[0034] Term 2: Sliding Surface. The sliding surface is a specific hyperplane in the state space of the controlled object, once the dynamic behavior of the controlled object reaches this hyperplane, it will slide along it to reach the desired state or output.
[0035] Term 3: Control Law. The control law is a set of rules or equations that define how to calculate the control input based on the current state or output of the controlled object and the desired target state or output. That is, the control law is the logic or mathematical expression used by the sliding mode controller to determine its control action.
[0036] Term 4: sliding mode controller. The sliding mode controller is a virtual component for implementing the sliding mode control method, which includes a sliding surface and a control law. The input parameter of the sliding mode controller is the target flight control parameter, and the output parameter is the control instruction of the flight platform power system. In this specification, when describing that the flight control system implements the sliding mode control method, it means that the flight control system implements the sliding mode control method through the sliding mode controller.
[0037] Term 5: chattering. Chattering refers to high-frequency vibration caused by high-frequency switching control of the target control law on the flight platform when the flight state parameter reaches the target sliding surface in the sliding mode control method.
[0038] The application scenario of the present specification is introduced as follows.
[0039] The technical solution provided by the present specification is applicable to the scenario of sliding mode control of a compound wing flight platform. In this scenario, the current technical solution is that the sliding surface is pre-configured according to experience when the compound wing flight platform is controlled by the sliding mode control method. The flight control effect of the pre-set sliding surface depends on the experience of the developer. When the developer lacks experience, the pre-set sliding surface cannot provide good flight control effect. Moreover, the pre-set sliding surface cannot be dynamically adjusted according to the flight mode of the compound wing flight platform, which may affect the stability and maneuverability of the compound wing flight platform.
[0040] The present specification provides a flight control method, which can be applied to the scenario of sliding mode control of a compound wing flight platform. In this scenario, the flight control system can obtain the target flight control parameter (which can include sliding mode control parameters such as sliding surface parameters, control laws, etc.) of the flight platform based on the flight state parameter of the flight platform and the target flight mode through a pre-trained reinforcement learning agent. Then, the flight platform is controlled according to the target flight control parameter, so that the flight state parameter of the flight platform approaches the target flight state parameter indicated by the target flight control parameter. The control system can obtain the target flight control parameter of the flight platform through the reinforcement learning agent according to the flight state parameter of the flight platform and the target flight mode. The flight control parameter does not need to be set manually. Moreover, the target flight control parameter obtained through the reinforcement learning agent is a control strategy with good effect under the current flight state and the current flight mode. The target flight control parameter can be dynamically updated as the flight state or flight mode changes, ensuring the stability and maneuverability of the flight platform.
[0041] Figure 1 An application scenario diagram of a flight control method provided by an embodiment of the present specification is shown. As shown in FIG. 1, the flight control system of the compound wing flight platform includes a flight state parameter acquisition module, a target flight mode input module, a reinforcement learning agent, and a flight control module. Figure 1As shown, the scene 100 includes a flight platform 11. In the scene 100, the flight platform 11 is shown in the form of a compound wing flying car. It should be noted that in the present specification, the flight platform 11 can be any type of aircraft, for example, the flight platform 11 can also be a manned drone, a cargo drone, or a special industry drone, etc.
[0042] Figure 2 A structural schematic diagram of a flight platform provided according to an embodiment of the present specification is shown.
[0043] In some embodiments, with reference to Figure 2 , the flight platform 11 can include a flight control system, a power system, and a power battery. The flight control system can include a sensor group, a reinforcement learning agent, and a sliding mode controller. The sensor group includes a plurality of sensors for collecting flight state parameters of the flight platform 11, the flight control system can determine a target flight mode of the flight platform 11 according to the flight state parameters, and then obtain the flight state of the flight platform. The reinforcement learning agent is used to determine a target flight control parameter according to the flight state, and the sliding mode controller is used to output a control instruction of the power system according to the target flight control parameter. The power system can be powered by the power battery and execute the control instruction output by the sliding mode controller.
[0044] In some embodiments, for the flight platform 11 being a compound wing flying car, the power system includes a power system required for fixed wing flight (such as a turbofan engine) and a power system required for multi-rotor flight (such as a motor equipped with a propeller). Alternatively, in some embodiments, a set of power systems (such as a motor equipped with a propeller) can be shared for fixed wing flight and multi-rotor flight, and the power system can adjust the angle according to the flight mode to adapt to flight. In this case, when the power system executes the control instruction output by the sliding mode controller, it also includes the control of the angle of the power system.
[0045] In some embodiments, the flight platform 11 can also be configured with different bearing systems according to the type of aircraft. For example, when the flight platform 11 is a compound wing flying car, the bearing system can be a crew cabin, and when the flight platform 11 is a cargo drone, the bearing system can be a cargo compartment.
[0046] Referring to Figure 1 , the flight process of the flight platform 11 taking off from address A and landing at address B can include a takeoff phase, a takeoff transition phase, a cruising phase, a landing transition phase, and a landing phase, each phase corresponding to a flight mode.
[0047] During the whole flight process, the flight platform 11 can perform flight control on each stage based on the built-in flight control system and the flight control method provided in the specification. The flight control method provided in the specification can determine the corresponding target flight control parameter according to the flight state (including the flight state parameter and the target flight mode) of the flight platform 11 in different stages, and perform flight control on the flight platform based on the target flight control parameter.
[0048] In some embodiments, the flight control method provided in the specification can be executed by the flight control system of the flight platform 11. At this time, the flight control system of the flight platform 11 can store data or instructions for executing the flight control method described in the specification, and can execute or be used to execute the data or instructions. In some embodiments, the flight control system of the flight platform 11 can include a hardware device with data information processing function and the necessary program required to drive the hardware device to work.
[0049] The flight control system of the flight platform 11 can correspond to a single computing device, or a computing cluster composed of multiple computing devices. For example, the flight control system can be deployed in the computing device on the flight platform 11; or the flight control system can also be distributed and deployed in multiple computing devices on the flight platform 11; or it can also be partially deployed in the computing device on the flight platform 11, and the other part is deployed on the cloud server connected with the computing device through network.
[0050] Figure 3 A hardware structure diagram of a computing system provided according to an embodiment of the specification is shown. The computing system can be used as Figure 1 the flight control system of the flight platform 11, to execute the flight control method described in the specification.
[0051] As Figure 3 shown, the computing system 200 can include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 can also include a communication port 250 and an internal communication bus 210. The computing system 200 can also include an I / O component 260.
[0052] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250 and the I / O component 260, etc.
[0053] The I / O component 260 supports the input / output between the computing system 200 and other components.
[0054] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and a network. The communication port 250 can be a wired communication port or a wireless communication port.
[0055] The storage medium 230 can include a data storage device. The data storage device can be a non-transitory storage medium or a transitory storage medium. For example, the data storage device can include one or more of a disk 232, a read-only memory (ROM) 234, or a random access memory (RAM) 235. The storage medium 230 also includes at least one set of instructions stored in the data storage device. The set of instructions can include computer program code, which can include programs, routines, objects, components, data structures, procedures, modules, and the like.
[0056] The at least one processor 220 can be communicatively connected to the at least one storage medium 230. When the computing system 200 is running, the at least one processor 220 reads the at least one set of instructions and executes the flight control method provided in the present specification according to the instructions of the at least one set of instructions. The processor 220 can execute the steps included in the flight control method. The processor 220 can be in the form of one or more processors, and in some embodiments, the processor 220 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, or the like, or any combination thereof.
[0057] For the sake of illustration only, the computing system 200 in the accompanying drawings shows only one processor 220. However, it should be noted that the computing system 200 in the present specification can also include multiple processors, and therefore, the operations and / or method steps disclosed in the present specification can be executed by one processor or jointly executed by multiple processors. For example, if the processor 220 of the computing system 200 is described in the present specification as executing step A and step B, it should be understood that step A and step B can also be executed jointly or separately by two different processors 220 (e.g., a first processor executes step A and a second processor executes step B, or the first and second processors jointly execute steps A and B).
[0058] Figure 4A flowchart of a flight control method according to an embodiment of the present specification is shown. As before, the flight control system deployed in the flight platform 11 can perform the flight control method of the present specification.
[0059] As shown in Figure 4 , the flight control method can include:
[0060] S410: Obtain the flight state of the flight platform, the flight state including flight state parameters and a target flight mode.
[0061] In some embodiments, the flight state parameters at least include attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters.
[0062] In some embodiments, the flight control system can obtain the attitude parameters and the acceleration parameters by reading the measurement results of the sensors arranged on the flight platform. For example, the flight control system can read the measurement results of the attitude angle sensors arranged on the flight platform to obtain the attitude parameters such as the pitch angle, the roll angle, and the heading angle. The flight control system can also read the parameters of the accelerometers arranged on the flight platform to obtain the acceleration parameters.
[0063] In some embodiments, the flight control system can obtain the energy consumption parameters such as the remaining power and the discharge power of the power battery corresponding to the power battery by the battery state sensors arranged on the power battery in the flight platform.
[0064] In some embodiments, the flight control system can obtain the speed parameters based on the measurement results of the speed sensors arranged on the flight platform.
[0065] As an example, the speed sensor can be an airspeed tube. At least one airspeed tube can be arranged on the flight platform, and the flight control system can obtain the flight speed of the flight platform based on the measurement results of the airspeed tube. Then, the flight control system can calculate the longitudinal speed, the lateral speed, and the vertical speed of the flight platform according to the attitude parameters and the flight speed to obtain the speed parameters.
[0066] In some embodiments, the flight mode of the flight platform at least includes a take-off mode, a cruising mode, a transition flight mode, and a landing mode. Different stages in the flight process of the flight platform can each correspond to a flight mode. As an example, referring to Figure 1 , the take-off stage of the flight platform can correspond to the take-off mode, the take-off transition stage and the landing transition stage of the flight platform can correspond to the transition mode, the cruising stage of the flight platform can correspond to the cruising mode, and the landing stage of the flight platform can correspond to the landing mode.
[0067] In the take-off mode, the flight platform needs to provide a large thrust in the vertical direction to support its vertical take-off, and the flight control system can adjust the attitude, speed and acceleration of the flying car to quickly respond to the instantaneous demand of the flying car and avoid unstable dynamic behavior of the flying car during take-off.
[0068] In some embodiments, the flight control system obtains a target flight mode of the flight platform, including: determining the target flight mode of the flight platform according to a corresponding route of the flight platform, geographical position information of the flight platform, and a flight height of the flight platform, the target flight mode being one of the flight modes of the flight platform.
[0069] In some embodiments, the flight control system obtains a target flight mode of the flight platform, including: determining the target flight mode of the flight platform according to a corresponding route of the flight platform, geographical position information of the flight platform, and a flight height of the flight platform, the target flight mode being one of the flight modes of the flight platform.
[0070] In some embodiments, the flight control system obtains a target flight mode of the flight platform, including: determining the target flight mode of the flight platform according to a corresponding route of the flight platform, geographical position information of the flight platform, and a flight height of the flight platform, the target flight mode being one of the flight modes of the flight platform.
[0071] For example, if the flight control system obtains the geographical position of the flight platform is the same or similar to the geographical position of the starting point, and the flight height of the flight platform is less than or equal to the preset take-off height, it can be determined that the flight platform is in the take-off stage, i.e., the target flight mode is the take-off mode.
[0072] For example, if the flight control system obtains the geographical position of the flight platform is between the starting point and the first waypoint, and the flight height of the flight platform is greater than the preset take-off height and less than the cruising height, it can be determined that the flight platform is in the take-off transition stage, i.e., the target flight mode is the transition mode.
[0073] For example, if the flight control system obtains the geographical position of the flight platform is between the first waypoint and the last waypoint, and the flight height of the flight platform is the same or similar to the cruising height, it can be determined that the flight platform is in the cruising stage, i.e., the target flight mode is the cruising mode.
[0074] For example, if the flight control system obtains the geographical position of the flight platform is between the last waypoint and the ending point, and the flight height of the flight platform is greater than the preset take-off height and less than the cruising height, it can be determined that the flight platform is in the landing transition stage, i.e., the target flight mode is the transition mode.
[0075] As an example, if the flight control system obtains the geographic position of the flight platform and the geographic position of the destination is the same or similar, and the flight height of the flight platform is less than or equal to the preset take-off height, it can be determined that the flight platform is in the landing stage, that is, the target flight mode is the landing mode.
[0076] It should be noted that the two geographic positions are the same or similar, that is, the coordinates corresponding to the two geographic positions are the same, or the distance between the coordinates corresponding to the two geographic positions is less than a preset distance threshold.
[0077] S420: According to the flight state of the flight platform, the target flight control parameter of the flight platform is obtained through the pre-trained reinforcement learning agent. The target flight control parameter of the flight platform at least includes a sliding mode control parameter, which is at least used to represent the corresponding target flight state parameter when the flight platform is in the target flight mode. The reinforcement learning agent is trained by taking the flight state of the flight platform as the state input, the control parameter of the flight platform as the action output, and the reward value of the preset reward function meeting the preset condition as the target.
[0078] In some embodiments, when creating the reinforcement learning agent, it is necessary to first construct the state space, action space and reward function of the reinforcement learning agent. The state space is the input of the reinforcement learning agent, which can include all types of parameters in the flight state in this specification. The action space is the output of the reinforcement learning agent, which can include the target flight control parameter in this specification.
[0079] In some embodiments, after the flight control system inputs the obtained flight state into the reinforcement learning agent, the reinforcement learning agent can input the target flight control parameter of the flight platform corresponding to the flight state. Then, the flight control system can control the flight platform according to the target flight control parameter. The reward function can be used to represent the control effect of the target flight control parameter on the flight platform, and the higher the reward value of the reward function, the better the control effect of the target flight control parameter on the flight platform.
[0080] In some embodiments, the flight control system can train the constructed reinforcement learning agent, and can perform multiple rounds of iteration on the reinforcement learning agent. The goal of multiple rounds of iteration is to make the reward value of the preset reward function meet the preset condition. As an example, the preset condition can be to maximize the reward value of the preset reward function.
[0081] That is, after multiple rounds of iteration, the reward values obtained by two consecutive rounds of iteration are the same or similar. The reward values obtained by two rounds of iteration can be the same or similar in value, or the difference between the two reward values is less than a preset difference threshold.
[0082] Figure 5 A schematic diagram of reinforcement learning agent training is shown according to an embodiment of the present specification.
[0083] Reference Figure 5 In the tthiteration, the flight control system can adjust the configuration parameters of the reinforcement learning agent according to the reward value of the (t-1) thiteration, and input the flight state obtained in the (t-1) thiteration into the reinforcement learning agent to generate the target flight control parameters of the tthiteration. Then, the flight control system generates the flight state of the tthiteration by interacting with the environment based on the target flight control parameters of the tthiteration, and obtains the reward value of the tthiteration based on the flight state of the tthiteration and the preset reward function. Finally, the flight state of the tthiteration and the reward value of the tthiteration are used for the (t+1) thiteration.
[0084] In some embodiments, Figure 5 The environment shown in FIG. 1 can simulate the flight state of the flight platform, and the environment state of the target space where the flight platform is located.
[0085] For example, the flight control system generates the flight state of the tthiteration by interacting with the environment based on the target flight control parameters of the tthiteration, which can first input the target flight control parameters of the tthiteration into the environment, and then simulate the flight state of the flight platform when flying in the target space according to the target flight control parameters through the environment, thereby generating the flight state of the tthiteration.
[0086] In some embodiments, the flight state parameters at least include attitude parameters, acceleration parameters, speed parameters, and energy consumption parameters. The preset reward function can at least include attitude error reward functions, acceleration error reward functions, speed error reward functions, and energy consumption reward functions. Among them, the reward values of the attitude error reward functions, the acceleration error reward functions and the speed error reward functions are negatively correlated with the errors, and the reward value of the energy consumption reward function is negatively correlated with the energy consumption.
[0087] As an example, the reward value obtained by the flight control system according to the preset reward function can be the sum of the reward value corresponding to the attitude error reward function, the reward value corresponding to the acceleration error reward function, the reward value corresponding to the speed error reward function, and the reward value corresponding to the energy consumption reward function.
[0088] In some embodiments, the flight control system controls the flight platform based on a sliding mode control method, in which case the target flight control parameter can include a sliding mode control parameter, a control gain parameter corresponding to the target flight mode, and a change rate of the control input. The sliding mode control parameter can include parameters of a target sliding surface and parameters of a target control law, and the control gain parameter is a weight coefficient in the target control law, used to represent a target response rate of the flight platform in response to the target control law in the target flight mode. The control gain parameter is positively correlated with the target response rate. That is, the greater the control gain parameter, the faster the response rate of the flight platform in response to the target control law. The change rate of the control input is used to represent the change rate of the control input (such as the control direction machine, the control throttle / door).
[0089] In some embodiments, when the pre-trained reinforcement learning agent is applied to the flight control system of the flight platform, the method used is similar to the iteration method during training. The difference is that when the pre-trained reinforcement learning agent is applied to the flight control system of the flight platform, the flight control system is based on the flight state obtained by the sensors arranged on the flight platform.
[0090] In this embodiment, the flight control system trains the reinforcement learning agent through the reinforcement learning scheme, and obtains the target flight control parameter of the flight platform through the pre-trained reinforcement learning agent based on the flight state of the flight platform, so as to realize the adaptive adjustment of the target flight control parameter and optimize the robustness of the flight platform in different flight modes.
[0091] S430: Control the flight platform according to the target flight control parameter, so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter is less than or equal to the first threshold value.
[0092] In some embodiments, the flight control system can obtain the target sliding surface and the target control law of the flight platform according to the sliding mode control parameter, the target sliding surface being used to represent the target flight state parameter of the flight platform in the target flight mode, and the target control law being used to represent the control logic when controlling the flight platform. In addition, the flight control system controls the flight platform based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding surface is less than or equal to the first threshold value.
[0093] In some embodiments, a sliding mode controller in the flight control system can determine a target sliding surface and a target control law for the flight platform according to the sliding mode control parameters. The target sliding surface can be a linear or nonlinear combination of the target flight state parameters. For example, the target sliding surface can include a multi-dimensional function of the control errors (e.g., attitude errors and velocity errors) and derivatives of the control errors (e.g., angular velocity, acceleration errors). The target control law can be a function of the flight state parameters for making the flight state parameters approach the target sliding surface in a finite time from an arbitrary initial value. For example, the target control law can include a sign function based on the target sliding surface and a linear function of the derivative of the target sliding surface, where the linear coefficient of the linear function is an adjustment gain parameter.
[0094] In some embodiments, when the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is less than or equal to a first threshold value, it indicates that the flight state parameters of the flight platform approach the target sliding surface. The first threshold value is determined by the reinforcement learning agent according to the flight state.
[0095] In some embodiments, when the flight state parameters of the flight platform approach the target sliding surface, chattering phenomenon occurs. Here, the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface is defined as a sliding variable, i.e., the closer the flight state parameters to the target sliding surface, the closer the absolute value of the sliding variable to the first threshold value.
[0096] To improve the chattering phenomenon, the flight control system can introduce a boundary layer near the target sliding surface, and when the corresponding sliding variable of the flight platform is outside the boundary layer, the target control law is used to switch control of the flight platform. When the corresponding sliding variable of the flight platform enters the boundary layer, the flight platform is continuously controlled based on the target control law.
[0097] To achieve the above scheme, in some embodiments, the flight control system can first determine a second threshold value corresponding to the target sliding surface. Then, when the flight control system detects that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the second threshold value, the flight platform is continuously controlled based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameters of the flight platform and the target flight state parameters represented by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the first threshold value, and the absolute value of the second threshold value is greater than the absolute value of the first threshold value.
[0098] The second threshold is the thickness of the boundary layer. When the flight control system detects that the absolute value of the sliding mode variable is less than or equal to the second threshold, it can be determined that the flight state parameter of the flight platform enters the boundary layer. In this case, the flight control system can continuously control the flight platform based on the target control law and the control gain parameter, so that the flight state parameter of the flight platform smoothly approaches the target sliding mode surface, and the chattering phenomenon is significantly reduced.
[0099] In some embodiments, the flight control system can determine the second threshold according to the target flight mode corresponding to the target sliding mode surface.
[0100] For example, the flight control system can obtain the target response rate corresponding to the target flight mode, determine the second threshold according to the target response rate, and the second threshold is negatively correlated with the target response rate.
[0101] When the boundary layer is thicker (i.e., the second threshold is larger), the flight control system takes longer to control the flight state parameter of the flight platform to smoothly approach the target sliding mode surface through the target control law. That is, the thickness of the boundary layer affects the response rate of the flight platform to the sliding mode control. Different flight modes have different requirements for the response rate.
[0102] In some embodiments, the flight control system can configure a corresponding second threshold for different flight modes. For example, when the flight mode requires a faster response rate, the flight control system can configure a smaller second threshold for the flight mode. When the flight control system determines the target flight mode corresponding to the flight platform, it can determine the second threshold corresponding to the target flight mode as the second threshold corresponding to the target sliding mode surface.
[0103] In some embodiments, the second threshold corresponding to the target sliding mode surface can also be determined directly by the reinforcement learning agent. For example, when the reinforcement learning agent is iteratively trained, the second threshold corresponding to the target sliding mode surface can be used as a parameter in the action space, that is, the target flight control parameter output by the reinforcement learning agent also includes the second threshold.
[0104] The flight control system can output the target flight control parameter including the second threshold according to the flight state parameter and the target flight mode through the reinforcement learning agent, and then determine the second threshold corresponding to the target sliding mode surface.
[0105] In this embodiment, the flight control system can dynamically adjust the thickness of the boundary layer according to the target flight mode corresponding to the flight platform, so that when the flight platform is controlled by the sliding mode, the chattering phenomenon can be significantly reduced, and the response rate of the flight platform to the sliding mode control can be guaranteed, and the stability of the flight platform can be improved.
[0106] In some embodiments, when the flight control system detects that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter characterized by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the second threshold value, the control frequency of the target control law can be gradually increased until continuous control is performed on the flight platform.
[0107] The gradual increase of the control frequency of the target control law by the flight control system indicates that the sliding mode controller will respond more quickly to changes in the flight state parameter. Moreover, as the control frequency increases, the sliding mode controller can adjust the interval of the output control instructions to be shorter, and the control granularity of the output control action instructions to be finer, and the control precision to be higher. When the control frequency is high enough, the control performed by the sliding mode controller on the flight platform can be considered as continuous control.
[0108] In some embodiments, the target control law includes a first control law and a second control law, and the second control law corresponds to a continuous function. When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter characterized by the target sliding surface (i.e., the absolute value of the sliding variable) is greater than the second threshold value, the flight platform is controlled based on the first control law and the control gain parameter. When it is detected that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter characterized by the target sliding surface (i.e., the absolute value of the sliding variable) is less than or equal to the second threshold value, the flight platform is continuously controlled based on the second control law and the control gain parameter.
[0109] The flight control system can use different control laws outside and inside the boundary layer, respectively. For example, the flight control system can use a first control law (e.g., a switching control law) when the sliding variable is outside the boundary layer, and use a second control law (e.g., a continuous control law) when the sliding variable is inside the boundary layer.
[0110] As an example, the target control law u(t) can be represented by the following formula:
[0111]
[0112] u(t) = -K - sat(φS(x)) (second control law)
[0113] wherein sat() is a saturation function, K is a control gain parameter, φ represents a second threshold value, and S(x) is a sliding mode variable. In this specification, S(x) represents a sliding mode variable. When |S(x)| is greater than φ (i.e., the sliding mode variable is outside the boundary layer), the flight control system controls the flight platform based on the first control law and the control gain parameter. When |S(x)| is less than or equal to φ (i.e., the sliding mode variable is inside the boundary layer), the flight platform is continuously controlled based on the second control law and the control gain parameter. In this case, the control input is saturated, thereby avoiding the occurrence of chattering.
[0114] In summary, the present specification provides a flight control method system and a storage medium. The flight control system can obtain target flight control parameters (which can include sliding mode control parameters such as sliding surface parameters, control laws, etc.) of the flight platform based on the flight state parameters and the target flight mode of the flight platform through a pre-trained reinforcement learning agent. Then, the flight platform is controlled according to the target flight control parameters, so that the flight state parameters of the flight platform tend to the target flight state parameters indicated by the target flight control parameters. The control system can obtain the target flight control parameters of the flight platform through the reinforcement learning agent according to the flight state parameters and the target flight mode of the flight platform. The flight control parameters do not need to be set manually. Moreover, the target flight control parameters obtained through the reinforcement learning agent are control strategies that are effective under the current flight state and the current flight mode. The target flight control parameters can be dynamically updated as the flight state or the flight mode changes, thereby ensuring the stability and maneuverability of the flight platform.
[0115] Another aspect of the present specification provides a computer-readable non-transitory storage medium storing at least one set of instructions for performing flight control. When the at least one set of instructions is executed by a processor, the at least one set of instructions directs the processor to implement steps of the flight control method described in the present specification. In some possible implementations, various aspects of the present specification can also be implemented as a program product in the form of a computer readable medium having program code portions stored therein. When the program product is run on the computing system 200, the program code portions cause the computing system 200 to perform steps of the flight control method described in the present specification. The program product for implementing the above-described method can include the program code portions in a portable compact disc read-only memory (CD-ROM) and can be run on the computing system 200. However, the program product of the present specification is not limited to this, and in the present specification, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system. The program product can take any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer readable storage medium can include a data signal carried in a baseband or as part of a carrier wave, in which readable program code is borne. Such a propagated data signal can take on many forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The readable storage medium can also be any readable medium that is not a storage medium that can send, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The program code contained in the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, and the like, or any suitable combination of the above. The program code for performing the operations of the present specification can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, and the like, and a conventional procedural programming language such as the "C" programming language or similar programming languages. The program code can be executed entirely on the computing system 200, partially on the computing system 200, as an independent software package, partially on the computing system 200 and partially on a remote computing device, or entirely on a remote computing device.
[0116] The above described embodiments of the disclosure have been described. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.
[0117] In light of the above, those skilled in the art will appreciate that the foregoing detailed description of the present disclosure is susceptible to various modifications and / or revisions without departing from the spirit and scope of the present disclosure. Although specific embodiments have been illustrated and described herein, it should be appreciated that any arrangement which is calculated to achieve the same purpose can be substituted for the specific embodiments shown. This disclosure is intended to cover any adaptations or variations of the present disclosure. Therefore, it is intended that the application be protected by: the broadest interpretation of the appended claims to take into account unforeseen equivalents and alternatives based on current knowledge, or future knowledge.
[0118] In addition, certain terminology has been used to describe embodiments of the disclosure. For example, "one embodiment," "an embodiment," and / or "some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. Therefore, it is understood that the use of "embodiment" or "one embodiment" or "an embodiment" or "some embodiments” in various places throughout the specification is not intended to be interpreted as a referral to the same embodiment, unless otherwise indicated. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0119] It should be understood that in the foregoing description of embodiments of the disclosure, various features are sometimes grouped together in a single embodiment, figure, or description of a figure for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various aspects, embodiments, and / or features. However, this should not be interpreted as a requirement that these features must be provided together in order to form an embodiment of the disclosure. In fact, some embodiments of the disclosure can provide one or more features while omitting others. In addition, the description of the embodiments of the disclosure is intended to cover any adaptations or variations of the specific embodiments discussed in this disclosure. Accordingly, the disclosure is intended to be broad in scope and the claims should be afforded with the broadest interpretation available.
[0120] Each patent, patent application, publication of a patent application, and other material, for example articles, books, specifications, publications, documents, things, and / or the like which can have been cited or referred to in this document, are incorporated by reference into the present document, and are hereby made part of the present document, for all purposes, to the same extent as if each individual publication or item had been individually cited or incorporated by reference. Additionally, where a definition or use of a term in an incorporated reference is inconsistent or contrary to the definition or use of that term in the present document, the definition or use that appears in the present document applies and the opposite definition or use does not apply.
[0121] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the present specification. Other modifications that fall within the scope of the present specification can also be made. Accordingly, the present specification discloses embodiments only as examples. Substantially any arrangement, which is neither specifically described nor explicitly illustrated, can be substituted for the specific embodiments disclosed, without departing from the scope of the present specification. Accordingly, the present specification discloses embodiments only as examples.
Claims
1. A flight control method applied to a flight control system of a flight platform, the method comprising: obtaining a flight state of the flight platform, the flight state comprising a flight state parameter and a target flight mode; obtaining a target flight control parameter of the flight platform according to the flight state of the flight platform by a pre-trained reinforcement learning agent, the target flight control parameter of the flight platform at least comprising a sliding mode control parameter and a control gain parameter corresponding to the target flight mode, the sliding mode control parameter at least being used to represent a target flight state parameter corresponding to the target flight mode of the flight platform, the reinforcement learning agent being obtained by reinforcement learning with the flight state of the flight platform as a state input, the control parameter of the flight platform as an action output, and a reward value of a preset reward function meeting a preset condition as a target; obtaining a target sliding mode surface and a target control law of the flight platform according to the sliding mode control parameter, the target sliding mode surface being used to represent the target flight state parameter corresponding to the target flight mode of the flight platform, and the target control law being used to represent a control logic when the flight platform is controlled; obtaining the target response rate corresponding to the target flight mode, determining a second threshold value according to the target response rate, the second threshold value being negatively correlated with the target response rate; or the target flight control parameter comprising the second threshold value, the second threshold value being determined by the reinforcement learning agent according to the flight state parameter and the target flight mode; and when detecting that an absolute value of a difference between the flight state parameter of the flight platform and a target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value, continuously controlling the flight platform based on the target control law and the control gain parameter, so that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to a first threshold value, the absolute value of the second threshold value being greater than the absolute value of the first threshold value, the control gain parameter being used to represent a target response rate of the flight platform in the target flight mode in response to the target control law, and the control gain parameter being positively correlated with the target response rate.
2. The method of claim 1, wherein, The continuously controlling the flight platform based on the target control law and the control gain parameter when detecting that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value comprises: when detecting that the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value, gradually increasing a control frequency of the target control law until the flight platform is continuously controlled.
3. The method of claim 2, wherein, The target control law comprises a first control law and a second control law, and a corresponding control function of the second control law is a continuous function. the first control law and the control gain parameter when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is greater than a second threshold value; the target control law and the control gain parameter when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value. the second control law and the control gain parameter when the absolute value of the difference between the flight state parameter of the flight platform and the target flight state parameter represented by the target sliding mode surface is less than or equal to the second threshold value.
4. The method of claim 1, wherein, The flight mode of the flight platform includes at least a take-off mode, a cruising mode, a transition flight mode and a landing mode. The target flight mode of the flight platform is obtained by: determining the target flight mode of the flight platform according to the corresponding flight route of the flight platform, the geographical position information of the flight platform and the flight height of the flight platform, the target flight mode being one of the flight modes of the flight platform.
5. The method of claim 1, wherein, The flight state parameter includes at least an attitude parameter, an acceleration parameter, a speed parameter and an energy consumption parameter. The preset reward function includes at least an attitude error reward function, an acceleration error reward function, a speed error reward function and an energy consumption reward function, wherein the reward values of the attitude error reward function, the acceleration error reward function and the speed error reward function are negatively correlated with errors, and the reward value of the energy consumption reward function is negatively correlated with energy consumption.
6. The method of claim 5, wherein, The reward value of the preset reward function meets a preset condition, including: The reward value of the preset reward function is maximized.
7. A flight control system, comprising: at least one storage medium storing at least one instruction set for flight control; and at least one processor in communication connection with the at least one storage medium, wherein when the flight control system is running, the at least one processor reads the at least one instruction set, and executes the flight control method according to the indication of the at least one instruction set.
8. A computer-readable non-transitory storage medium, wherein, The computer readable non-volatile storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to realize the flight control method according to any one of claims 1-6.