Helicopter flight control method, system and equipment

By optimizing the training of the imitation learning network, the problem of relying on manual adjustments in the traditional helicopter flight control system has been solved, achieving more efficient and stable flight control and improving the safety and handling performance of the helicopter.

CN120802613AActive Publication Date: 2025-10-17NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510893800.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Traditional helicopter flight control systems rely on manual adjustments, which are inefficient and prone to human error in complex environments, affecting flight safety and stability. Existing intelligent control methods have not yet been able to fully improve performance.

Method used

The model is trained using an imitation learning network, combined with a behavior cloning algorithm and an adversarial network. The model is then optimized to output the control actions of the helicopter, and flight control is performed through the optimized model.

Benefits of technology

It improves the flight safety and stability of helicopters in dynamic environments and enhances the accuracy and reliability of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802613A_ABST
    Figure CN120802613A_ABST
Patent Text Reader

Abstract

The invention discloses a flight control method, system and device for a helicopter, and relates to the technical field of helicopter flight control, and the method comprises the steps: inputting the flight state information of the helicopter into an optimized imitation learning network, and outputting the control action of the helicopter to control the flight of the helicopter; the training process of the optimized imitation learning network comprises the following steps of: training an initial imitation learning network by adopting a behavior cloning algorithm and utilizing an expert demonstration data set to obtain a first imitation learning network; performing joint adversarial training on the initial discrimination network and the first imitation learning network by using an expert demonstration data set by adopting an adversarial network training method to obtain a trained discrimination network and a second imitation learning network; optimizing the second imitation learning network based on the first data set and the trained discrimination network to obtain a third imitation learning network; and based on the second data set, optimizing the third imitation learning network to obtain an optimized imitation learning network. The flight safety and stability of the helicopter are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of helicopter flight control, in particular to a flight control method, system and device of a helicopter. BACKGROUND

[0002] With the development of helicopter related technologies, helicopters are increasingly widely used in various fields. Unlike other controlled objects of control systems, a helicopter is a complex controlled object with strong coupling, nonlinearity, multiple inputs and multiple outputs, under-actuation and other characteristics, and has many complex flight modes, including forward flight, backward flight, side flight, hovering and vertical take-off and landing, which requires the control system of the helicopter to have better performance and be able to flexibly adjust and switch according to different modes to ensure the stability and maneuverability of the helicopter in various flight states.

[0003] As the core component of the helicopter to realize autonomous flight function, the traditional helicopter control technology still needs a lot of human involvement. In a high dynamic environment, the adaptability of the model is low, and the controller parameters need to be adjusted manually to adapt to the changes of the flight state. This way of relying on manual adjustment not only is inefficient, but also is prone to human error in complex and variable flight environments, affecting the safety and stability of the helicopter flight. In order to deal with the problem of flight autonomous control, researchers have proposed various intelligent control methods, such as neural network control, deep learning control, etc. These methods enable the helicopter to maintain stable control in dynamic and unstructured scenarios, but how to further improve the performance of the helicopter flight controller remains to be studied.

[0004] Therefore, it is necessary to provide a flight control method of a helicopter to solve the above problems. SUMMARY

[0005] The purpose of the present application is to provide a flight control method, system and device of a helicopter to improve the safety and stability of the helicopter flight.

[0006] To achieve the above purpose, the present application provides the following solutions:

[0007] In a first aspect, the present application provides a flight control method of a helicopter, the flight control method of the helicopter comprising:

[0008] obtaining flight state information of the helicopter at a current time;

[0009] inputting the flight state information of the helicopter at the current time into an optimized imitation learning network, outputting a control action of the helicopter at the current time, and controlling the flight of the helicopter by using the control action of the helicopter at the current time; wherein the training process of the optimized imitation learning network comprises:

[0010] The initial imitation learning network is trained by using an expert demonstration data set to obtain a first imitation learning network by using a behavior cloning algorithm; the expert demonstration data set includes a plurality of original flight state information-control action pairs;

[0011] The initial discrimination network and the first imitation learning network are jointly trained by using an expert demonstration data set by using an adversarial network training method to obtain a trained discrimination network and a second imitation learning network;

[0012] A first data set is constructed, and the second imitation learning network is optimized based on the first data set and the trained discrimination network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set;

[0013] A second data set is constructed, and the third imitation learning network is optimized based on the second data set to obtain an optimized imitation learning network; the second data set is constructed based on a new flight state information-score label pair and the expert demonstration data set; the new flight state information-score label pair includes: new flight state information and a score label, the new flight state information is the flight state information of the helicopter at the time when the control action output by the third imitation learning network is executed, and the difference between the new flight state information and the original flight state information is greater than a preset threshold, and the score label is the score of the control action output by the third imitation learning network.

[0014] Optionally, the flight state information includes: flight state, special situation information, and virtual control signal.

[0015] Optionally, the initial imitation learning network is trained by using an expert demonstration data set to obtain a first imitation learning network by using a behavior cloning algorithm, specifically including:

[0016] The expert demonstration data set is constructed;

[0017] The initial imitation learning network is trained by using a behavior cloning algorithm, taking the flight state information of the helicopter at the current time of the sample as input and taking the control action of the helicopter at the current time of the sample as output, until the loss function reaches a minimum value or the training round reaches a maximum value, the training is stopped, and the first imitation learning network is obtained.

[0018] Optionally, the expert demonstration data set is constructed, specifically including:

[0019] The control action at the current time of the sample is obtained based on the flight state information of the helicopter at the current time of the sample by using an expert strategy;

[0020] The control action at the current time of the sample is applied to the helicopter to obtain the flight state information of the helicopter at the next time of the sample and the control action at the next time of the sample.

[0021] Based on the flight state information of the plurality of helicopters at the sample next moment and the control action at the sample next moment, a state-action sequence of the helicopter is obtained, and the state-action sequence of the helicopter is taken as expert demonstration data;

[0022] The state-action sequence in the expert demonstration data is split into a plurality of original flight state information-control action pairs, and an expert demonstration data set is obtained.

[0023] Optionally, an adversarial network training method is adopted, and the initial discriminator network and the first imitation learning network are jointly adversarially trained by using the expert demonstration data set to obtain the trained discriminator network and the second imitation learning network, specifically including:

[0024] The initial discriminator network is pre-trained by using the expert demonstration data set to obtain the pre-trained discriminator network;

[0025] The expert demonstration data is input into the first imitation learning network to obtain generated data;

[0026] The generated data and the expert demonstration data are respectively input into the pre-trained discriminator network to determine the discrimination error of the pre-trained discriminator network on the generated data;

[0027] Based on the discrimination error of the pre-trained discriminator network on the generated data, the update gradient of the first imitation learning network and the update gradient of the pre-trained discriminator network are respectively calculated;

[0028] The weight of the first imitation learning network is updated and iterated based on the update gradient of the first imitation learning network until the update gradient of the first imitation learning network reaches the corresponding preset threshold, and the updating and iteration is stopped to obtain the trained discriminator network;

[0029] The weight of the pre-trained discriminator network is updated and iterated based on the update gradient of the pre-trained discriminator network until the update gradient of the pre-trained discriminator network reaches the corresponding preset threshold, and the updating and iteration is stopped to obtain the second imitation learning network.

[0030] Optionally, a first data set is constructed, specifically including:

[0031] Noise is added to the input set in the expert demonstration data set to obtain a first input set;

[0032] Based on the first input set and the second imitation learning network, a first output set is obtained;

[0033] The first data set is constructed by using the first input set and the first output set.

[0034] Optionally, based on the first data set and the trained discriminant network, the second imitation learning network is optimized to obtain a third imitation learning network, specifically comprising:

[0035] The first data set is input into the trained discriminant network to obtain a discriminant error of the trained discriminant network for the first data set;

[0036] Based on the discriminant error of the trained discriminant network for the first data set, an update gradient of the second imitation learning network is calculated;

[0037] The update gradient of the second imitation learning network is used to update and iterate the weight of the second imitation learning network until the update gradient of the second imitation learning network reaches a corresponding preset threshold, and the update iteration is stopped to obtain the third imitation learning network.

[0038] Optionally, based on the second data set, the third imitation learning network is optimized to obtain an optimized imitation learning network, specifically comprising:

[0039] Based on the score label and the expert demonstration data in the second data set, an update gradient of the third imitation learning network is calculated;

[0040] Based on the second data set, the update gradient of the third imitation learning network is iterated until the update gradient of the third imitation learning network reaches a corresponding preset threshold, and the update iteration is stopped to obtain the optimized imitation learning network.

[0041] In a second aspect, the application provides a flight control system of a helicopter, which is used to implement the flight control method of the helicopter, and comprises:

[0042] A data acquisition unit is configured to acquire flight state information of the helicopter at a current time;

[0043] A control action determination and flight control unit is configured to input the flight state information of the helicopter at the current time into the optimized imitation learning network, output a control action of the helicopter at the current time, and control the flight of the helicopter by using the control action of the helicopter at the current time; wherein the training process of the optimized imitation learning network comprises:

[0044] An action cloning algorithm is used to train an initial imitation learning network by using an expert demonstration data set to obtain a first imitation learning network; the expert demonstration data set comprises a plurality of original flight state information-control action pairs;

[0045] An adversarial network training method is used to jointly train an initial discriminant network and the first imitation learning network by using the expert demonstration data set to obtain a trained discriminant network and a second imitation learning network;

[0046] construct a first data set, and based on the first data set and the trained discriminant network, optimize the second imitation learning network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set;

[0047] construct a second data set, and based on the second data set, optimize the third imitation learning network to obtain an optimized imitation learning network; the second data set is constructed based on a new flight state information-score label pair and the expert demonstration data set; the new flight state information-score label pair includes new flight state information and a score label, the new flight state information is flight state information of the helicopter at a time after the helicopter executes the control action output by the third imitation learning network, and a difference between the new flight state information and the original flight state information is greater than a preset threshold, and the score label is a score of the control action output by the third imitation learning network.

[0048] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the flight control method of the helicopter.

[0049] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0050] The present application discloses a flight control method, system and device of a helicopter, an optimized imitation learning network is obtained by training, the optimized imitation learning network is used to output a control action of the helicopter at a current time, and the helicopter is controlled by using the control action of the helicopter at the current time. Wherein, first, a behavior cloning algorithm is used to imitate an excellent control strategy of an expert to obtain a first imitation learning network; then, the fitting ability of the first imitation learning network is further optimized by using the idea of joint adversarial training to obtain a second imitation learning network; then, the generalization and the robustness of the second imitation learning network are improved by generalization training, and finally, a high-performance intelligent enhanced controller (i.e. the optimized imitation learning network) that can adapt to dynamic environment changes is obtained, thereby improving the safety and stability of the helicopter flight. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0052] Figure 1A flight control method flowchart of a helicopter is provided for an embodiment of the present application.

[0053] Figure 2 A construction process schematic diagram of an expert demonstration data set is provided for an embodiment of the present application.

[0054] Figure 3 A process schematic diagram of training an initial imitation learning network using a behavior cloning algorithm is provided for an embodiment of the present application.

[0055] Figure 4 A guided learning training process schematic diagram is provided for an embodiment of the present application.

[0056] Figure 5 A training process schematic diagram of a second imitation learning network is provided for an embodiment of the present application.

[0057] Figure 6 A performance generalization training process schematic diagram of a third imitation learning network is provided for an embodiment of the present application.

[0058] Figure 7 A structural schematic diagram of a computer device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0060] In the research of intelligent flight control of a helicopter, how to train a high-performance controller is an important issue. Therefore, the present application introduces an intelligent enhanced control method, which continuously trains an imitation learning network using excellent flight data, and uses the optimized imitation learning network as the core algorithm of the intelligent enhanced controller. The optimized imitation learning network is the core implementation module of the intelligent enhanced controller, and the intelligent enhanced controller is the application carrier of the optimized imitation learning network. The two are interdependent and mutually promote each other, and together improve the flight performance and safety of the helicopter, and realize the coordination of safety and control performance of the helicopter. The intelligent enhanced flight control is the imitation of the pilot and the high-performance flight control instruction, which can improve the accuracy and reliability of the helicopter control system in dynamic environment, and the stability and safety of the flight. The intelligent enhanced control design and research of the helicopter have important theoretical significance and engineering value.

[0061] In order to make the above objectives, characteristics and advantages of the present application more apparent, further specific embodiments will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] In one exemplary embodiment, as shown in Figure 1 a flight control method of a helicopter is provided, comprising the following steps. Wherein:

[0063] Step S1, obtaining flight state information of the helicopter at the current time.

[0064] As an optional implementation, the flight state information includes: flight state, special situation information and virtual control signal.

[0065] Step S2, inputting the flight state information of the helicopter at the current time into the optimized imitation learning network, outputting the control action of the helicopter at the current time, and using the control action of the helicopter at the current time to control the flight of the helicopter; wherein the training process of the optimized imitation learning network comprises the following steps:

[0066] Step S21, using a behavior cloning algorithm to train an initial imitation learning network using an expert demonstration data set to obtain a first imitation learning network; the expert demonstration data set includes a plurality of original flight state information-control action pairs.

[0067] Specifically, the expert behavior is copied and imitated through teaching learning, and the initial imitation learning network is trained using a behavior cloning algorithm in learning to fit the control strategy under the expert behavior.

[0068] As an optional implementation, step S21 specifically includes:

[0069] Step S211, constructing an expert demonstration data set.

[0070] As an optional implementation, step S211 specifically includes:

[0071] Step S2111, using an expert strategy to obtain a control action at a sample current time based on flight state information of the helicopter at the sample current time.

[0072] Step S2112, applying the control action at the sample current time to the helicopter to obtain flight state information of the helicopter at a sample next time and a control action at the sample next time.

[0073] Step S213, obtaining a state-action sequence of the helicopter based on the flight state information of the helicopter at the sample next time and the control action at the sample next time, and taking the state-action sequence of the helicopter as expert demonstration data.

[0074] In step S2114 , the state-action sequence in the expert demonstration data is split into multiple original flight state information-control action pairs to obtain an expert demonstration data set.

[0075] Specifically, in order to learn the control experience of a helicopter, it is necessary to obtain demonstration data under expert strategy control, where the demonstration data can come from human experts or excellent pilots. Figure 2 As shown in the figure, the construction principle of the expert demonstration dataset is as follows: the sample flight status information s of the helicopter is n Input to the expert strategy π * , including helicopter flight status, special information, etc., output sample control action a based on the sample flight status information of the helicopter n , after acting on the environment, new flight status information s is obtained n+1 , and output control action a n+1 , so that the state action sequence τ can be obtained by repeating it, that is, the expert demonstration data; then, the expert demonstration data is split into multiple original flight state information-control action pairs, and the expert demonstration data set D = {(s1, a1), (s2, a2), …, (s n ,a n )}, n represents the sequence number of flight status information or control action, n∈(1,N).

[0076] In step S212, a behavioral cloning algorithm is used to train the initial imitation learning network with the flight state information of the helicopter at the current moment of the sample as input and the control action of the helicopter at the current moment of the sample as output until the loss function reaches the minimum value or the number of training rounds reaches the maximum value. The training is stopped to obtain the first imitation learning network.

[0077] Specifically, based on the expert demonstration dataset D, the behavior cloning algorithm is used to transform the helicopter’s flight status information s at the current moment into t ( Figure 3 The control action a of the helicopter at the current moment is taken as the feature as the input of the initial imitation learning network. t ( Figure 3 (denoted as action in the text) is regarded as a label as the target output of the initial imitation learning network for fitting training. The process of the behavior cloning algorithm is as follows Figure 3 As shown. In the learning process, the expert demonstration data set D( Figure 3 The dataset is divided into a training set (70%) and a validation set (30%). The training objectives are as follows:

[0078]

[0079] in, is the parameter of the first imitation learning network, E is the expectation; θ is the parameter of the initial imitation learning network, π(st ,θ) is the actual output of the initial imitation learning network, and the loss function loss is represented as:

[0080]

[0081] wherein a t is the predicted value of the control action at the t time, a′ n is the true value of the control action at the t time.

[0082] An optimal strategy model π θ is obtained using a regression learning method, and finally a state-to-action simulation mapping relationship based on a neural network is obtained:

[0083] a t = π(s t ,θ) (3)

[0084] After the error of the verification set converges, the training is stopped, then the trained imitation learning network (i.e. the first imitation learning network) is tested in the environment, the current state (i.e. the flight state information) s t is obtained from the environment in real time, and the action (i.e. the control action) a t at the current time is calculated and applied to the environment to check the training effect.

[0085] In step S22, an adversarial network training method is used to jointly adversarially train the initial discriminant network and the first imitation learning network using the expert demonstration data set, to obtain a trained discriminant network and a second imitation learning network.

[0086] As an optional implementation, step S22 specifically includes:

[0087] In step S221, the initial discriminant network is pre-trained using the expert demonstration data set to obtain a pre-trained discriminant network.

[0088] In step S222, the expert demonstration data is input into the first imitation learning network to obtain generated data.

[0089] In step S223, the generated data and the expert demonstration data are respectively input into the pre-trained discriminant network to determine the discriminant error of the pre-trained discriminant network on the generated data.

[0090] In step S224, based on the discriminant error of the pre-trained discriminant network on the generated data, the update gradient of the first imitation learning network and the update gradient of the pre-trained discriminant network are respectively calculated.

[0091] Step S225, the weight of the first imitation learning network is updated based on the update gradient of the first imitation learning network, and the update iteration is stopped until the update gradient of the first imitation learning network reaches the corresponding preset threshold, and a trained discriminant network is obtained.

[0092] Step S226, the weight of the pre-trained discriminant network is updated based on the update gradient of the pre-trained discriminant network, and the update iteration is stopped until the update gradient of the pre-trained discriminant network reaches the corresponding preset threshold, and a second imitation learning network is obtained.

[0093] Specifically, as shown in the figure, Figure 4 Based on the teaching learning, the discriminant network is introduced to develop guided learning, and the performance of the discriminant network and the first imitation learning network is improved simultaneously by developing joint adversarial training of the discriminant network and the first imitation learning network, and the first imitation learning network is guided to optimize again. The goal of the discriminant network in the training process is to successfully judge whether the flight state information-control action pair is from the expert demonstration data set or the first imitation learning network, and the goal of the first imitation learning network is to make the discriminant network judge the data generated by the first imitation learning network as coming from the expert demonstration data set, so that the data generated by the first imitation learning network is close enough to the real expert demonstration data.

[0094] The judgment result of the discriminant network for the generated data is represented as The judgment result of the pre-trained discriminant network for the expert demonstration data x t is represented as p t (x t ,w p ), wherein the generated data is x t =(s t ,a t ), w p is the weight of the pre-trained discriminant network, is the control action generated by the first imitation learning network. In the pre-training phase, that is, the judgment result of the generated data in the subsequent phase is closer to 1, the generated data is closer to the expert demonstration data x t , is closer to 0, and the corresponding input data is farther away from the expert demonstration data. The output range of the pre-trained discriminant network is 0≤p t (x t ,w p ), The discriminant error of the pre-trained discriminant network for the generated data is represented as:

[0095]

[0096] During the learning process, the update targets of the pre-trained discriminant network and the first imitation learning network are e t = 0, e t = 1.

[0097] Based on the discrimination error of the pre-trained discriminant network on the generated data, the update gradient Aw i of the first imitation learning network and the update gradient Aw p of the pre-trained discriminant network are calculated.

[0098]

[0099]

[0100] where w i is the weight of the first imitation learning network, and a1(t), a2(t) > 0 are network parameter update steps. Considering that the error between the actual value and the target value is large at the beginning of network training, a large step can quickly reduce the error. As the training progresses, the error gradually decreases. In order to avoid the vibration phenomenon caused by overfitting problem in the later training, the update step should be weakened as the discrimination error decreases. The update steps are represented as:

[0101] a1(t) = exp(-y1e t - e1)(7)

[0102] a2(t) = exp(-y2e t - e2)(8)

[0103] where a1(t) is the first update step at the t-th time, a2(t) is the second update step at the t-th time, y1, y2, 1, 2 are constants, and y1, y2, 1,2 > 0.

[0104] Step S23, constructing a first data set, and based on the first data set and the trained discriminant network, optimizing the second imitation learning network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set.

[0105] As an optional implementation, in step S23, the first data set is constructed, specifically including:

[0106] Step S231, adding noise to the input set in the expert demonstration data set to obtain a first input set.

[0107] Step S232, based on the first input set and the second imitation learning network, obtaining a first output set.

[0108] Step S233, a first data set is constructed using the first input set and the first output set.

[0109] Specifically, as shown in Figure 5 In order to improve the generalization of the second imitation learning network, based on the results of guided learning, noise is added to the input set in the expert demonstration data set to construct a new input set (i.e. the first input set), see equation (9), and data generalization training is performed. The trained discriminant network is used to optimize the performance of the second imitation learning network on the new data set (i.e. the first data set), and the generalization of the third imitation learning network is improved.

[0110]

[0111] wherein, is the first input set, and σ is a multidimensional Gaussian distribution with a zero vector as the mean.

[0112] The second imitation learning network calculates the corresponding new output set based on the new input set as Then, a new first data set is constructed based on the new input set and the new output set

[0113] As an optional implementation, in step S23, based on the first data set and the trained discriminant network, the second imitation learning network is optimized to obtain the third imitation learning network, which specifically includes:

[0114] Step S234, input the first data set to the trained discriminant network to obtain the discrimination error of the trained discriminant network for the first data set.

[0115] Step S235, based on the discrimination error of the trained discriminant network for the first data set, the update gradient of the second imitation learning network is calculated.

[0116] Step S236, the weight of the second imitation learning network is updated and iterated using the update gradient of the second imitation learning network, until the update gradient of the second imitation learning network reaches the corresponding preset threshold, the update iteration is stopped, and the third imitation learning network is obtained.

[0117] Specifically, the discrimination error of the first data set is calculated by the trained discriminant network as

[0118] The update gradient Aw of the second imitation learning network is calculated based on the discrimination error of the trained discriminant network for the first data set j and update, the update gradient Aw of the second imitation learning network j The calculation result is: ​

[0119]

[0120] Similar to formula (7), wherein It is also a variable step size update method.

[0121] In step S24, a second data set is constructed, and a third imitation learning network is optimized based on the second data set to obtain an optimized imitation learning network; the second data set is constructed based on a new flight state information-score label pair and an expert demonstration data set; the new flight state information-score label pair includes new flight state information and a score label, the new flight state information is flight state information at a time when the helicopter executes a control action output by the third imitation learning network, and a difference between the new flight state information and the original flight state information is greater than a preset threshold, and the score label is a score of the control action output by the third imitation learning network.

[0122] As an optional implementation, in step S24, the third imitation learning network is optimized based on the second data set to obtain the optimized imitation learning network, and specifically includes:

[0123] In step S241, an update gradient of the third imitation learning network is calculated based on the score label in the second data set and the expert demonstration data.

[0124] In step S242, the update gradient of the third imitation learning network is iterated based on the second data set until the update gradient of the third imitation learning network reaches a corresponding preset threshold, the update iteration is stopped, and the optimized imitation learning network is obtained.

[0125] Specifically, the performance generalization training is to use actual flight trajectory data to ensure the effectiveness of learning from the expert demonstration data, and the essence of the process is to correct the ground learning result by using actual flight trajectory data, and in the correction process, the effective rules in the expert demonstration data are extracted by using the increasing data, and the interference in the expert demonstration data is eliminated. Based on the result of the data generalization training, the actual flight data is used to perform performance generalization training on the controller. In each round of training, a score label is added to the new flight state information according to the performance evaluation function and is put into the expert demonstration data set, and the third imitation learning network is trained by using the new and old aggregated data to optimize the robustness of the controller.

[0126] The performance generalization training process of the third imitation learning network is as shown in Figure 6 . Wherein, the label data is to add a score label y t ,a t ,s t+1 ) in the flight state information-control action pair and the new flight state information (s t+1The score label is calculated by a manually designed scoring rule, which mainly considers the tracking accuracy and safety of the flight trajectory, and the partial derivative of the value with respect to the state of the helicopter is continuous.

[0127] In each round of training, when the third imitation learning network is used to control the flight of the helicopter, if a state that is not in the expert demonstration data set is encountered, the state (s t ,a t ,s t+1 ) is collected, a control score label y t+1 is given, and the state (s t ,a t ,s t+1 ,y t+1 ) is put into the second data set D i . The third imitation learning network is retrained, and the third imitation learning network is trained using multiple sets of new and old aggregated data (i.e., the second data set) to optimize the robustness of the controller, and then the next round of training is entered. For the labeled data (s t ,a t ,s t+1 ,y t+1 ), the calculation result of the update gradient of the corresponding third imitation learning network is represented as:

[0128]

[0129] wherein a represents an update step size, represents the partial derivative of the score with respect to the flight state information. In this way, more flight trajectory data can be collected during the training of the third imitation learning network, so that the expert demonstration database is more perfect, and the actual control performance of the controller will also show a trend of robustness improvement as the data increases.

[0130] Before the first round of training, the performance generalization training result is defined as the initialization strategy, denoted as Then, the helicopter flight trajectory data under the strategy is collected and added to the data set D, which is used to train the strategy In order to fully utilize the expert strategy, the data aggregation queries the expert strategy during the learning phase, and updates the strategy π i for collecting flight trajectory data. i In the i-th iteration, the update rule of the strategy π

[0131]

[0132] wherein β i represents the weight of the initialization strategy, which is a constant.

[0133] The above formula shows that the learning strategy In the triggered state, part of the expert demonstration is required. Data aggregation can alleviate the performance difference caused by the difference between the state triggered by the learning strategy and the state distribution in the initial data set by collecting the control behavior of the learning strategy in the state.

[0134] Based on the same inventive concept, the embodiments of the present application also provide a flight control system of a helicopter for implementing the flight control method of the helicopter as described above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more flight control system embodiments of the helicopter provided below can refer to the limitations of the flight control method of the helicopter described above, and will not be repeated here.

[0135] In one exemplary embodiment, a flight control system of a helicopter is provided, comprising:

[0136] A data acquisition unit is configured to acquire flight state information of the helicopter at a current time.

[0137] A control action determination and flight control unit is configured to input the flight state information of the helicopter at the current time into the optimized imitation learning network, output a control action of the helicopter at the current time, and perform flight control on the helicopter using the control action of the helicopter at the current time; wherein the training process of the optimized imitation learning network comprises:

[0138] An behavior cloning algorithm is used to train the initial imitation learning network using the expert demonstration data set to obtain a first imitation learning network; the expert demonstration data set includes a plurality of original flight state information-control action pairs.

[0139] An adversarial network training method is used to jointly train the initial discriminator network and the first imitation learning network using the expert demonstration data set to obtain a trained discriminator network and a second imitation learning network.

[0140] A first data set is constructed, and the second imitation learning network is optimized based on the first data set and the trained discriminator network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set.

[0141] The second data set is constructed based on the second data set, and the third imitation learning network is optimized based on the second data set to obtain an optimized imitation learning network; the second data set is constructed based on a new flight state information-score label pair and an expert demonstration data set; the new flight state information-score label pair includes: new flight state information and a score label, the new flight state information is flight state information at a time when the helicopter executes the control action output by the third imitation learning network, and a difference between the new flight state information and the original flight state information is greater than a preset threshold, and the score label is a score of the control action output by the third imitation learning network.

[0142] In an exemplary embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the flight control method of the helicopter.

[0143] In an exemplary embodiment, a computer device is provided, which can be a server or a terminal, and an internal structure diagram thereof can be as shown in Figure 7 The computer device comprises a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a flight control method of a helicopter.

[0144] Those skilled in the art can understand that Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0145] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0146] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0147] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0148] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0149] The principles and implementations of the present application are described in the specific examples in this article, and the above examples are only used to help understand the method, system and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A helicopter flight control method, characterized in that: The flight control method of the helicopter comprises: Get the current flight status information of the helicopter; The flight state information of the helicopter at the current moment is input into the optimized imitation learning network, the control action of the helicopter at the current moment is output, and the flight of the helicopter is controlled by using the control action of the helicopter at the current moment. The training process of the optimized imitation learning network includes: The behavioral cloning algorithm is used to train the initial imitation learning network using an expert demonstration dataset to obtain the first imitation learning network; the expert demonstration dataset includes multiple original flight state information-control action pairs; Adopting the adversarial network training method, the initial discriminant network and the first imitation learning network are jointly trained adversarially using the expert demonstration dataset to obtain the trained discriminant network and the second imitation learning network. Constructing a first data set, and optimizing the second imitation learning network based on the first data set and the trained discriminant network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set; A second data set is constructed, and based on the second data set, the third imitation learning network is optimized to obtain an optimized imitation learning network; the second data set is constructed based on the new flight state information-scoring label pair and the expert demonstration data set; the new flight state information-scoring label pair includes: new flight state information and a scoring label, the new flight state information is the flight state information of the helicopter at the moment after the control action output by the third imitation learning network is executed, and the difference between the new flight state information and the original flight state information is greater than a preset threshold, and the scoring label is the score of the control action output by the third imitation learning network.

2. The helicopter flight control method according to claim 1, characterized in that: Flight status information includes: flight status, special situation information and virtual control signals.

3. The helicopter flight control method according to claim 1, characterized in that: The behavioral cloning algorithm is used to train the initial imitation learning network using the expert demonstration dataset to obtain the first imitation learning network, which specifically includes: Construct expert demonstration dataset; The behavioral cloning algorithm is used to take the flight state information of the helicopter at the current moment of the sample as input and the control action of the helicopter at the current moment of the sample as output. The initial imitation learning network is trained until the loss function reaches the minimum value or the training round reaches the maximum value. The training is stopped to obtain the first imitation learning network.

4. The helicopter flight control method according to claim 3, characterized in that: Construct an expert demonstration dataset, including: Using the expert strategy, based on the helicopter's flight status information at the current moment of the sample, the control action of the sample at the current moment is obtained; Apply the control action of the sample at the current moment to the helicopter to obtain the flight state information of the helicopter at the next moment and the control action of the sample at the next moment; Based on the flight state information and control actions of multiple helicopters at the next moment of the sample, the state action sequence of the helicopters is obtained, and the state action sequence of the helicopters is used as expert demonstration data; The state-action sequence in the expert demonstration data is split into multiple original flight state information-control action pairs to obtain the expert demonstration dataset.

5. The helicopter flight control method according to claim 4, characterized in that: Adopting the adversarial network training method, the initial discriminant network and the first imitation learning network are jointly trained adversarially using the expert demonstration dataset to obtain the trained discriminant network and the second imitation learning network. Specifically, the following steps are performed: Use the expert demonstration data set to pre-train the initial discriminant network to obtain a pre-trained discriminant network; Input the expert demonstration data into the first imitation learning network to obtain generated data; The generated data and expert demonstration data are respectively input to the pre-trained discriminant network to determine the discrimination error of the pre-trained discriminant network on the generated data; Based on the discrimination error of the pre-trained discriminant network on the generated data, the updated gradient of the first imitation learning network and the updated gradient of the pre-trained discriminant network are calculated respectively; The weights of the first imitation learning network are updated and iterated based on the updated gradient of the first imitation learning network until the updated gradient of the first imitation learning network reaches a corresponding preset threshold, and the update iteration is stopped to obtain a trained discriminant network; The weights of the pre-trained discriminant network are updated and iterated based on the updated gradient of the pre-trained discriminant network until the updated gradient of the pre-trained discriminant network reaches the corresponding preset threshold, and the update iteration is stopped to obtain the second imitation learning network.

6. The helicopter flight control method according to claim 5, characterized in that: Construct the first data set, specifically including: Add noise to the input set in the expert demonstration data set to obtain a first input set; Based on the first input set and the second imitation learning network, obtaining a first output set; A first data set is constructed using the first input set and the first output set.

7. The helicopter flight control method according to claim 6, characterized in that: Based on the first data set and the trained discriminant network, the second imitation learning network is optimized to obtain a third imitation learning network, which specifically includes: Inputting the first data set into the trained discriminant network to obtain the discrimination error of the trained discriminant network for the first data set; Calculate the update gradient of the second imitation learning network based on the discrimination error of the trained discriminant network for the first data set; The weights of the second imitation learning network are updated and iterated using the updated gradient of the second imitation learning network until the updated gradient of the second imitation learning network reaches a corresponding preset threshold, and the updating iteration is stopped to obtain a third imitation learning network.

8. The helicopter flight control method according to claim 7, characterized in that: Based on the second data set, the third imitation learning network is optimized to obtain an optimized imitation learning network, which specifically includes: Calculate the update gradient of the third imitation learning network based on the score labels and expert demonstration data in the second dataset; Based on the second data set, the update gradient of the third imitation learning network is iterated until the update gradient of the third imitation learning network reaches a corresponding preset threshold, and the update iteration is stopped to obtain an optimized imitation learning network.

9. A flight control system for a helicopter, characterized in that: The helicopter flight control system is used to implement the helicopter flight control method according to any one of claims 1 to 8, and the helicopter flight control system includes: A data acquisition unit is used to obtain the flight status information of the helicopter at the current moment; The control action determination and flight control unit is used to input the current flight state information of the helicopter into the optimized imitation learning network, output the current control action of the helicopter, and use the current control action of the helicopter to control the flight of the helicopter. The training process of the optimized imitation learning network includes: The behavioral cloning algorithm is used to train the initial imitation learning network using an expert demonstration dataset to obtain the first imitation learning network; the expert demonstration dataset includes multiple original flight state information-control action pairs; Adopting the adversarial network training method, the initial discriminant network and the first imitation learning network are jointly trained adversarially using the expert demonstration dataset to obtain the trained discriminant network and the second imitation learning network. Constructing a first data set, and optimizing the second imitation learning network based on the first data set and the trained discriminant network to obtain a third imitation learning network; the first data set is obtained by adding noise to the expert demonstration data set; A second data set is constructed, and based on the second data set, the third imitation learning network is optimized to obtain an optimized imitation learning network; the second data set is constructed based on the new flight state information-scoring label pair and the expert demonstration data set; the new flight state information-scoring label pair includes: new flight state information and a scoring label, the new flight state information is the flight state information of the helicopter at the moment after the control action output by the third imitation learning network is executed, and the difference between the new flight state information and the original flight state information is greater than a preset threshold, and the scoring label is the score of the control action output by the third imitation learning network.

10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the flight control method for a helicopter according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Defense method oriented to deep reinforcement learning model confrontation attack

    CN110968866A

  • Distance Estimation Using Machine Learning

    US20200074674A1

  • Quantum, biological, computer vision, and neural network systems for industrial internet of things

    US20230176550A1