Formation Flight Control Method and System Based on Cost-Sensitive Meta-Reinforcement Learning

Through the formation flight control method based on cost-sensitive meta-reinforcement learning, the problem that traditional methods are difficult to cope with complex environments and diverse mission needs is solved, and the formation flight is efficient, economical and safe, and has good scalability and generalization capabilities are achieved.

CN119645082BActive Publication Date: 2025-06-10NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162972.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-10
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Traditional formation flight control methods are difficult to cope with complex and changing flight environments and diverse and detailed mission needs, especially when uncertain factors such as airflow disorders, electromagnetic interference and sudden obstacles occur frequently.

Method used

The formation flight control method based on cost-sensitive meta-reinforcement learning is adopted, and the final meta-strategy parameters are obtained by establishing a three-dimensional kinematic model of formation flight, constructing a control model, and conducting meta-reinforcement learning training to achieve efficient and economical formation control.

Benefits of technology

It realizes the independent learning and response of aircraft in complex environments, ensures the safety and efficiency of formation flights, and reduces the dependence on communication links, with excellent scalability and cross-mission generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645082B_ABST
    Figure CN119645082B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of formation flight control, and provides a formation flight control method and system based on cost-sensitive meta-reinforcement learning. The method mainly includes the following steps: establishing a three-dimensional kinematic model of formation flight; constructing a set target of the formation flight control method; constructing a formation flight control model according to the three-dimensional kinematic model of formation flight and the set target of the formation flight control method; training the formation flight control model according to cost-sensitive meta-reinforcement learning to obtain final meta-policy parameters. The present invention realizes fast adaptation and safety guarantee in a complex formation control dynamic environment by adopting a hierarchical optimization strategy within meta-reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of formation flight control, and provides a formation flight control method and system based on cost-sensitive meta-reinforcement learning. Background Art

[0002] In today's aerospace field, multi-aircraft formation flight has become a key technology attracting much attention, and is widely applied to many scenarios such as military reconnaissance, disaster relief, and geographical mapping. However, traditional formation flight control methods are increasingly difficult to meet the complex and changeable real-world requirements. On the one hand, the flight environment is becoming increasingly complex and harsh, and uncertain factors such as turbulent airflows, electromagnetic interference, and sudden obstacles frequently appear; on the other hand, the mission requirements are becoming increasingly diverse and refined, requiring the formation to be able to change its formation quickly, flexibly, and precisely. Traditional control means are stretched, and traditional fixed-model-based methods are difficult to cope with such complex situations.

[0003] With the development of artificial intelligence technology, meta-reinforcement learning can be used to accurately solve control schemes that meet multiple constraints such as aircraft performance, obstacle avoidance, and communication, ensuring the stability of the formation. In the face of the need for formation change, reinforcement learning quickly iterates strategies, can quickly adjust and cooperate to respond, and also reduces the dependence on communication links, with excellent scalability, providing an innovative solution to the multi-aircraft formation flight problem. Summary of the Invention

[0004] The present invention aims to at least solve one of the technical problems existing in the related art. For this reason, the present invention provides a formation flight control method and system based on cost-sensitive meta-reinforcement learning, which optimize the energy consumption of rudder surface operation while ensuring the safety of formation flight, and achieve the efficiency and economy of formation control.

[0005] The present invention provides a formation flight control method based on cost-sensitive meta-reinforcement learning, including:

[0006] S1: Establish a three-dimensional kinematic model of formation flight;

[0007] S2: Construct the set goals of the formation flight control method;

[0008] S3: Construct a formation flight control model according to the three-dimensional kinematic model of formation flight and the set goals of the formation flight control method;

[0009] S4: Train the formation flight control model according to cost-sensitive meta-reinforcement learning to obtain the final meta-policy parameters;

[0010] S5: Control the formation flight according to the final meta-policy parameters.

[0011] A formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, step S1 is specifically as follows:

[0012] S11: Define There is 1 leader aircraft and follower aircraft in the formation system of the aircraft, and design the state vectors of each follower aircraft during formation flight:

[0013]

[0014] Among them, is the state vector of the th follower aircraft, is the ordinal number of the follower aircraft, , is the total number of aircraft, is the abscissa of the th follower aircraft in the inertial coordinate system, is the ordinate of the th follower aircraft in the inertial coordinate system, The th follower aircraft's vertical coordinate in the inertial coordinate system, represents the speed of the th follower aircraft, represents the flight inclination angle of the th follower aircraft, represents the flight deviation angle of the th follower aircraft;

[0015] S12: Design the control vectors of each follower aircraft during formation flight:

[0016]

[0017] Among them, is the control vector of the th follower aircraft, is the overload of the th follower aircraft in the axis direction, is the overload of the th follower aircraft in the axis direction, is the overload of the th follower aircraft in the axis direction; is the matrix transpose;

[0018] S13: Construct the three-dimensional kinematic model of the formation flight according to the state vector and the control vector :

[0019]

[0020]

[0021]

[0022] Among them, is the acceleration due to gravity, is the linearization matrix for formation flight control of the aircraft, is the basic model for formation flight control of the aircraft.

[0023] According to a formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, step S2 is specifically as follows:

[0024] For any formation initial condition:

[0025]

[0026] Among them, is the position vector of the leader aircraft, is the th position vector of the follower aircraft, is the expected distance between the leader aircraft and the th follower aircraft, is the velocity vector of the leader aircraft, is the th velocity vector of the follower aircraft, is the transformation matrix from the ballistic coordinate system of the leader aircraft to the inertial reference coordinate system, is the allowable error of the position vector, is the allowable error of the velocity vector.

[0027] According to a formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, step S3 is specifically as follows:

[0028] S31: Define the adjacency matrix of the leader aircraft:

[0029]

[0030] Among them, is the communication state between the th follower aircraft and the leader aircraft. If it is in a communicable state , if it is in a non-communicable state , is a diagonal matrix;

[0031] S32: Design the control command of the th follower aircraft:

[0032]

[0033]

[0034] Among them, defines the sign function, is the first independent variable for defining the sign function, is the second independent variable for defining the sign function, is the sign function, is to take the absolute value, is the ordinal number of the adjacent following aircraft following the aircraft, is the set of adjacent following aircraft of the th following aircraft, is the expected position of the th following aircraft in the leader ballistic coordinate system, represents the position vector of the th following aircraft, is the first feedback gain of the formation flight controller, is the second feedback gain of the formation flight controller, is the first order gain of the formation flight controller, is the second order gain of the formation flight controller; is the modulo operation; is the first derivative of, is the first derivative of, is the first derivative of, is the first derivative of, is the second derivative of;

[0035] S33: Design the control vector of the th following aircraft:

[0036]

[0037]

[0038]

[0039] Among them, is the first feedback matrix of the formation flight controller state, is the second feedback matrix of the formation flight controller state;

[0040] According to the control vector calculate and .

[0041] A formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, step S4 includes:

[0042] S41: Randomly generate meta-policy parameters ;

[0043] S42: Set the local policy parameters Perform local update for the maximum number of iterations times of local update to obtain the updated local policy parameters ;

[0044] S43: Update the meta-policy parameters to obtain the updated meta-policy parameters ;

[0045] S44: Determine whether the update times of the meta-policy parameters reach the iteration number threshold , if not, return to step S42; if so, execute step S45;

[0046] S45: Output the finally obtained meta-policy parameters .

[0047] A formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, the local update in step S42 includes the following steps:

[0048] S421: For each time step , select and record the action space , the state space , the state space of the next time step , the reward function and the constraint cost :

[0049]

[0050]

[0051]

[0052]

[0053]

[0054] wherein, is the position tracking error of the i-th following aircraft, is the error threshold of the reward function, is the reward threshold of the reward function, is the penalty threshold of the reward function, is the cost function threshold, is the abscissa of the leader aircraft in the inertial coordinate system, is the ordinate of the leader aircraft in the inertial coordinate system, is the vertical coordinate of the leader aircraft in the inertial coordinate system;

[0055] S422: Construct a training data set according to the action space , state space , next time step state space , reward function , constraint cost : :

[0056]

[0057] S423: Calculate the gradient of the reward advantage function and the gradient of the cost advantage function ;

[0058]

[0059]

[0060] where, is the th local policy function, is the expected cumulative reward obtained by following the policy starting from the state in the state space ; is the expected cumulative reward obtained by following the policy , starting from the state in the action space and the action space ; is the cost function associated with taking the action space in the state space ; is the cost value function related to the safety constraint of the state space ; is the parameter of the gradient of and the gradient of which is the partial derivative when is ; is the parameter of the gradient of and the gradient of which is the partial derivative when is ;

[0061] S424: Update the -th local policy function as the local policy function for the -th iteration

[0062]

[0063]

[0064]

[0065]

[0066] where is the constraint condition, represents taking the maximum value, is the Wasserstein distance between the old and new policies, is the residual, is 's decision cost, is the safety threshold.

[0067] According to a formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention, step S43 includes:

[0068] S431: Calculate the objective function gradient of the meta-policy parameter and the constraint function gradient :

[0069]

[0070]

[0071] where is the total number of tasks, is the expected value of the reward function of the -th task after local update, is the -th task's expected value of the cost function after local update;

[0072] S432: Update the meta-policy parameter to obtain the updated meta-policy parameter :

[0073]

[0074]

[0075]

[0076] where is the residual for meta-update cost constraint, is the Hessian matrix, is the size of the allowed meta-update trust region, is to obtain the Hessian matrix. (Here, argmax means to maximize this function)

[0077] The present invention also provides a formation flight control system based on cost-sensitive meta-reinforcement learning, including: a preprocessing module: establishing a three-dimensional kinematic model of formation flight, constructing a set target of a formation flight control method, and constructing a formation flight control model according to the three-dimensional kinematic model of formation flight and the set target of the formation flight control method;

[0078] A meta-policy parameter training module: training the formation flight control model according to cost-sensitive meta-reinforcement learning to obtain final meta-policy parameters.

[0079] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0080] A formation flight control method and system based on cost-sensitive meta-reinforcement learning provided by the present invention combine cost-sensitive meta-reinforcement learning with multi-aircraft formation control, enabling the aircraft to perceive and autonomously learn to cope with complex environmental disturbances in real time and accumulate flight experience; through meta-reinforcement learning iterative update and joint Wasserstein distance constraint, it can accurately solve a reasonable formation control scheme for the aircraft while satisfying the constraint control of the safety distance; in addition, by adding BLO (Bi-Level Optimization) to the meta-reinforcement learning algorithm for end-to-end gradient calculation to optimize the meta-policy parameters, the generalization ability of the present invention across tasks can be further improved. It provides an innovative solution to the problem of multi-aircraft formation flight.

[0081] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0083] Figure 1 is a schematic flowchart of a formation flight control method based on cost-sensitive meta-reinforcement learning provided by the present invention;

[0084] Figure 2 It is a schematic diagram of the position tracking error in the embodiment of the present invention;

[0085] Figure 3 It is a schematic diagram of the speed error of the aircraft in the embodiment of the present invention;

[0086] Figure 4 It is a schematic diagram of the ballistic inclination error of the aircraft in the embodiment of the present invention;

[0087] Figure 5 It is a schematic diagram of the ballistic deflection error of the aircraft in the embodiment of the present invention;

[0088] Figure 6 It is a schematic diagram of the flight trajectory of the aircraft in the embodiment of the present invention;

[0089] Figure 7 It is a structural block diagram of a formation flight control system based on cost-sensitive meta-reinforcement learning provided by the present invention.

[0090] Reference numerals:

[0091] 101, preprocessing module; 102, meta-policy parameter training module; 103, formation flight module. Specific implementation manners

[0092] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0093] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art may combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0094] The following will describe the present invention in conjunction with Figures 1 to 7 Describe the present invention.

[0095] Embodiment

[0096] As Figure 1 shown, an embodiment of the present invention provides a formation flight control method based on cost-sensitive meta-reinforcement learning, including the following steps:

[0097] S1: Establish a three-dimensional kinematic model of formation flight;

[0098] S2: Construct the set goals of the formation flight control method;

[0099] S3: Construct a formation flight control model according to the three-dimensional kinematic model of formation flight and the set goals of the formation flight control method;

[0100] S4: Train the formation flight control model according to cost-sensitive meta-reinforcement learning to obtain the final meta-policy parameters;

[0101] S5: Control the formation flight according to the final meta-policy parameters.

[0102] Specifically, step S1 includes:

[0103] S11: Design the state vectors of each aircraft in formation flight:

[0104]

[0105] Among them, is the state vector of the th aircraft, is the aircraft ordinal number, , is the total number of aircraft, is the abscissa of the th aircraft in the inertial coordinate system, is the ordinate of the th aircraft in the inertial coordinate system, The th vertical coordinate of the aircraft in the inertial coordinate system, represents the th aircraft speed, represents the th flight inclination angle of the aircraft, represents the th flight deviation angle of the aircraft.

[0106] S12: Design the control vectors of each following aircraft in formation flight:

[0107]

[0108] Among them, is the control vector of the th following aircraft, is the overload of the th following aircraft in the axis direction, is the overload of the th following aircraft in the axis direction, is the overload of the th following aircraft in the axis direction; is the matrix transpose;

[0109] S13: Construct the three-dimensional kinematic model of the formation flight according to the state vector and the control vector :

[0110]

[0111]

[0112]

[0113] wherein, is the gravitational acceleration, is the linearization matrix of the aircraft formation control, is the basic model of the aircraft formation control.

[0114] Specifically, step S2 includes:

[0115] Considering a formation system including N aircraft, we can define a leader aircraft and N - 1 following aircraft. The leader aircraft usually provides the reference state, while the following aircraft adjust according to the state of the leader aircraft to maintain a reasonable formation structure.

[0116] For any formation initial condition:

[0117]

[0118] wherein, is the position vector of the leader aircraft, is the th position vector of the following aircraft, is the expected distance between the leader aircraft and the th following aircraft, is the velocity vector of the leader aircraft, is the th velocity vector of the following aircraft, is the transformation matrix from the ballistic coordinate system of the leader aircraft to the inertial reference coordinate system, is the allowable error of the position vector, is the allowable error of the velocity vector.

[0119] Specifically, step S3 includes:

[0120] It is known that the formation control system includes a leader aircraft and N - 1 follower aircraft. In the inertial coordinate system, since each channel is independent, the control laws in the three directions of x, y, and z need to be designed separately. Control commands are designed for the i-th follower aircraft based on position and velocity information

[0121] S31: Define the adjacency matrix of the leader aircraft:

[0122]

[0123] Among them, is the communication state between the -th follower aircraft and the leader aircraft. If it is in a communicable state , if it is in a non - communicable state , is a diagonal matrix;

[0124] S32: Design the control command of the -th follower aircraft :

[0125]

[0126]

[0127] Among them, is the defined sign function, is the first independent variable of the defined sign function, is the second independent variable of the defined sign function, is the sign function, is to take the absolute value, is the ordinal number of the adjacent follower aircraft of the follower aircraft, is the -th follower aircraft's set of adjacent follower aircraft, is the -th follower aircraft's expected position in the leader ballistic coordinate system, represents the -th follower aircraft's position vector, is the first feedback gain of the formation flight controller, is the second feedback gain of the formation flight controller, is the first - order gain of the formation flight controller, is the second - order gain of the formation flight controller, is the damping gain of the formation flight controller; the damping gain can help absorb oscillations in the system and improve the system's stability. In aircraft control, this can reduce unstable behavior caused by rapid maneuvers or environmental changes; is the modulo operation; is the first derivative of, is the first derivative of, is the first derivative of, is the first derivative of, is the second derivative of; all function differentiations are with respect to time;

[0128] S33: Design the control vector for the th follower aircraft:

[0129]

[0130]

[0131]

[0132] where, is the first feedback matrix of the formation flight controller state, is the second feedback matrix of the formation flight controller state,

[0133] Calculate and and .

[0134] In reinforcement learning, the PPO (Proximal Policy Optimization) algorithm, which is representative of online policy algorithms, requires a large number of samples for learning and has the problem of low sample efficiency. The DDPG algorithm (deep deterministic policy gradient) is an off-policy algorithm. Although it improves the utilization rate of samples, DDPG is a deterministic policy, that is, only the best action is considered in each state, resulting in problems such as insufficient exploration ability, unstable training, poor convergence, sensitivity to hyperparameters, and difficulty in adapting to different complex environments. Compared with ordinary reinforcement learning, the CSMRL (Cost-sensitive Meta-reinforcement learning) algorithm of the embodiment of the present invention shows significant advantages in multiple aspects: it realizes the ability to quickly adapt to new tasks through a meta-learning framework, ensuring efficient learning in a dynamically changing environment; at the same time, the constraint policy optimization mechanism integrated in the algorithm naturally considers safety constraints during the policy optimization process, ensuring that the learning policy does not violate the preset safety limits while pursuing performance improvement, and overcomes the end-to-end differentiability problem in traditional cost-sensitive reinforcement learning, making the entire learning process smoother and more efficient.

[0135] Specifically, step S4 includes the following steps:

[0136] S41: Randomly generate meta-policy parameters ;

[0137] S42: Set the local policy parameters Perform the maximum number of local update iterations times of local updates to obtain the updated local policy parameters ;

[0138] S43: Update the meta-policy parameters to obtain the updated meta-policy parameters ;

[0139] S44: Determine whether the number of updates of the meta-policy parameters reaches the iteration number threshold , if not, return to step S42; if so, execute step S45;

[0140] S45: Output the finally obtained meta-policy parameters .

[0141] Among them, the local update in step S42 includes the following steps:

[0142] S421: For each time step , select and record the action space , the state space , the state space at the next time step , the reward function and the constraint cost :

[0143]

[0144]

[0145]

[0146]

[0147]

[0148] wherein, is the position tracking error of the i-th follower aircraft, is the error threshold of the reward function, is the reward threshold of the reward function, is the penalty threshold of the reward function, is the threshold of the cost function, is the abscissa of the leader aircraft in the inertial coordinate system, is the ordinate of the leader aircraft in the inertial coordinate system, is the vertical coordinate of the leader aircraft in the inertial coordinate system;

[0149] S422: Construct a training data set according to the action space , the state space , the state space at the next time step , the reward function , the constraint cost : :

[0150]

[0151] S423: Calculate the gradient of the reward advantage function and the gradient of the cost advantage function ;

[0152]

[0153]

[0154] wherein, is the -th local policy function, is the expected cumulative reward obtained by following the policy starting from the state in the state space ; Starting from the state space , the action space The expected cumulative reward obtained by following the policy is To represent the cost function associated with taking actions in the action space under the state space is The cost value function related to the safety constraint in the state space is is The gradient of and the gradient parameter is The partial derivative when is The gradient of and the gradient parameter is The partial derivative when

[0155] S424: Update the -th local policy function to be the local policy function of the -th iteration ,

[0156]

[0157]

[0158]

[0159]

[0160] where is the constraint condition is the matrix transpose denotes taking the maximum value The Wasserstein distance between the old and new policies is used to measure the step size of policy update, ensuring that the update does not deviate too far from the current policy is the maximum allowed Wasserstein distance, which defines the size of the trust region is the residual is 's decision cost (the expected value of the cost function of the task under the policy decided by ), is the safety threshold

[0161] Local update mainly applies the BLO (Bi-Level Optimization) method. BLO is used to calculate gradients, which allows the algorithm to optimize the meta-policy parameters through end-to-end gradient calculation and the local update results of all tasks, thus achieving bi-level optimization. The Wasserstein distance is used as a metric in local update to avoid instability caused by overly aggressive updates. In local update, the cost-sensitive reinforcement learning method optimizes the objective function to maximize the reward while ensuring the safety and feasibility of the policy through constraints.

[0162] Step S43 includes the following steps:

[0163] S431: Calculate the gradient of the objective function of the meta-policy parameters and the gradient of the constraint function :

[0164]

[0165]

[0166] where is the total number of tasks, is the -th task's expected value of the reward function after local update, is the -th task's expected value of the cost function after local update;

[0167] S432: Update the meta-policy parameters to obtain the updated meta-policy parameters :

[0168]

[0169]

[0170]

[0171] where is the residual of the meta-update cost constraint, is the Hessian matrix, that is, or the second-order partial derivative matrix of with respect to is the size of the allowed meta-update trust region, is to obtain the Hessian matrix.

[0172] During meta-update, the BLO method is mainly applied. BLO is used to calculate the gradient of the meta-policy parameters, which allows the algorithm to optimize the meta-policy parameters through end-to-end gradient calculation and the local update results of all tasks, thus achieving two-layer optimization. During meta-update, adjustments are made according to the local update results of all tasks to improve the generalization ability across tasks. Meta-update provides gradient information for BLO when evaluating the performance and cost on different tasks, helping to optimize the meta-policy.

[0173] S5: According to the final meta-policy parameters Control formation flight. The final meta-policy parameters Include formation flight tasks, including control values for speed, flight inclination, and flight deviation angle. Specifically, the final meta-policy parameters trained in the embodiments of the present invention Act on the flight formation as follows:

[0174] First, in the simulation environment, initialize the local policy parameters for each test task to the obtained final meta-policy parameters , and then fine-tune the policy to adapt to specific formation flight requirements. Execute the formation flight task in the simulation environment, control its speed, flight inclination, and flight deviation angle, and at the same time monitor the coordination and formation accuracy of the aircraft, ensuring that all aircraft comply with safety constraints. Evaluate the policy performance according to the test results and further adjust the policy if necessary.

[0175] In the embodiments of the present invention, as shown in Table 1 and Table 2, Table 1 is the Monte Carlo simulation environment parameters, and Table 2 is the setting of network parameters in the intelligent network CSMRL algorithm:

[0176] Table 1: Monte Carlo simulation environment parameters

[0177]

[0178] Among them, Is the speed of the leader aircraft, Is the first parameter of the leader aircraft speed, Is the second parameter of the leader aircraft speed, Is the third parameter of the leader aircraft speed, Is the first parameter of the leader aircraft angle, The second parameter of the leader aircraft angle, The third parameter of the leader aircraft angle, The fourth parameter of the leader aircraft angle, Is the abscissa of the expected distance between the leader aircraft and the first following aircraft, Is the ordinate of the expected distance between the leader aircraft and the first following aircraft, is the vertical coordinate of the desired distance between the leader aircraft and the first follower aircraft; is the abscissa of the desired distance between the leader aircraft and the second follower aircraft, is the ordinate of the desired distance between the leader aircraft and the second follower aircraft, is the vertical coordinate of the desired distance between the leader aircraft and the second follower aircraft.

[0179] Table 2 CSMRL network parameter settings

[0180]

[0181] As Figure 2 shown, Figure 2 is the schematic diagram of the position tracking error of the embodiment of the present invention. According to Figure 2 it can be seen that after training, the tracking error is less than 1.5 meters, and the speed of error reduction is also very fast, which takes about 1500 seconds.

[0182] Figure 3 is the schematic diagram of the speed error of the aircraft in the embodiment of the present invention. According to Figure 3 it can be seen that after about 1500 seconds, the speed change rules and periods of the two follower aircraft are similar to those of the leader aircraft.

[0183] Figure 4 is the schematic diagram of the ballistic inclination error of the aircraft in the embodiment of the present invention. According to Figure 4 it can be seen that after about 1500 seconds, the inclination change rules and periods of the two follower aircraft are similar to those of the leader aircraft.

[0184] Figure 5 is the schematic diagram of the ballistic deflection error of the aircraft in the embodiment of the present invention. According to Figure 5 it can be seen that after about 1500 seconds, the elevation change rules and periods of the two follower aircraft are similar to those of the leader aircraft.

[0185] Figure 6 is the schematic diagram of the flight trajectory of the aircraft in the embodiment of the present invention; According to Figure 6 it can be seen that the flight trajectories of the two follower aircraft quickly become consistent with that of the leader aircraft after flying for a period of time.

[0186] As Figure 7 shown, Figure 7 is the structural block diagram of a formation flight control system based on cost-sensitive meta-reinforcement learning provided by the present invention. It is used to execute a formation flight control method based on cost-sensitive meta-reinforcement learning, including:

[0187] Preprocessing module 101: Establish a three-dimensional kinematic model for formation flight, construct the set goals of the formation flight control method, and construct a formation flight control model according to the three-dimensional kinematic model of formation flight and the set goals of the formation flight control method;

[0188] Meta-policy parameter training module 102: Train the formation flight control model according to cost-sensitive meta-reinforcement learning to obtain the final meta-policy parameters;

[0189] Formation flight module 103: Control formation flight according to the final meta-policy parameters.

[0190] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0191] It should be noted that the embodiments of the present disclosure can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic: The software part can be stored in a memory and executed by a suitable instruction execution system such as a microprocessor or dedicated design hardware. Those skilled in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a programmable memory or a data carrier such as an optical or electronic signal carrier.

[0192] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can change the execution order. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of one device described above can be further divided and embodied by multiple devices.

Claims

1. A formation flight control method based on cost-sensitive meta-reinforcement learning, characterized in that: include: S1: Establish a three-dimensional kinematic model of formation flying; S2: Setting goals for constructing formation flight control methods; S3: constructing a formation flight control model according to the formation flight three-dimensional kinematic model and the set target of the formation flight control method; S4: Train the formation flight control model based on cost-sensitive meta-reinforcement learning to obtain the final meta-strategy parameters, including: S41: Randomly generate meta-strategy parameters ; S42: Setting local policy parameters Maximum number of iterations for local update The local strategy parameters are updated by the local update. , local update includes the following steps: S421: For each time step , select and record the action space , state space , the state space of the next time step , the reward function and constraint cost : in, is the position tracking error of the ith following vehicle, is the reward function error threshold, is the reward threshold of the reward function, is the penalty threshold of the reward function, is the cost function threshold, is the horizontal coordinate of the leading aircraft in the inertial coordinate system, is the ordinate of the leading aircraft in the inertial coordinate system, is the vertical coordinate of the leading aircraft in the inertial coordinate system, For the The horizontal coordinate of the following aircraft in the inertial coordinate system, For the The ordinate of the following aircraft in the inertial coordinate system, No. The vertical coordinate of the following aircraft in the inertial coordinate system, To follow the aircraft sequence, , is the total number of aircraft, Indicates The speed of the following aircraft, Indicates The flight inclination angle of the following aircraft, Indicates The flight angle of the following aircraft, For the Follow the aircraft in Overload in the axial direction, For the Follow the aircraft in Overload in the axial direction, For the Follow the aircraft in Overload in the axial direction, For the The position vector of the following vehicle, It is the leading aircraft and the The desired distance between the following aircraft, To take the absolute value, is the first feedback gain of the formation flight controller, is the second feedback gain of the formation flight controller, is the first-order gain of the formation flight controller; S422: Based on the action space , state space , the state space of the next time step , the reward function , constraint cost Building a training dataset : S423: Calculate the gradient of the reward advantage function and the cost advantage function gradient ; in, For the The sub-local policy function, From the state space Status starts following policy The expected cumulative rewards obtained; From the state space , Action Space Status starts following policy The expected cumulative reward obtained is To represent in state space Take action space The associated cost function, is the state space The cost value function associated with the security constraint, for The gradient and The gradient parameter for The partial derivative when ; for The gradient and The gradient parameter for The partial derivative when ; S424: Update The local strategy function is The local policy function of the iteration , in, As constraints, Indicates taking the maximum value, is the Wasserstein distance between the new and old strategies, is the maximum allowed Wasserstein distance, is the residual, for The decision cost, is the safety threshold, Transpose the matrix; S43: Update the meta-strategy parameters to obtain updated meta-strategy parameters ; S44: Determine whether the update times of the meta-strategy parameter reaches the iteration times threshold If not, return to step S42; if reached, execute step S45; S45: Output the final meta-strategy parameters ; S5: Controlling formation flight according to the final meta-strategy parameters.

2. A formation flight control method based on cost-sensitive meta-reinforcement learning according to claim 1, characterized in that: Step S1 is specifically as follows: S11: Definition In a formation system of aircraft, there is a leading aircraft and There are a number of follower aircraft, and the state vectors of each follower aircraft in formation flying are designed as follows: in, For the A state vector of the following vehicle; S12: Design the control vectors of each follower aircraft in formation flying: in, For the A control vector for the following vehicle; S13: Constructing the formation flight three-dimensional kinematic model according to the state vector and the control vector : in, is the acceleration due to gravity, is the linearization matrix for the flight formation control, A basic model for aircraft formation control.

3. A formation flight control method based on cost-sensitive meta-reinforcement learning according to claim 2, characterized in that: Step S2 is specifically as follows: For any formation, the initial conditions are: in, is the position vector of the leader vehicle, is the velocity vector of the leading vehicle, For the A velocity vector of the following vehicle, is the transformation matrix from the ballistic coordinate system of the leading aircraft to the inertial reference coordinate system, is the position vector allowable error, is the allowable error of the velocity vector.

4. The formation flight control method based on cost-sensitive meta-reinforcement learning according to claim 3, characterized in that: Step S3 is specifically as follows: S31: Define the adjacency matrix of the leader aircraft: in, For the The communication status between the follower aircraft and the leader aircraft, if it is in a communication state, If it is in a non-communication state, , is a diagonal matrix; S32: Design Follow the control commands of the aircraft : in, To define a symbolic function, is the first argument of the symbolic function, To define the second argument of the symbolic function, is the symbolic function, is the ordinal number of the adjacent following aircraft of the following aircraft, For the The set of adjacent following aircraft of a following aircraft, It is The expected position of the follower vehicle in the leader's trajectory coordinate system, Indicates The position vector of the following vehicle, is the second-order gain of the formation flight controller, is the formation flight controller damping gain; It is a modulo operation; for The first derivative of for The first derivative of for The first derivative of for The first derivative of for The second derivative of S33: Design The control vector of the following vehicle : in, is the first feedback matrix of the formation flight controller state, is the second feedback matrix of the formation flight controller state; According to the control vector calculate and .

5. A formation flight control method based on cost-sensitive meta-reinforcement learning according to claim 4, characterized in that: Step S43 includes: S431: Calculate the objective function gradient of the meta-strategy parameters and the constraint function gradient : in, is the total number of tasks, For the The expected value of the reward function after local update for each task, It is The expected value of the cost function of each task after local update; S432: Update the meta-policy parameters to obtain updated meta-policy parameters : in, is the residual of the meta-update cost constraint, is the Hessian matrix, is the size of the allowed meta-update trust domain, To obtain the Hessian matrix.

6. A formation flight control system based on cost-sensitive meta-reinforcement learning, used to execute a formation flight control method based on cost-sensitive meta-reinforcement learning as claimed in any one of claims 1 to 5, characterized in that: include: Preprocessing module: establishing a three-dimensional kinematic model of formation flight, constructing a set target of a formation flight control method, and constructing a formation flight control model according to the three-dimensional kinematic model of formation flight and the set target of the formation flight control method; Meta-strategy parameter training module: trains the formation flight control model based on cost-sensitive meta-reinforcement learning to obtain the final meta-strategy parameters; Formation flying module: controls the formation flying according to the final meta-strategy parameters.

Citation Information

Patent Citations

  • Unmanned aerial vehicle flight decision-making method based on meta-reinforcement learning parallel training algorithm

    CN114895697A

  • Unmanned aerial vehicle formation collaborative decision-making method based on meta learning and MADDPG

    CN117111632A