A method and system for adaptive control of aircraft cluster communication topology

Through the aircraft cluster control method that is coupled with the control strategy and communication strategy built by deep neural network, the communication topology is adaptively adjusted, which solves the problems of low robustness and large communication volume of aircraft cluster communication topology in the prior art, and realizes efficient control in complex environments.

CN116107346BActive Publication Date: 2025-05-13HARBIN INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310285168.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-05-13
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

The communication topology of existing aircraft clusters is low in complex dynamic environments and has a large communication volume, making it difficult to adapt to the dynamic changes in the number of aircraft clusters and different flight environments.

Method used

Deep neural network is used to build a vehicle cluster control strategy coupled with control strategies and communication strategies. By obtaining the interaction information between the aircraft cluster and the environment, the network is trained to adaptively adjust the communication topology structure.

Benefits of technology

It realizes adaptive adjustment of communication topology in complex environments, reduces communication volume, improves robustness, and adapts to changes in the number of aircraft clusters and different flight environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116107346B_ABST
    Figure CN116107346B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive control method and system for aircraft cluster communication topology, which relates to the field of aircraft cluster control technology and is used to solve the problems of low robustness and large communication volume required for high-speed aircraft cluster control strategies based on traditional communication mechanisms. The technical highlights of the present invention include: obtaining interaction information between an aircraft cluster and the environment; using the interaction information between the aircraft cluster and the environment to train an aircraft cluster control strategy network based on a deep neural network; and using the trained aircraft cluster control strategy network to control the movement and communication topology of the aircraft cluster. The aircraft cluster control strategy implemented by the present invention has the ability to control the movement behavior of aircraft and the communication topology of aircraft clusters, and can adaptively adjust the communication topology of the cluster in a complex cluster mission environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft cluster control, and in particular to an aircraft cluster communication topology adaptive control method and system. Background Art

[0002] In the process of aircraft cluster control, mutual communication between aircraft is one of the foundations of cluster collaboration. Communication between cluster aircraft is mainly used for the exchange of status and load information between aircraft. Due to the large number of nodes, many types of tasks, fast flight speed, frequent changes in relative time and space relationships, and the immediacy and suddenness of information transmission, the communication and networking between clusters are very challenging. At present, the commonly used communication topologies for aircraft clusters include centralized communication topology, hierarchical communication topology and distributed communication topology. In the centralized communication topology, all nodes are connected to a central node, which is responsible for maintaining the communication of the entire network, which is suitable for the situation with a small number of nodes; in the hierarchical communication topology, nodes are connected according to certain rules to form a tree structure, which is suitable for the situation with a central coordination unit; in the distributed communication topology, nodes are arbitrarily connected to form a mesh-like structure, which is suitable for the situation with a large number of nodes.

[0003] Although a lot of research has been done on the design of communication topology for aircraft clusters, the pre-designed communication topology network has certain limitations in practical applications. In the two communication mechanisms of centralized communication and hierarchical communication, the "leader" aircraft has a greater communication pressure, and once the "leader" aircraft fails, the entire aircraft cluster will not be able to continue to complete the mission. That is, the centralized communication mechanism and the hierarchical communication mechanism have low robustness and cannot be applied to complex and dynamic confrontation environments. In the distributed communication mechanism, it is usually designed that the aircraft interacts with all aircraft within the communication range. This interaction mechanism based on fixed rules does not take into account the actual environment faced by the aircraft cluster. It will not only add additional communication burden to the aircraft, but also affect the performance of the aircraft cluster control. Too much or too little communication with neighboring aircraft will have an adverse effect on the control process of the aircraft cluster. Therefore, it is extremely necessary to develop a cluster communication control strategy that can adaptively adjust the communication topology according to the environmental state for the complex and changeable aircraft cluster mission environment. Summary of the invention

[0004] To this end, the present invention proposes an aircraft cluster communication topology adaptive control method and system to solve the problems of low robustness and large required communication volume of high-speed aircraft cluster control strategy based on traditional communication mechanism.

[0005] According to one aspect of the present invention, a method for adaptive control of aircraft cluster communication topology is provided, which uses a deep neural network to construct an aircraft cluster control strategy that couples a control strategy with a communication strategy, and whose output includes an overload instruction for controlling the movement of an aircraft and the number of communications with neighboring aircraft; the adaptive control method comprises the following steps:

[0006] Acquire interaction information between the aircraft cluster and the environment, wherein the interaction information includes observation states, executed control instructions, and reward values ​​of the aircraft at multiple times;

[0007] Using the interactive information between the aircraft cluster and the environment to train an aircraft cluster control strategy network based on a deep neural network;

[0008] The trained aircraft cluster control strategy network is used to control the movement and communication topology of the aircraft cluster.

[0009] Furthermore, the observed state includes the state information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the obstacle area th , information of the aircraft relative to the cluster target point o ta ;in,

[0010] Aircraft status information relative to the group formation Contains the formation deviation status of individual aircraft relative to the initial formation and the speed deviation state of the flight speed of the individual aircraft relative to the speed of the aircraft cluster

[0011] Information about the aircraft relative to the obstacle zone Contains the relative position of the aircraft relative to the obstacle area closest to the aircraft in the flight environment, vector ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the obstacle area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the obstacle area, k is the flag position;

[0012] Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V ; ΔR t (ΔR t ∈R3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient.

[0013] Furthermore, the control instructions are obtained by inputting the observed state of the aircraft cluster at each moment into an aircraft cluster control strategy network based on a deep neural network for processing, and the control instructions include overload instructions and communication instructions.

[0014] Furthermore, the reward value is decomposed as follows:

[0015]

[0016] In the formula, Indicates position reward; Indicates target strike reward; Represents the distance change reward; r ost It indicates a threat zone avoidance bonus; represents the relative speed reward; Indicates relative formation reward.

[0017] Furthermore, the aircraft cluster control strategy network based on deep neural network is constructed as follows: including value function network and strategy network, the network structure is as follows:

[0018] The observation state of the aircraft is spliced ​​into an observation vector. The observation vector and the control command are spliced ​​and input into the value function network. Then, they are passed through a two-layer 128-node fully connected network and divided into two information paths: one information path is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 2 nodes and activation functions of Linear and Tanh respectively, to obtain the mean and logarithmic standard deviation of the overload command for controlling the movement of the aircraft, and the overload command of the aircraft is obtained by sampling the Gaussian distribution. Finally, the information is processed by a fully connected layer with 2 nodes and an activation function of Linear and Tanh. The information is processed by a fully connected layer with a Tanh function to limit the overload instructions to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 1 node and an activation function of Linear and Tanh respectively to obtain the mean and logarithmic standard deviation of the communication control instructions, and the communication instructions of the aircraft are obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 1 node and an activation function of Tanh to limit the communication instructions to (-1, 1).

[0019] Furthermore, the process of training the aircraft cluster control strategy network based on the deep neural network includes:

[0020] Initialize the cluster control strategy network parameters θ, cluster control strategy value function network parameters φ1 and φ2 and the experience pool D;

[0021] Set the value function target network parameter φ targ,1 ,φ targ,2 The same as parameters φ1 and φ2, that is, φ targ,1 ←φ1,φ targ,2 ←φ2;

[0022] Repeat the following steps:

[0023] 1) Observe the state of the cluster simulation environment, that is, observe the state s, s = [o f ,o th ,o ta ], output cluster control instructions according to the control strategy;

[0024] 2) Execute control instruction a in the cluster simulation environment;

[0025] 3) Observe the next state s', the feedback reward r and the round end flag d;

[0026] 4) Store the experience group (s, a, r, s', d) in the experience pool D;

[0027] 5) If the round ends, reset the environment state; if the update cycle is reached, execute steps 6) to 10);

[0028] 6) Randomly sample a set of experiences B = {(s, a, r, s', d)} from the experience pool D;

[0029] 7) Calculate the true value estimate of the value function by the following formula:

[0030]

[0031] Where: γ represents the discount rate, γ∈[0,1]; α represents a positive coefficient; Represents state s′, action The objective value function of represents the policy network; Representation Policy Network The action of sampling;

[0032] 8) Update the value function network parameters φ by minimizing the following loss function i :

[0033]

[0034] Where: |B| represents the number of batch data; Represents the target value function of state s and action a;

[0035] 9) Update the cluster control strategy network parameters θ by minimizing the following loss function:

[0036]

[0037] Where: Indicates state s, action The objective value function of represents the policy network; Representation Policy Network The action of sampling;

[0038] 10) Update the target network:

[0039] φ targ,i ←ρφ targ,i +(1-ρ)φ i ,i=1,2

[0040] Where: ρ represents a positive coefficient.

[0041] According to another aspect of the present invention, there is provided an aircraft cluster communication topology adaptive control system, the system comprising:

[0042] A data acquisition module configured to acquire interaction information between the aircraft cluster and the environment, wherein the interaction information includes observation states, executed control instructions, and reward values ​​of the aircraft at multiple times;

[0043] A strategy network training module, configured to train an aircraft cluster control strategy network based on a deep neural network using interaction information between the aircraft cluster and the environment;

[0044] A control and communication module is configured to control the movement and communication topology of the aircraft cluster using the trained aircraft cluster control strategy network.

[0045] Furthermore, the observed state includes the state information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the obstacle area th , information of the aircraft relative to the cluster target point o ta ;in,

[0046] Aircraft status information relative to the group formation Contains the formation deviation status of individual aircraft relative to the initial formation and the speed deviation state of the flight speed of the individual aircraft relative to the speed of the aircraft cluster

[0047] Information about the aircraft relative to the obstacle zone Contains the relative position of the aircraft relative to the obstacle area closest to the aircraft in the flight environment, vector ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the obstacle area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the obstacle area, k is the flag position;

[0048] Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V ; ΔR t (ΔR t ∈R 3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient;

[0049] The control instructions are obtained by inputting the observed state of the aircraft cluster at each moment into an aircraft cluster control strategy network based on a deep neural network for processing, and the control instructions include overload instructions and communication instructions.

[0050] Furthermore, the reward value is decomposed as follows:

[0051]

[0052] In the formula, Indicates position reward; Indicates target strike reward; Represents the distance change reward; r ost It indicates a threat zone avoidance bonus; represents the relative speed reward; Indicates relative formation reward.

[0053] Furthermore, the aircraft cluster control strategy network based on deep neural network is constructed as follows: including value function network and strategy network, the network structure is as follows:

[0054] The observation state of the aircraft is spliced ​​into an observation vector. The observation vector and the control command are spliced ​​and input into the value function network. Then, they are passed through a two-layer 128-node fully connected network and divided into two information paths: one information path is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 2 nodes and activation functions of Linear and Tanh respectively, to obtain the mean and logarithmic standard deviation of the overload command for controlling the movement of the aircraft, and the overload command of the aircraft is obtained by sampling the Gaussian distribution. Finally, the information is processed by a fully connected layer with 2 nodes and an activation function of Linear and Tanh. The information is processed by a fully connected layer with a Tanh function to limit the overload instructions to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 1 node and an activation function of Linear and Tanh respectively to obtain the mean and logarithmic standard deviation of the communication control instructions, and the communication instructions of the aircraft are obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 1 node and an activation function of Tanh to limit the communication instructions to (-1, 1).

[0055] The beneficial technical effects of the present invention are:

[0056] Compared with the aircraft cluster control strategy based on the traditional communication mechanism, the aircraft cluster control strategy implemented by the present invention has the ability to control both the aircraft movement behavior and the aircraft cluster communication topology. It can adaptively adjust the cluster communication topology in a complex cluster mission environment, control the aircraft cluster to avoid the threat area in the environment under low communication volume, and quickly reach the cluster target point. At the same time, the aircraft cluster communication topology adaptive control strategy proposed by the present invention can adapt to the dynamic changes in the number of aircraft clusters and different cluster flight environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings, which together with the following detailed description are included in this specification and form a part of this specification, and are used to further illustrate the preferred embodiments of the present invention and explain the principles and advantages of the present invention.

[0058] Figure 1 It is a schematic diagram of the network structure of the aircraft cluster control strategy in an embodiment of the present invention.

[0059] Figure 2 Schematic diagram of the Q-value function network structure in an embodiment of the present invention.

[0060] Figure 3 It is a curve diagram of reward value variation during the aircraft cluster control strategy training process in an embodiment of the present invention.

[0061] Figure 4 1 is a trajectory curve generated by the aircraft cluster in the training environment in an embodiment of the present invention, wherein Figure (a), Figure (b), Figure (c), and Figure (d) correspond to the flight trajectories of the aircraft cluster at different target point positions, respectively. DETAILED DESCRIPTION

[0062] In order to enable those skilled in the art to better understand the scheme of the present invention, exemplary implementations or embodiments of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described implementations or embodiments are only implementations or embodiments of a part of the present invention, not all of them. Based on the implementations or embodiments of the present invention, all other implementations or embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.

[0063] The present invention proposes an aircraft cluster control method with the ability to adaptively adjust the cluster communication topology. A deep neural network is used to construct a cluster control strategy that couples the control strategy with the communication strategy under a deep reinforcement learning architecture. The output includes an overload instruction for controlling the movement of the aircraft and the number of communications with neighboring aircraft. Through continuous interaction with the mission environment, the trained cluster control strategy can autonomously adjust the communication topology structure according to environmental information, thereby ensuring the robustness of cluster control and low communication volume.

[0064] Consider a high-speed aircraft cluster U(u i ∈U,i=1,2,…N) consists of N high-speed aircraft, and there are M obstacle areas B (b i ∈B,i=1,2,…M), the communication relationship between each high-speed aircraft and the surrounding high-speed aircraft is represented by the vector set C(c i ∈C,c i =[1 1 ... 0] N , i=1,2,…N) indicates that 1 represents communicating with it, 0 represents not communicating with it, and the default aircraft communicates with itself.

[0065] The goal of swarm control is to establish a swarm controller that enables high-speed aircraft u i The following conditions are met:

[0066] (1) The high-speed aircraft cluster should maintain its initial formation as much as possible during the flight;

[0067] (2) Minimize the u of each high-speed aircraft i From the initial position Arrival at the target location Time

[0068] (3) Each high-speed aircraft is able to avoid threat areas in the environment.

[0069] Therefore, an embodiment of the present invention provides an aircraft cluster communication topology adaptive control method, the method comprising the following steps:

[0070] Step 1: Obtain interaction information between the aircraft cluster and the environment The interactive information includes the observation status of the aircraft at multiple times, the executed control instructions and the reward value. Specifically, represents the observed state of the i-th aircraft at time t, represents the control command executed by the i-th aircraft at time t, represents the observed state of the i-th aircraft at time t+1, r i t+1 represents the reward value obtained by the i-th aircraft at time t+1;

[0071] Step 2: Use the interactive information between the aircraft cluster and the environment to train the aircraft cluster control strategy network based on the deep neural network;

[0072] Step three: Use the trained aircraft cluster control strategy network to control the movement and communication topology of the aircraft cluster; specifically, the obtained aircraft control instructions control the movement of the aircraft, and adjust the communication topology of the aircraft cluster by changing the number of neighboring aircraft communicating with the aircraft.

[0073] In step 1, each aircraft in the aircraft cluster obtains its own observation state, and the observation state of each aircraft in the aircraft cluster is o=[o f ,o th ,o ta ], including the status information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the threat zone th , the information of the aircraft relative to the cluster target point o ta .

[0074] 1) Aircraft status information relative to the cluster formation Contains the formation error state of the individual aircraft relative to the initial formation and the deviation of the flight speed of the individual aircraft relative to the speed of the aircraft cluster Among them, the formation error state is expressed as:

[0075]

[0076] In the formula, ΔR i (ΔR i ∈R 3, i=1,2,……,n) represents the deviation of the relative position relationship between the current aircraft and the i-th aircraft relative to the relative position relationship between the two in the initial formation, R f is the position normalization coefficient, n represents the number of neighboring aircraft communicating with the current aircraft, and n max Indicates the maximum number of neighboring aircraft for the set aircraft communication.

[0077] The speed deviation state is expressed as Where ΔV represents the deviation between the current aircraft’s flight speed and the average flight speed of the aircraft cluster, V f Represents the speed normalization coefficient.

[0078] 2) Information about the aircraft relative to the obstacle area Contains the relative position of the aircraft relative to the threat area closest to the aircraft in the flight environment, ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the threat area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the threat area. k is a flag, and its value is as follows:

[0079]

[0080] Where, d th Indicates the distance between the current aircraft and the nearest threat area.

[0081] 3) Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V , where ΔR t (ΔR t ∈R 3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient.

[0082] The observed states of each aircraft are input into the aircraft cluster control strategy network based on deep neural network for processing, and the control instructions of each aircraft can be obtained. The control instructions include overload instructions and communication instructions.

[0083] In step 2, if Figure 1 As shown in the figure, the structure of the aircraft cluster control strategy network based on deep neural network includes a value function network (such as Figure 2 As shown in the figure, the observation information of the aircraft is concatenated to obtain a joint observation state, which is first processed by two layers of fully connected layers with 128 nodes and ReLU activation function, and then divided into two paths of information. One path of information is processed by a fully connected layer with 128 nodes and ReLU activation function, and then processed by two fully connected layers with Linear and Tanh node activation functions respectively, to obtain the mean value of the overload command for controlling the movement of the aircraft. and the logarithmic standard deviation By sampling the Gaussian distribution Get the overload command of the aircraft Finally, it is processed by a fully connected layer with 2 nodes and the activation function is Tanh, and the overload instruction is limited to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and the activation function is ReLU, and then processed by a fully connected layer with 1 node and the activation function is Linear and Tanh respectively, and the mean of the communication control instruction is obtained. and log standard deviation By sampling the Gaussian distribution Get the aircraft's communication command n c , and finally processed by a fully connected layer with 1 node and the activation function Tanh, limiting the communication instructions to (-1,1).

[0084] Suppose the observation state space of each aircraft in the cluster formation is S, the output action is the cluster control instruction a∈A, and A is the action space. The cluster control agent starts from the initial state s0~p(s0) and distributes according to the strategy a t ~π(·|s t ) Sample and output a control instruction a t Acting on the cluster environment, the cluster updates the state according to the input command and obtains a reward feedback r(s t ,a t ) and based on the cluster environment model s t+1 ~p(·|s t ,a t ) to a new state s t+1 , and keep repeating this process until the stopping condition is met. A cycle is called an episode. The cumulative reward of an episode is defined as:

[0085]

[0086] Where: γ∈[0,1] is the discount rate, γ→1 makes the learned strategy more focused on long-term returns, and T represents the total number of steps in which the agent interacts with the environment in one round. The training goal is to improve the strategy π so that G t maximize.

[0087] The basic principle of reward value, i.e. reward function, is to reward good states and behaviors and punish states and behaviors that are contrary to the optimization goal. Corresponding to the observed state, the reward function design also includes three sources: cluster formation, threat area avoidance, and target strike.

[0088] 1) The relative formation reward in the cluster formation reward is expressed as

[0089]

[0090] Where: λ is a negative coefficient, which makes the reward decrease as the formation deviation increases, and limits the error to When the formation deviation is 0, the reward reaches the maximum value of 0, i min The ID of the aircraft closest to the current aircraft. Indicates the current moment of the aircraft relative to aircraft i min The position vector of the aircraft relative to aircraft i at the initial moment min The difference in position vectors, It represents the difference between the position vector of the aircraft relative to the center of the aircraft cluster at the current moment and the position vector of the aircraft relative to the center of the cluster at the initial moment, R fmt is the distance scaling factor.

[0091] The relative speed reward is expressed as

[0092]

[0093] Where: μ is a negative coefficient, is the upper limit of the speed error amplitude; The vector representing the difference between the current aircraft speed and the center speed of the aircraft cluster in the current aircraft speed coordinate system, V fmt is the speed scaling factor.

[0094] 2) The threat zone avoidance reward is represented by r ost :

[0095]

[0096] Where: η and σ are negative reward coefficients. When the distance between the aircraft and the nearest threat area is greater than When the distance is less than d0 but greater than the radius of the threat zone, the reward is 0. As the distance decreases, the reward also decreases; when the aircraft enters the threat area, it gets the smallest reward, that is, the penalty is the largest.

[0097] 3) Target strike reward is expressed as

[0098]

[0099] Where: ρ is a positive coefficient; V atk V represents the dot product of the aircraft velocity vector and the unit vector pointing to the target point. tgt Represents the velocity scaling factor.

[0100] In order to provide additional hints to the agent to assist training convergence, position rewards and distance change rewards are added to the above rewards. tgt Get extra position bonus

[0101]

[0102] Where: κ is a positive reward coefficient.

[0103] The distance change reward is expressed as is the reward for the aircraft moving toward the target point, which is obtained by calculating the change in the distance between the aircraft and the target point at each simulation step:

[0104]

[0105] In the formula, and are the distances between the aircraft and the target at step t+1 and step t, respectively. ζ is a positive reward coefficient. If the aircraft moves toward the target point, the distance decreases. You will get a positive reward, otherwise you will get a negative reward.

[0106] In summary, the total reward function is:

[0107]

[0108] The training process of the aircraft cluster control strategy network proposed in the embodiment of the present invention includes the following steps:

[0109] Initialize the cluster control strategy network parameters θ, cluster control strategy value function network parameters φ1 and φ2 and the experience pool D;

[0110] Set the value function target network parameter φ targ,1 ,φ targ,2 The same as parameters φ1 and φ2 respectively: φ targ,1 ←φ1,φ targ,2←φ2;

[0111] Repeat the following steps:

[0112] 1) Observe the cluster simulation environment state s and output cluster control instructions a~π according to the control strategy θ (·|s);

[0113] 2) Execute control instruction a in the cluster simulation environment;

[0114] 3) Observe the next state s' and the feedback reward r and the round end flag d;

[0115] 4) Store the experience group (s, a, r, s', d) in the experience pool D;

[0116] 5) If the round ends, reset the environment state; if the update cycle is reached, execute steps 6) to 10);

[0117] 6) Randomly sample a set of experiences B = {(s, a, r, s', d)} from the experience pool D;

[0118] 7) Calculate the true value estimate of the value function by the following formula:

[0119]

[0120] Where: γ represents the discount rate, γ∈[0,1]; α represents a positive coefficient; Represents state s′, action The objective value function of represents the policy network; Representation Policy Network The action of sampling;

[0121] 8) Update the value function network parameters φ by minimizing the following loss function i :

[0122]

[0123] Where: |B| represents the number of batch data; Q φi (s,a) represents the target value function of state s and action a;

[0124] 9) Update the cluster control strategy network parameters θ by minimizing the following loss function:

[0125]

[0126] Where: Indicates state s, action The objective value function of represents the policy network; Representation Policy Network The action of sampling;

[0127] 10) Update the target network:

[0128] φ targ,i ←ρφ targ,i +(1-ρ)φ i ,i=1,2

[0129] Where: ρ represents a positive coefficient.

[0130] Then the continuous communication control instruction n c After discretization, the integer communication control instruction is obtained as follows:

[0131]

[0132] In the formula, ceil(·) is the upward rounding function, n max Indicates the maximum number of adjacent aircraft for aircraft communication.

[0133] The technical effect of the present invention is further verified through experiments.

[0134] Digital simulation is used to verify the correctness and rationality of the present invention. Python language is used to construct a high-speed aircraft cluster simulation environment, in which 9 threat areas of the same size are evenly distributed in a circular shape. Each threat area is represented by a circular area with a radius of 10km, and the distance between the center of each threat area and the center of the distribution circle is 60km. During the training process, an aircraft formation consisting of 10 aircraft is used to train the cluster control strategy. The initial speed of each aircraft is set to 1km / s, and the initial altitude is set to 10km. The maximum number of neighboring aircraft communicating with the aircraft is set to n max =6. Secondly, the simulation test software environment of the present invention is Windows 10+Python3.7, and the hardware environment is AMDRyzen 5 3550H CPU+16.0GB RAM.

[0135] Figure 3 The figure shows the reward value curve obtained by the cluster control strategy during the training process. The curve shown in the figure is the average and variance of the reward value obtained by the cluster control strategy in every 100 adjacent training cycles. As can be seen from the figure, the training of the cluster control strategy has gone through 2000 training cycles. After 500 trainings, the cumulative reward value received in each round remains basically stable, indicating that the training of the cluster control strategy is gradually converging. Figure 3 The results shown prove that the adaptive cluster communication topology control strategy proposed in the present invention can be stably trained.

[0136] Figure 4The figure shows the trajectory curve generated by the cluster control strategy trained above to control the aircraft cluster in the training environment. As can be seen from the figure, in all four scenarios, the aircraft cluster can safely avoid the threat area in the mission environment, indicating that the cluster control strategy proposed in the present invention has high safety. In addition, it is noted that after the aircraft cluster encounters the threat area, its formation will be deformed due to the need to avoid the threat area. After leaving the threat area, its formation tends to restore the initial formation. This shows that under the adaptive communication mechanism strategy proposed in the present invention, the aircraft cluster can autonomously adjust the formation to avoid the threat area, which also shows that the cluster control strategy proposed in the present invention has a good formation maintenance ability. In addition, in the above four test scenarios, the average communication volume of the cluster is 0.39, 0.35, 0.30, and 0.35, respectively, which is much lower than the average communication volume of the cluster under the distributed communication mechanism (under the distributed communication mechanism, the aircraft communicates with the six adjacent aircraft in a fixed manner, and the average communication volume of the cluster in this mode is 1). It shows that the adaptive cluster communication topology control strategy proposed in the present invention can greatly reduce the communication volume of the cluster and alleviate the communication burden of the cluster compared with the distributed communication mechanism. f Defined as:

[0137]

[0138] In the formula, T represents the time it takes for the aircraft cluster to travel from the initial point to the target point, N represents the number of aircraft in the aircraft cluster, represents the number of neighboring aircraft communicating with the i-th aircraft, n max Indicates the maximum number of communications that is set.

[0139] According to the method of the present invention, an adaptive communication topology control strategy for an aircraft cluster can be constructed to control the aircraft cluster to autonomously adjust the communication topology of the aircraft cluster according to the mission environment. While ensuring the safe and rapid completion of the established tasks, the communication volume of the cluster is reduced as much as possible, providing a feasible technical approach for the practical application of aircraft clusters in a confrontational environment.

[0140] Another embodiment of the present invention provides an aircraft cluster communication topology adaptive control system, the system comprising:

[0141] A data acquisition module configured to acquire interaction information between the aircraft cluster and the environment, wherein the interaction information includes observation states, executed control instructions, and reward values ​​of the aircraft at multiple times;

[0142] A strategy network training module, configured to train an aircraft cluster control strategy network based on a deep neural network using interaction information between the aircraft cluster and the environment;

[0143] A control and communication module is configured to control the movement and communication topology of the aircraft cluster using the trained aircraft cluster control strategy network.

[0144] In this embodiment, preferably, the observed state includes the state information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the obstacle area th , information of the aircraft relative to the cluster target point o ta ;in,

[0145] Aircraft status information relative to the group formation Contains the formation deviation status of individual aircraft relative to the initial formation and the speed deviation state of the flight speed of the individual aircraft relative to the speed of the aircraft cluster

[0146] Information about the aircraft relative to the obstacle zone Contains the relative position of the aircraft relative to the obstacle area closest to the aircraft in the flight environment, vector ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the obstacle area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the obstacle area, k is the flag position;

[0147] Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V ; ΔR t (ΔR t ∈R 3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient;

[0148] The control instructions are obtained by inputting the observed state of the aircraft cluster at each moment into an aircraft cluster control strategy network based on a deep neural network for processing, and the control instructions include overload instructions and communication instructions.

[0149] In this embodiment, preferably, the reward value is decomposed as follows:

[0150]

[0151] In the formula, Indicates position reward; Indicates target strike reward; Represents the distance change reward; r ost It indicates a threat zone avoidance bonus; represents the relative speed reward; Indicates relative formation reward.

[0152] In this embodiment, preferably, the aircraft cluster control strategy network based on the deep neural network is constructed as follows: including a value function network and a strategy network, and the network structure is as follows:

[0153] The observation state of the aircraft is spliced ​​into an observation vector. The observation vector and the control command are spliced ​​and input into the value function network. Then, they are passed through a two-layer 128-node fully connected network and divided into two information paths: one information path is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 2 nodes and activation functions of Linear and Tanh respectively, to obtain the mean and logarithmic standard deviation of the overload command for controlling the movement of the aircraft, and the overload command of the aircraft is obtained by sampling the Gaussian distribution. Finally, the information is processed by a fully connected layer with 2 nodes and an activation function of Linear and Tanh. The information is processed by a fully connected layer with a Tanh function to limit the overload instructions to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 1 node and an activation function of Linear and Tanh respectively to obtain the mean and logarithmic standard deviation of the communication control instructions, and the communication instructions of the aircraft are obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 1 node and an activation function of Tanh to limit the communication instructions to (-1, 1).

[0154] The functions of an aircraft cluster communication topology adaptive control system according to an embodiment of the present invention can be described by the aforementioned aircraft cluster communication topology adaptive control method. Therefore, for the parts not described in detail in the system embodiment, please refer to the above method embodiment and will not be repeated here.

[0155] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments are contemplated within the scope of the invention thus described. The disclosure of the present invention is intended to be illustrative rather than restrictive of the scope of the invention, which is defined by the appended claims.

Claims

1. A method for adaptive control of aircraft cluster communication topology, characterized in that: A deep neural network is used to construct an aircraft cluster control strategy that couples the control strategy with the communication strategy, and its output includes an overload instruction for controlling the movement of the aircraft and the number of communications with neighboring aircraft; the adaptive control method includes the following steps: Acquire interaction information between the aircraft cluster and the environment, wherein the interaction information includes observation states, executed control instructions, and reward values ​​of the aircraft at multiple times; The interactive information between the aircraft cluster and the environment is used to train the aircraft cluster control strategy network based on the deep neural network; the aircraft cluster control strategy network based on the deep neural network is constructed as follows: it includes a value function network and a strategy network, and the network structure is as follows: the observation state of the aircraft is spliced ​​into an observation vector, and the observation vector and the control instruction are spliced ​​and input into the value function network, and then passed through a two-layer 128-node fully connected network to be divided into two information paths: one information path is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 2 nodes and activation functions of Linear and Tanh respectively, to obtain the overload for controlling the movement of the aircraft. The mean and logarithmic standard deviation of the instruction are obtained by sampling Gaussian distribution to obtain the overload instruction of the aircraft, and finally processed by a fully connected layer with 2 nodes and the activation function of Tanh to limit the overload instruction to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and the activation function of ReLU, and then processed by a fully connected layer with 1 node and the activation function of Linear and Tanh respectively to obtain the mean and logarithmic standard deviation of the communication control instruction, and the communication instruction of the aircraft is obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 1 node and the activation function of Tanh to limit the communication instruction to (-1, 1); The trained aircraft cluster control strategy network is used to control the movement and communication topology of the aircraft cluster.

2. The method for adaptive control of aircraft cluster communication topology according to claim 1, characterized in that: The observed state includes the state information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the obstacle area th , information of the aircraft relative to the cluster target point o ta ;in, Aircraft status information relative to the group formation Contains the formation deviation status of individual aircraft relative to the initial formation And the speed deviation state of the flight speed of the individual aircraft relative to the speed of the aircraft cluster Information about the aircraft relative to the obstacle zone Contains the relative position of the aircraft relative to the obstacle area closest to the aircraft in the flight environment, vector ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the obstacle area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the obstacle area, k is the flag position; Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V ; ΔR t (ΔR t ∈R 3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient.

3. The method for adaptive control of aircraft cluster communication topology according to claim 2, characterized in that: The control instructions are obtained by inputting the observed state of the aircraft cluster at each moment into an aircraft cluster control strategy network based on a deep neural network for processing, and the control instructions include overload instructions and communication instructions.

4. The method for adaptive control of aircraft cluster communication topology according to claim 3, characterized in that: The reward values ​​break down as follows: In the formula, Indicates position reward; Indicates target strike reward; Represents the distance change reward; r ost indicates a threat zone avoidance bonus; represents the relative speed reward; Indicates relative formation reward.

5. The method for adaptive control of aircraft cluster communication topology according to claim 4, characterized in that: The process of training a deep neural network-based aircraft swarm control strategy network includes: Initialize the cluster control strategy network parameters θ, cluster control strategy value function network parameters φ1 and φ2 and the experience pool D; Set the value function target network parameter φ targ,1 ,φ targ,2 The same as parameters φ1 and φ2, that is, φ targ,1 ←φ1,φ targ,2 ←φ2; Repeat the following steps: 1) Observe the state of the cluster simulation environment, i.e., observe the state s, and output cluster control instructions according to the control strategy; 2) Execute control instruction a in the cluster simulation environment; 3) Observe the next state s', the feedback reward r and the round end flag d; 4) Store the experience group (s, a, r, s', d) in the experience pool D; 5) If the round ends, reset the environment state; if the update cycle is reached, execute steps 6) to 10); 6) Randomly sample a set of experiences B = {(s, a, r, s', d)} from the experience pool D; 7) Calculate the true value estimate of the value function by the following formula: Where: γ represents the discount rate, γ∈[0,1]; α represents a positive coefficient; Represents state s′, action The objective value function of represents the policy network; Representation Policy Network The action of sampling; 8) Update the value function network parameters φ by minimizing the following loss function i : Where: |B| represents the number of batch data; Represents the target value function of state s and action a; 9) Update the cluster control strategy network parameters θ by minimizing the following loss function: Where: Indicates state s, action The objective value function of represents the policy network; Representation Policy Network The action of sampling; 10) Update the target network: f targ,i ←rf targ,i +(1-r)φ i ,i=1.2 Where: ρ represents a positive coefficient.

6. An aircraft cluster communication topology adaptive control system, characterized by: include: A data acquisition module configured to acquire interaction information between the aircraft cluster and the environment, wherein the interaction information includes observation states, executed control instructions, and reward values ​​of the aircraft at multiple times; A strategy network training module is configured to use the interactive information between the aircraft cluster and the environment to train an aircraft cluster control strategy network based on a deep neural network; the aircraft cluster control strategy network based on a deep neural network is constructed as follows: it includes a value function network and a strategy network, and the network structure is as follows: the observation state of the aircraft is spliced ​​into an observation vector, and the observation vector and the control instruction are spliced ​​and input into the value function network, and then passed through a two-layer 128-node fully connected network to be divided into two paths of information: one path of information is processed by a fully connected layer with 128 nodes and an activation function of ReLU, and then processed by a fully connected layer with 2 nodes and activation functions of Linear and Tanh respectively, to obtain a control flight state. The mean and logarithmic standard deviation of the overload command of the vehicle movement are obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 2 nodes and the activation function of Tanh to limit the overload command to (-1, 1); the other information is processed by a fully connected layer with 128 nodes and the activation function of ReLU, and then processed by a fully connected layer with 1 node and the activation function of Linear and Tanh respectively to obtain the mean and logarithmic standard deviation of the communication control command, and the communication command of the aircraft is obtained by sampling Gaussian distribution, and finally processed by a fully connected layer with 1 node and the activation function of Tanh to limit the communication command to (-1, 1); A control and communication module is configured to control the movement and communication topology of the aircraft cluster using the trained aircraft cluster control strategy network.

7. The aircraft cluster communication topology adaptive control system according to claim 6, characterized in that: The observed state includes the state information of the aircraft relative to the cluster formation. f , information about the aircraft relative to the obstacle area th , information of the aircraft relative to the cluster target point o ta ;in, Aircraft status information relative to the group formation Contains the formation deviation status of individual aircraft relative to the initial formation And the speed deviation state of the flight speed of the individual aircraft relative to the speed of the aircraft cluster Information about the aircraft relative to the obstacle zone Contains the relative position of the aircraft relative to the obstacle area closest to the aircraft in the flight environment, vector ΔR th (ΔR th ∈R 3 ) represents the relative position of the current aircraft relative to the obstacle area closest to the current aircraft in the flight environment, ||ΔR th || represents the vector ΔR th The modulus length, d s Indicates the safe distance of the aircraft relative to the obstacle area, k is the flag position; Information about the aircraft relative to the cluster target point Contains the position information of the aircraft relative to the cluster target point The aircraft's own flight speed information And the gravity overload vector n in the current aircraft velocity coordinate system V ; ΔR t (ΔR t ∈R 3 ) represents the relative position of the current aircraft relative to the cluster target point, R t is the normalization coefficient, V represents the current speed of the aircraft, V t is the normalization coefficient; The control instructions are obtained by inputting the observed state of the aircraft cluster at each moment into an aircraft cluster control strategy network based on a deep neural network for processing, and the control instructions include overload instructions and communication instructions.

8. The aircraft cluster communication topology adaptive control system according to claim 7, characterized in that: The reward values ​​break down as follows: In the formula, Indicates position reward; Indicates target strike reward; Represents the distance change reward; r ost indicates a threat zone avoidance bonus; represents the relative speed reward; Indicates relative formation reward.

Citation Information

Patent Citations

  • Unmanned system cluster control method based on deep reinforcement learning

    CN112068549A