A Control Method and System for the Autonomous Behavior of a UAV Cluster
By building an autonomous behavior decision model in the drone cluster and performing simulation training, the problem that drone clusters are difficult to communicate efficiently and make autonomous decisions in the environment of scarce communication resources is solved, and effective task execution under different bandwidth conditions is achieved.
Patent Information
- Application Number
- CN202210607478.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In the battlefield environment where communication resources are scarce, traditional drone clusters are difficult to achieve efficient communication and independent decision-making, and cannot effectively deal with complex and dynamic combat environments.
The autonomous behavioral decision-making model based on part of the considerable Markov decision-making process is adopted, combined with convolutional neural networks and recursive neural networks, and the decision-making of the drone cluster is simulated and trained and iteratively updated through the state evaluation function Q and the cumulative return expectation value function J to optimize the communication efficiency of the drone.
In the environment of scarce communication resources, the communication efficiency of drone clusters in the behavioral decision-making process is improved, ensuring that drones can effectively perform tasks under different bandwidth conditions.
Smart Images

Figure CN114895710B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of UAV control, and particularly relates to a control method and system for autonomous behavior of UAV swarms. Background Art
[0002] In an increasingly complex combat environment and combat missions, the human-computer interaction technology of traditional unmanned systems cannot support operators / commanders to make real-time decisions and control on swarms. UAVs need to have the ability to autonomously and intelligently complete tasks and cooperate to cope with the complexity and dynamics of the battlefield. How to achieve autonomous response to battlefield situation changes in an uncertain combat environment will be the key for UAV swarms to complete complex tasks.
[0003] At the same time, how to analogize the decision-making process of commanders or drivers to study the autonomous behavior and decision-making mechanism of UAVs is of great significance for understanding, designing, and implementing UAV autonomous systems. Communication is the basis for cooperative decision-making control of UAV swarms. It is of great significance to achieve efficient communication of UAV swarms in a battlefield environment with scarce communication resources. At present, multi-agent reinforcement learning methods are widely used in the research of autonomous cooperative strategies for UAV swarms, but most methods do not consider the impact brought by limited communication resources. Summary of the Invention
[0004] The purpose of the present invention is to provide a control method and system for autonomous behavior of UAV swarms, which can improve the communication efficiency of UAVs in the process of behavior decision-making in a battlefield environment with scarce communication resources and ensure that UAVs execute tasks under different bandwidth conditions.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] The first aspect of the present invention provides a control method for autonomous behavior of UAV swarms, including:
[0007] Receiving the observation information m sent by other UAVs i And collecting the perception information o of the surrounding environment i , to obtain the global situation information;
[0008] Inputting the global situation information into the trained autonomous behavior decision-making model to obtain the UAV action a i ; Using the perception information o i As the observation information m i+1 To other UAVs;
[0009] The training process of the autonomous behavior decision-making model includes:
[0010] Constructing an autonomous behavior decision-making model based on the partially observable Markov decision process;
[0011] Simulate and train the autonomous behavior decision-making model through a convolutional neural network, evaluate the decisions of the autonomous behavior decision-making model during the training process using the state evaluation function Q, and obtain the task reward R of the UAV cluster task , calculate the broadband reward R of the UAV cluster according to the channel capacity constraint conditions comm ;
[0012] According to the training status information of the UAV cluster, broadband reward R comm and task reward R task Establish a loss function L(θ Qi ) and a cumulative return expectation value function J(μ i );
[0013] During the training process, use the loss function L(θ Qi ) to iteratively update the state evaluation function Q, and use the policy gradient of the cumulative return expectation value function J(μ i ) to iteratively update the autonomous behavior decision-making model.
[0014] Preferably, the method of sending the perception information o i as the observation information m i+1 to other UAVs includes:
[0015] Set a sequence number for the routing of each UAV through the DSDV protocol, and propagate the observation information m i+1 in the UAV cluster along a directed tree network without intersections;
[0016] The channel capacity constraint conditions include: the link between UAVs is a unidirectional link, the maximum number of times each UAV sends the observation information m at the same moment interval is 1, and the time delay from the sending of the observation information m to the last UAV in the UAV cluster receiving the observation information m i+1 is less than one moment interval.
[0017] Preferably, calculate the broadband reward R of the UAV cluster according to the channel capacity constraint conditions comm , and the expression formula is:
[0018]
[0019] In the formula, g comm,i represents the communication resource allocation amount of the i-th UAV, g comm represents the communication resource allocation amount of the UAV cluster, R comm,i represents the broadband reward of the i-th UAV; k comm represents the number of symbol discrete levels.
[0020]
[0021]
[0022] In the formula, B represents the channel bandwidth between UAVs; N represents the number of UAVs in the UAV cluster; L represents the number of symbols in the observation information; N b represents the number of bits occupied by each symbol; n m represents the number of UAVs sending the observation information.
[0023] Preferably, the state training information of the UAV cluster, the broadband reward R comm and the task reward R task are trained by the recurrent neural network LSTM and stored in the experience pool D;
[0024] The state training information of the UAV cluster includes the self-state s of each UAV in the UAV cluster i , action a i , perception information o i , observation information m i , the parameter θ of the state evaluation function Q Q and the parameter θ of the autonomous behavior decision-making model μ ;
[0025] The historical state of the parameter θ of the state evaluation function Q in the experience pool D Q is denoted as h Q ; The historical state of the parameter θ of the autonomous behavior decision-making model in the experience pool D μ is denoted as h μ .
[0026] Preferably, the method for simulating and training the autonomous behavior decision-making model by a convolutional neural network includes: using the Recurrent Actor-Critic neural network to simulate and train the autonomous behavior decision-making model, the Recurrent Actor sub-neural network simulates the autonomous behavior decision-making model; the Recurrent Critic network simulates the state evaluation function Q.
[0027] Preferably, the method for evaluating the decision of the autonomous behavior decision-making model during the training process by using the state evaluation function Q includes:
[0028] By inputting the global situation information into the autonomous behavior decision-making model, the decision of the UAV action a i is obtained;
[0029] The UAV action a i is executed through the motion model; the state evaluation function Q evaluates according to the execution result;
[0030] The expression formula of the motion model is:
[0031]
[0032] In the formula, x i ′ represents the horizontal coordinate of the drone's own state s i ′ after executing action a i ′; y i ′ represents the vertical coordinate of the drone's own state s i ′ after executing action a i ′; x i represents the horizontal coordinate of the drone's own state s i before executing action a i ; y i represents the vertical coordinate of the drone's own state s i before executing action a i ; v i represents the speed of the drone when executing action a i ; represents the heading angle of the drone when executing action a i .
[0033] Preferably, the method for iteratively updating the state evaluation function Q using the loss function L(θ Qi ) includes:
[0034] Randomly extract T samples from the experience pool D; the samples include the drone's own state s at the j-th moment j , the action a of the drone at the j-th moment j , the drone's own state s j ′ after executing action a at the j-th moment j and the reward value of the i-th drone at the j-th moment
[0035] Calculate the loss values of the T samples through the loss function L(θ Qi ), and iteratively update the state evaluation function Q according to the loss values;
[0036] The expression formula of the loss function L(θ Qi ) is
[0037]
[0038]
[0039] In the formula, h′ μ represents the parameters θ of the autonomous behavior decision-making model in the updated experience pool D μ historical state; h′ Q represents the parameters θ of the state evaluation function Q in the updated experience pool D Q historical state; Denoted as the reward value of the $i$-th UAV at time $j$; Denoted as the state evaluation function $Q$ for evaluating the task execution of the $i$-th UAV $\mu$ i ; $\mu$ i $(\cdot)$ is denoted as the task executed by the $i$-th UAV; $\gamma$ is denoted as the discount factor, $\gamma\in[0,1]$.
[0040] Preferably, the expression formula of the cumulative return expectation function $J(\mu$ i ) is
[0041]
[0042] $R$ i $=(R$ comm,i $+R$ task,i )
[0043] In the formula, $E$ is denoted as the reward value weight $R$ i , and $t$ is denoted as the number of training times of the autonomous behavior decision model.
[0044] The second aspect of the present invention provides a control system for the autonomous behavior of a UAV cluster, including:
[0045] A global situation information acquisition module, which receives the observation information $m$ sent by other UAVs i and acquires the perception information $o$ of the surrounding environment i , and obtains the global situation information;
[0046] A UAV decision module; used to input the global situation information into the trained autonomous behavior decision model to obtain the UAV action $a$ i ; use the perception information $o$ i as the observation information $m$ i+1 to other UAVs;
[0047] A model construction module, which constructs an autonomous behavior decision model based on the partially observable Markov decision process;
[0048] A model training module, used to simulate and train the autonomous behavior decision model through a convolutional neural network. During the training process, the state evaluation function $Q$ is iteratively updated using the loss function $L(\theta$ Qi ), and the policy gradient of the cumulative return expectation function $J(\mu$ i ) is used to iteratively update the autonomous behavior decision model;
[0049] A training evaluation module, which uses the state evaluation function $Q$ to evaluate the decision of the autonomous behavior decision model during the training process to obtain the task reward $R$ of the UAV cluster task , and calculates the broadband reward $R$ of the UAV cluster according to the channel capacity constraint condition comm; Based on the training status information of the UAV swarm, the broadband reward R comm and the mission reward R task to establish the loss function L(θ Qi ) and the cumulative return expectation function J(μ i ).
[0050] The third aspect of the present invention provides a computer-readable storage medium, characterized in that a computer program is stored thereon, and when the program is executed by a processor, the steps of the control method are implemented.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] The present invention uses the state evaluation function Q to evaluate the decisions of the autonomous behavior decision model during the training process, and obtains the mission reward R of the UAV swarm task , calculates the broadband reward R of the UAV swarm according to the channel capacity constraint conditions comm ; uses the mission reward R task and the broadband reward R comm to iteratively update the state evaluation function Q and the autonomous behavior decision model. After training, the autonomous behavior decision model enables the UAV swarm to improve the communication efficiency of the UAVs during the behavior decision process in a battlefield environment with scarce communication resources, and ensures that the UAVs can execute tasks under different bandwidth conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the structural diagram of the UAV autonomous behavior decision model provided by the embodiment of the present invention;
[0054] Figure 2 is the structural diagram of the UAV motion model provided by the embodiment of the present invention;
[0055] Figure 3 is the path diagram of the UAV swarm communication provided by the embodiment of the present invention;
[0056] Figure 4 is the learning curve diagram of the UAVs under different bandwidth conditions provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be used to limit the protection scope of the present invention.
[0058] Embodiment 1
[0059] As Figure 1 shown, this embodiment provides a control method for the autonomous behavior of a UAV swarm, including:
[0060] Receiving the observation information m sent by other UAVsi and collect the perception information of the surrounding environment. i , and obtain the global situation information;
[0061] Input the global situation information into the trained autonomous behavior decision-making model to obtain the UAV action a i ; the autonomous behavior decision-making model of UAV i based on communication can be expressed as μ i (a i |o i ,m -i );
[0062] Take the perception information o i as the observation information m i+1 The method of sending it to other UAVs includes:
[0063] As Figure 3 shown, set the sequence number for the routing of each UAV through the DSDV protocol. By setting the sequence number for each routing, the generation of routing loops is avoided, and the information selects the transmission path according to the number of links passed; according to the channel capacity constraint conditions, the observation information m i+1 is propagated in the UAV cluster along a directed acyclic tree network without crossing; the channel capacity constraint conditions include: the link between UAVs is a one-way link, the maximum number of times each UAV sends the observation information m at the same moment is 1, and the time delay from the observation information m being sent to the last UAV in the UAV cluster receiving the observation information m i+1 is less than one time slot.
[0064] The UAVs adopt the Frequency Division Multiple Access (FDMA) protocol, and evenly divide the wireless channel resources into several sub-channels according to the number of links required at the current moment. Each physical link is assigned a sub-channel.
[0065] Table 1 is the routing protocol of the communication network
[0066]
[0067] The training process of the autonomous behavior decision-making model includes:
[0068] Construct an autonomous behavior decision-making model based on the partially observable Markov decision process; the method of simulating and training the autonomous behavior decision-making model through a convolutional neural network includes:
[0069] The Recurrent Actor-Critic neural network is used to simulate and train the autonomous behavior decision-making model. The Recurrent Actor sub-neural network simulates the autonomous behavior decision-making model; the Recurrent Critic network simulates the state evaluation function Q.
[0070] The state training information of the UAV cluster is trained through the recurrent neural network LSTM, and the broadband reward R comm and the task reward R task are stored in the experience pool D;
[0071] The state training information of the UAV cluster includes the self-state s of each UAV in the UAV cluster i , the action a i , the perception information o i , the observation information m i , the parameter θ of the state evaluation function Q Q and the parameter θ of the autonomous behavior decision-making model μ ;
[0072] The historical state of the parameter θ of the state evaluation function Q in the experience pool D Q is denoted as h Q ; the historical state of the parameter θ of the autonomous behavior decision-making model in the experience pool D μ is denoted as h μ .
[0073] The method for evaluating the decision-making of the autonomous behavior decision-making model during the training process by using the state evaluation function Q includes:
[0074] As Figure 2 shown, by inputting the global situation information into the autonomous behavior decision-making model, the decision-making of the UAV action a i is obtained; the UAV action a i is executed through the motion model; the state evaluation function Q evaluates according to the execution result to obtain the task reward R of the UAV cluster task .
[0075] Assume that the flight altitude of all UAVs is constant, and the self-state of UAV i is represented by s i = [x i , y i , and the expression formula of the motion model is:
[0076]
[0077] In the formula, x i ′ represents the horizontal coordinate of the self-state s i ′ of the UAV after executing the action a i ′; y i' represents the self - state s after the UAV executes action a i after the UAV executes action a i The longitudinal coordinate of'; x i represents the self - state s before the UAV executes action a i before the UAV executes action a i The lateral coordinate of; y i represents the self - state s before the UAV executes action a i before the UAV executes action a i The longitudinal coordinate of; v i represents the speed of the UAV after executing action a i The speed of; represents the heading angle of the UAV after executing action a i The heading angle of.
[0078] The method for calculating the broadband reward R of the UAV cluster according to the channel capacity constraint conditions includes: comm The method includes:
[0079] The broadband reward R comm The expression formula is:
[0080]
[0081]
[0082]
[0083] In the formula, g comm,i represents the communication resource allocation amount of the i - th UAV, g comm represents the communication resource allocation amount of the UAV cluster, R comm,i represents the broadband reward of the i - th UAV; k comm represents the number of symbol discrete levels; B represents the channel bandwidth between UAVs; N represents the number of UAVs in the UAV cluster; L represents the number of symbols in the observation information; N b represents the number of bits occupied by each symbol; n m represents the number of UAVs sending the observation information.
[0084] According to the training state information of the UAV cluster, the broadband reward R comm and the task reward R task establish the loss function L(θ Qi ) and the cumulative return expectation function J(μ i );
[0085] The method for iteratively updating the state evaluation function Q using the loss function L(θ Qi ) includes:
[0086] Randomly select T samples from the experience pool D; the samples include the self-state s of the UAV at the j-th moment j , the action a of the UAV at the j-th moment j , the self-state s j of the UAV after executing the action a at the j-th moment j ′ and the reward value of the i-th UAV at the j-th moment
[0087] Calculate the loss values of the T samples through the loss function L(θ Qi ), and iteratively update the state evaluation function Q according to the loss values;
[0088] The expression formula of the loss function L(θ Qi ) is
[0089]
[0090]
[0091] In the formula, h′ μ represents the parameter θ μ of the autonomous behavior decision model in the updated experience pool D Q historical state; h′ Q represents the parameter θ of the state evaluation function Q in the updated experience pool D historical state; i represents the reward value of the i-th UAV at the j-th moment; i represents the state evaluation function Q for evaluating the task μ
[0092] of the i-th UAV; μ i (·) represents the task executed by the i-th UAV; γ represents the discount factor, γ ∈ [0, 1].
[0093]
[0094] R i =(R comm,i +R task,i )
[0095] In the formula, E represents the reward value weight R i , and t represents the number of training times of the autonomous behavior decision model.
[0096] During the training process, use the loss function L(θ Qi ) to iteratively update the state evaluation function Q, and use the policy gradient of the cumulative return expected value function J(μ i ) to iteratively update the autonomous behavior decision model.
[0097] The present invention simulates the air confrontation of unmanned aerial vehicles (UAVs) in a bandwidth - limited combat scenario in the Swarmflow simulation platform for unmanned combat built by the research group. The simulation environment simulates a real airspace combat environment based on the satellite map of Dadong Mountain, and selects an airspace of 2000m×2000m as the engagement area. In this airspace, the UAV groups of both sides conduct confrontation with a force ratio of 2:4. The UAVs make decisions and take actions simultaneously at discrete time steps.
[0098] The air confrontation task is simplified to a collaborative attack of an adversarial nature. The combat objective of both sides is to obtain rewards by attacking the other side through collaboration as much as possible. It is assumed that the UAVs can visually estimate the azimuth angle between the enemy aircraft and themselves. If more than two UAVs of one side encounter one enemy UAV, the UAVs participating in the attack will obtain rewards, and the besieged enemy aircraft will be punished, and vice versa. At the same time, the closer the heading angle of the UAV is to the azimuth angle of the target enemy aircraft, the smaller the negative reward value obtained.
[0099] Due to the limited available channel bandwidth on the battlefield, the UAVs of both sides need to adopt an efficient communication method to avoid frequent communication. The available bandwidth size B for both sides on the battlefield is set, the discount factor γ is set to 0.9, the simulation time step is set to 0.1, the batch sample number is set to 64; the number of training rounds is set to 12000; the maximum number of simulation time steps per round is set to 3000000; to verify that the proposed method can reduce the bandwidth consumption while maintaining the collaborative ability of the UAVs, experiments were repeated in scenarios with different bandwidth sizes, as Figure 4 shown. The results show that the smaller the bandwidth, the slower the UAV strategy learning speed, and the smaller the reward value in the early stage of training.
[0100] Embodiment 2
[0101] This embodiment provides a control system for the autonomous behavior of a UAV swarm. The control system provided in this embodiment can be applied to the control method described in Embodiment 1. The control system includes:
[0102] A global situation information acquisition module, which receives the observation information m sent by other UAVs i and acquires the perception information o of the surrounding environment i , and obtains the global situation information;
[0103] A UAV decision - making module; used to input the global situation information into the trained autonomous behavior decision model to obtain the UAV action a i ; and use the perception information o i as the observation information m i+1 to other UAVs;
[0104] A model construction module, which constructs an autonomous behavior decision model based on the partially observable Markov decision process;
[0105] A model training module for simulating and training an autonomous behavior decision-making model through a convolutional neural network. During the training process, the loss function L(θ Qi ) is used to iteratively update the state evaluation function Q, and the policy gradient of the cumulative return expected value function J(μ i ) is used to iteratively update the autonomous behavior decision-making model;
[0106] A training evaluation module that uses the state evaluation function Q to evaluate the decisions of the autonomous behavior decision-making model during the training process to obtain the task reward R of the UAV cluster task , and calculates the broadband reward R of the UAV cluster according to the channel capacity constraint conditions comm ; according to the training status information of the UAV cluster, the broadband reward R comm and the task reward R task to establish the loss function L(θ Qi ) and the cumulative return expected value function J(μ i ).
[0107] Example 3
[0108] This embodiment provides a computer-readable storage medium, which is characterized in that a computer program is stored thereon, and when the program is executed by a processor, the steps of the control method for the autonomous behavior of the UAV cluster described in Example 1 are implemented.
[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0111] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0113] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster, characterized in that, comprising: Receive the observation information m sent by other drones i and collect the perception information o of the surrounding environment i , and obtain the global situation information; Input the global situation information into the trained autonomous behavior decision-making model to obtain the drone action a i ; Take the perception information o i as the observation information m i+1 to other drones; The training process of the autonomous behavior decision-making model includes: Constructing an autonomous behavior decision-making model based on the partially observable Markov decision process; The autonomous behavior decision-making model is simulated and trained through a convolutional neural network, and the state evaluation function Q is used to evaluate the decisions of the autonomous behavior decision-making model during the training process to obtain the task rewards of the UAV cluster , and the broadband rewards of the UAV cluster are calculated according to the channel capacity constraint conditions ; The expression formula is: ; ; ; In the formula, represents the communication resource allocation amount of the i-th UAV, represents the communication resource allocation amount of the UAV cluster, represents the broadband reward of the i-th UAV; represents the number of symbol discrete levels; B represents the channel bandwidth between UAVs; N represents the number of UAVs in the UAV cluster; L represents the number of symbols in the observation information; represents the number of bits occupied by each symbol; Based on the training status information of the UAV swarm, broadband rewards and mission rewards establish a loss function and an expected cumulative return function ; During the training process, use the loss function to iteratively update the state evaluation function Q, and use the policy gradient of the expected cumulative return function to iteratively update the autonomous behavior decision-making model.
2. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 1, characterized in that, Take the perception information o i As the observation information m i+1 The method of sending it to other drones includes: Set the sequence number for the routing of each UAV through the DSDV protocol, and transmit the observation information m in the UAV cluster along a directed tree network without intersections; the channel capacity constraint conditions include: the links between UAVs are one-way links, the maximum number of times each UAV can send the observation information m at the same moment with a gap is 1, and the time delay from the transmission of the observation information m to the last UAV in the UAV cluster receiving the observation information m i+1 is less than one moment gap. i+1 3. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 2, characterized in that, Training information on the state of the UAV swarm, broadband rewards, and mission rewards are memorized through the recurrent neural network LSTM and stored in the experience pool D; the state training information of the UAV swarm includes the self-state s of each UAV in the UAV swarm , action a i , perception information o i , observation information m i , the parameters of the state evaluation function Q i , and the parameters of the autonomous behavior decision-making model ; the historical state of the parameters of the state evaluation function Q in the experience pool D is denoted as ; the historical state of the parameters of the autonomous behavior decision-making model in the experience pool D is denoted as ; ; the historical state of the parameters of the autonomous behavior decision-making model in the experience pool D is denoted as ; .
4. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 3, characterized in that, The method for simulating and training the autonomous behavior decision-making model through a convolutional neural network includes: Using the Recurrent Actor-Critic neural network to simulate and train the autonomous behavior decision-making model, where the Recurrent Actor sub-neural network simulates the autonomous behavior decision-making model; the Recurrent Critic network simulates the state evaluation function Q.
5. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 3, characterized in that, The method for evaluating the decision-making of the autonomous behavior decision-making model during the training process using the state evaluation function Q includes: By inputting the global situation information into the autonomous behavior decision-making model, the decision of the UAV action a is obtained. i The UAV action a is executed through the motion model. i The state evaluation function Q evaluates according to the execution result. The expression formula of the motion model is: ; In the formula, represents the horizontal coordinate of the UAV's own state i after performing action a; ; represents the vertical coordinate of the UAV's own state i after performing action a; x ; i represents the horizontal coordinate of the UAV's own state s i before performing action a; y i ; i represents the vertical coordinate of the UAV's own state s i before performing action a; i ; represents the speed of the UAV i after performing action a; represents the heading angle of the UAV i after performing action a.
6. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 5, characterized in that, Using a loss function The method for iteratively updating the state evaluation function Q includes: Randomly select T samples from the experience pool D; the samples include the self-state of the drone at the j-th moment , the action of the drone at the j-th moment , the self-state of the drone after performing the action at the j-th moment and the reward value of the i-th drone at the j-th moment ; ; Through the loss function Calculate the loss values of T samples, and iteratively update the state evaluation function Q according to the loss values; The loss function The expression formula is ; ; In the formula, represents the parameters of the autonomous behavior decision-making model in the updated experience pool D historical state; represents the parameters of the state evaluation function Q in the updated experience pool D historical state; represents the reward value of the i-th UAV at the j-th moment; represents the evaluation of the i-th UAV performing the task state evaluation function Q; represents the task performed by the i-th UAV; represents the discount factor, .
7. The control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster according to claim 6, characterized in that, Cumulative return expected value function The expression formula is as follows: ; ; In the formula, represents the expected value of the reward value , and t represents the number of training times of the autonomous behavior decision-making model.
8. A control system for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster, characterized in that, comprising: The global situation information acquisition module receives the observation information m sent by other UAVs i and acquires the perception information o of the surrounding environment i , and obtains the global situation information; A UAV decision-making module; For inputting global situation information into a trained autonomous behavior decision-making model to obtain the action a of the UAV i ; taking the perception information o i as the observation information m i+1 to other UAVs; A model construction module that constructs an autonomous behavior decision-making model based on the partially observable Markov decision process; A model training module for simulating and training an autonomous behavior decision-making model through a convolutional neural network. During the training process, a loss function is used to iteratively update the state evaluation function Q, and the policy gradient of the cumulative return expected value function is used to iteratively update the autonomous behavior decision-making model; The training evaluation module uses the state evaluation function Q to evaluate the decisions of the autonomous behavior decision-making model during the training process, and obtains the task rewards of the UAV cluster , calculates the broadband rewards of the UAV cluster according to the channel capacity constraint conditions ; Based on the training status information, broadband rewards and task rewards to establish a loss function and the cumulative return expectation value function ; The training and evaluation module calculates the broadband reward of the UAV cluster according to the channel capacity constraint condition , and the expression formula is: ; ; ; In the formula, represents the communication resource allocation amount of the i-th unmanned aerial vehicle (UAV), represents the communication resource allocation amount of the UAV cluster, represents the broadband reward of the i-th UAV; represents the number of symbol discrete levels; B represents the channel bandwidth between UAVs; N represents the number of UAVs in the UAV cluster; L represents the number of symbols in the observation information; represents the number of bits occupied by each symbol.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the steps of the control method for the autonomous behavior of an unmanned aerial vehicle (UAV) cluster described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-agent group cooperation strategy automatic generation method
CN112488310A
Unmanned aerial vehicle cooperative control training method and system based on multi-agent reinforcement learning
CN113900445A