An Unmanned Aerial Vehicle Cluster Formation Method, Device, Equipment and Storage Medium

By using Markov decision-making process and Actor-Critic network model in the drone cluster formation, combining the maximum entropy model-free algorithm and formation reward function, the problem of difficulty in formation adjustment of the drone cluster in a changing environment is solved, and higher stability and task success rate are achieved.

CN119088076BActive Publication Date: 2025-05-27CHINA ORDNANCE EQUIP GRP AUTOMATION RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411149224.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-05-27
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The existing drone cluster formation method cannot adjust the formation in time when facing a changing environment, resulting in mission failure.

Method used

By establishing a drone motion model during Markov decision-making, and using the Actor-Critic network model combined with the maximum entropy model-free algorithm for training, dynamic obstacles and formation reward functions are set, the drone cluster formation is maintained.

Benefits of technology

It enhances the ability of drone clusters to respond in the face of environmental changes, and improves the stability of the formation and mission success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119088076B_ABST
    Figure CN119088076B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and storage medium for unmanned aerial vehicle cluster formation. This method has the ability to respond to environmental changes. At the same time, for the formation problem of the formation, a reward function is designed to improve the formation stability ability. The SAC algorithm is adopted to avoid the environment from repeatedly falling into local optima, enhancing the system's ability to respond to external environmental changes. The formation is carried out using the reward function, achieving the stability of the formation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) swarms, and particularly to a method, device, equipment, and storage medium for UAV swarm formation based on deep reinforcement learning. Background Art

[0002] In the execution of UAV formation tasks, formation control is an important research content. Existing formation methods mainly complete control through hierarchical dimensionality reduction. Commonly used algorithms are generally genetic algorithms, simulated annealing algorithms, ant colony optimization algorithms, etc.

[0003] However, the formation methods in the prior art rely on pre-acquired environmental information. When the UAV swarm faces a changing environment, it cannot adjust the formation in a timely manner, resulting in task failure. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method, device, equipment, and storage medium for UAV swarm formation to overcome or at least partially solve the above problems. Based on the establishment of a model in the Markov decision process, by setting dynamic obstacles, the ability of the system to face a changing environment is enhanced, and by setting a formation reward function, the formation shape maintenance of the UAV swarm is achieved.

[0005] The present invention provides the following solutions:

[0006] A method for UAV swarm formation, comprising:

[0007] Establishing a UAV motion model based on the Markov decision process; the UAV motion model includes environmental information, observation information, and an action set;

[0008] Bringing the information of the UAV motion model into the Actor-Critic network model with a centralized training - distributed decision-making architecture, and updating the network parameters of the Actor network and the Critic network according to the maximum entropy model-free algorithm;

[0009] Bringing the obtained network parameters into the reward function of the reward mechanism to obtain a reward value;

[0010] The Actor network and the Critic network perform iterative calculations until the reward value of the reward function converges, and the model ends the calculation to obtain the action parameters that each UAV needs to execute.

[0011] Preferably: the environmental information includes the two-dimensional positions of all UAVs, the headings of all UAVs, the speeds of all UAVs, and the target point positions of all UAVs;

[0012] The observation information includes the two-dimensional position of the current UAV; the heading of the current UAV; the speed of the current UAV; the distance between the current UAV and nearby UAVs;

[0013] The action set includes moving left, moving right, moving forward, moving backward, and staying still.

[0014] Preferably: The design principle of the reward mechanism is that the closer the UAV is to the target point, the greater the reward value; the closer the distance between two UAVs is to the optimal distance, the greater the reward value. The reward mechanism includes a reward function for guiding the UAV's heading, a reward function for the distance between the target point and the destination, a reward function for guiding the UAV to avoid obstacles, and a sparse reward function for the UAV arriving near the target point.

[0015] Preferably: The reward function for guiding the UAV's heading is expressed by the following formula:

[0016]

[0017] In the formula: θ represents the flight angle of the UAV with the target point as the orientation;

[0018] The reward function for the distance between the target point and the destination is expressed by the following formula:

[0019]

[0020] In the formula: d_targe represents the Euclidean distance between the UAV and the target point at the current moment, and d_pre represents the Euclidean distance between the UAV and the target point in the previous step;

[0021] The reward function for guiding the UAV to avoid obstacles is expressed by the following formula:

[0022]

[0023] The reward function for guiding the UAV to maintain a formation with its nearby UAVs is expressed by the following formula:

[0024]

[0025] In the formula: d min represents the safety distance between UAVs, d target represents the optimal distance maintained between UAVs, d max represents the maximum distance between UAVs, d ij represents the distance between the UAV and its nearby UAVs, d err represents the set error threshold between two UAVs; this reward mechanism is for maintaining the UAV formation;

[0026] The sparse reward function for the UAV arriving near the target point is expressed by the following formula:

[0027]

[0028] Wherein, δd represents the set distance threshold;

[0029] The overall reward function based on the UAV is as follows:

[0030] R i = k 1 r 1 + k 2 r 2 + k 3 r 3 + k 4 r 4 + k 5 r 5

[0031] Wherein, k 1 , k 2 , k 3 , k 4 , k 5 are weighting coefficients, and r 1 , r 2 , r 3 , r 4 , r 5 are reward functions.

[0032] Preferably: The parameters of the Actor network and the Critic network are updated according to the maximum entropy model-free algorithm. The Actor network inputs the observation information of the UAV and outputs the action set of the UAV. The Critic network inputs the observation information and the action set of the UAV and outputs the evaluated Q value.

[0033] Preferably: For the maximum entropy model-free algorithm, a centralized training and distributed execution architecture is adopted:

[0034] The calculation formula for the Critic network to input the observation information and the action set of the UAV and output the evaluated Q value is as follows:

[0035]

[0036] Wherein, i represents the i-th UAV, i = 1, 2, 3,..., N, j represents the number of Critic networks, j = 1, 2, r is the reward value, γ is the discount factor, d is the round termination signal, s ′ represents the environmental state of the current UAV at the next moment, represents the action set of other UAVs at the next moment, θ represents the network parameters of the Actor, represents the network parameters of the Critic, π θ(a|s) is the probability that the next - moment policy takes action a in state s. It represents the target Q - value calculated when the UAV executes action a in the environmental state s.

[0037] Parameter update is performed according to the gradient - descent method, and its loss - function calculation formula is:

[0038]

[0039] In the formula, a i ={…, a N-1 ,…}, It represents the Q - value calculated when the UAV adopts action a in the environment s.

[0040] The observation information of the UAV is input into the Actor network, and the action set of the UAV is output. Parameter update is performed according to the gradient - ascent method, and its loss - function calculation formula is:

[0041]

[0042] a other represents the action set of the UAV at the current moment.

[0043] The entropy - regularization coefficient α is updated using the gradient - descent method:

[0044]

[0045] In the formula, H represents the initial value of entropy.

[0046] Soft - update the parameters of the target network:

[0047]

[0048] In the formula, τ represents the soft - update parameter.

[0049] A UAV swarm formation device is used to execute the above - mentioned UAV swarm formation method. The device includes:

[0050] A motion - model establishment unit is used to establish a UAV motion model based on the Markov decision process; the UAV motion model includes environmental information, observation information, and an action set.

[0051] A parameter - input unit is used to input the information of the UAV motion model into the Actor - Critic network model with a centralized - training and distributed - decision - making architecture, and update the network parameters of the Actor network and the Critic network according to the maximum - entropy model - free algorithm.

[0052] A reward - function calculation unit is used to input the obtained network parameters into the reward function of the reward mechanism to obtain a reward value.

[0053] An action parameter acquisition unit is used to perform loop calculations on the Actor network and the Critic network until the reward value of the reward function converges, the model ends the calculation, and the actions required for each drone are obtained.

[0054] An unmanned aerial vehicle (UAV) cluster formation device, the device includes a processor and a memory:

[0055] The memory is used to store program code and transmit the program code to the processor;

[0056] The processor is used to execute the above-mentioned UAV cluster formation method according to the instructions in the program code.

[0057] A computer-readable storage medium, the computer-readable storage medium is used to store program code, and the program code is used to execute the above-mentioned UAV cluster formation method.

[0058] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0059] A UAV cluster formation method, device, equipment and storage medium provided by an embodiment of the present application have the ability to respond to environmental changes. At the same time, for the formation problem of the formation, a reward function is designed to improve the formation stability ability. The SAC algorithm is adopted to avoid the environment from repeatedly falling into local optima and enhance the system's ability to respond to external environmental changes. The formation is carried out by using the reward function, and the stability of the formation is achieved.

[0060] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0062] Figure 1 is a flowchart of a UAV cluster formation method based on reinforcement learning provided by an embodiment of the present invention;

[0063] Figure 2 is a structural diagram of a model-free deep learning algorithm with maximum entropy in an embodiment of the present invention;

[0064] Figure 3 is a schematic diagram of a UAV cluster formation device provided by an embodiment of the present invention;

[0065] Figure 4 It is a schematic diagram of a drone swarm formation device provided by an embodiment of the present invention. Specific implementation manner

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.

[0067] See Figure 1 - Figure 2 , a drone swarm formation method provided by an embodiment of the present invention, the method may include:

[0068] S101: Establish a drone motion model based on the Markov decision process; the drone motion model includes environmental information, observation information, and action set; specifically, in the implementation, the environmental information provided by the embodiments of the present application may include the two-dimensional positions of all drones; the headings of all drones; the speeds of all drones; the target point positions of all drones;

[0069] The observation information includes the two-dimensional position of the current drone; the heading of the current drone; the speed of the current drone; the distance between the current drone and nearby drones;

[0070] The action set includes moving left, moving right, moving forward, moving backward, and staying.

[0071] S102: Bring the drone motion model information in S101 into the Actor-Critic network model and adopt a centralized training-distributed decision-making architecture, and update the network parameters of the Actor network and the Critic network according to the maximum entropy model-free algorithm; the algorithm includes 1 Actor network, one Actor target network, and selects the smaller value among the two Critics as the target Q value; specifically, the Critic network inputs the observation information and action set of the drone and outputs the evaluation Q value. The calculation formula is as follows:

[0072]

[0073] In the formula, i represents the i-th drone, i = 1, 2, 3,..., N, j represents the number of Critic networks, j = 1, 2, r is the reward value, γ is the discount factor, d is the episode termination signal, s ′ represents the environmental state of the current drone at the next moment, represents the action set of other drones at the next moment, θ represents the network parameters of the Actor, Denote the network parameters of the Critic as π θ (a|s) is the probability of taking action a in state s by the next - moment policy, Denote the target Q - value calculated when the UAV executes action a in the environmental state s;

[0074] According to the gradient - descent method for parameter update, its loss - function calculation formula is:

[0075]

[0076] In the formula, a i ={…, a N-1 ,…}, Denote the Q - value calculated when the UAV adopts and executes action a in the environment s;

[0077] Input the observation information of the UAV into the Actor network, output the action set of the UAV, and update the parameters according to the gradient - ascent method. Its loss - function calculation formula is:

[0078]

[0079] a other Denote the action set of the UAV at the current moment;

[0080] Update the entropy - regularization coefficient α using the gradient - descent method:

[0081]

[0082] In the formula, H represents the initial value of entropy;

[0083] Soft - update the target - network parameters:

[0084]

[0085] In the formula, τ represents the soft - update parameter.

[0086] S103: Substitute the network parameters obtained in step S102 into the reward function of the said reward mechanism to obtain the reward value; The design principle of the reward mechanism is that the closer the UAV is to the target point, the greater the reward value; The closer the distance between two UAVs is to the optimal distance, the greater the reward value. It mainly includes the reward function for guiding the UAV's heading, the reward function for the distance between the target point and the destination, the reward function for guiding the UAV to avoid obstacles, and the sparse reward function when the UAV arrives near the target point. The specific implementation is as follows:

[0087] The reward function for guiding the UAV's heading is expressed by the following formula:

[0088]

[0089] where: θ represents the flight angle of the UAV guided by the target point;

[0090] The reward function guiding the distance of the target point from the destination is expressed by the following formula:

[0091]

[0092] where: d_target represents the Euclidean distance between the UAV and the target point at the current moment, and d_pre represents the Euclidean distance between the UAV and the target point in the previous step. When the UAV moves towards the target point, a reward is given; when the UAV moves away from the target point or remains stationary, a penalty is given;

[0093] The reward function guiding the UAV to avoid obstacles is expressed by the following formula:

[0094]

[0095] The reward function guiding the UAV to maintain a formation with nearby UAVs is expressed by the following formula:

[0096]

[0097] where: d min represents the safe distance between UAVs, d target represents the optimal distance maintained between UAVs, d max represents the maximum distance between UAVs, d ij represents the distance between the UAV and nearby UAVs, d err represents the set error threshold between two UAVs. This reward mechanism is to maintain the UAV formation.

[0098] The sparse reward function for the UAV to reach near the target point is expressed by the following formula:

[0099]

[0100] where δd represents the set distance threshold. If this reward is triggered, the UAV ends the current round of training.

[0101] In summary, the reward function based on by the UAV is as follows:

[0102] R i = k 1 r 1 + k 2 r 2 + k 3 r 3 + k 4 r 4 + k 5 r 5

[0103] In the formula, k 1 , k 2 , k 3 , k 4 , k 5 are weighting coefficients, and r 1 , r 2 , r 3 , r 4 , r 5 are reward functions. Every time the state of the UAV changes, a guiding reward will be obtained. r 4 is a sparsity reward. When triggered, the UAV will end the training of this round.

[0104] S104: The Actor network and the Critic network are calculated iteratively until convergence, and each UAV outputs the action parameters to be executed according to the policy function.

[0105] In summary, the UAV swarm formation method provided by this application has the ability to respond to environmental changes. At the same time, for the problem of autonomous change of the formation, the method adopts a reward function to improve the formation ability. The SAC algorithm is used to avoid the environment from repeatedly falling into local optima, enhancing the system's ability to respond to external environmental changes. The use of the reward function makes the UAV swarm formation more stable and improves the formation ability.

[0106] See Figure 3 , this embodiment of the application can also provide a UAV swarm formation device, as shown in Figure 4 , for executing the above UAV swarm formation method. The device may include:

[0107] A motion model establishment unit 301, configured to establish a UAV motion model based on a Markov decision process; the UAV motion model includes environmental information, observation information, and an action set;

[0108] A parameter input unit 302, configured to bring the information of the UAV motion model into the Actor-Critic network model, adopt a centralized training-distributed decision-making architecture, and update the network parameters of the Actor network and the Critic network according to the maximum entropy model-free algorithm;

[0109] A reward function calculation unit 303, configured to bring the obtained network parameters into the reward function of the reward mechanism to obtain a reward value;

[0110] An action parameter acquisition unit 304, configured to calculate the Actor network and the Critic network iteratively until the reward value of the reward function converges, the model ends the calculation, and obtain the actions to be executed by each UAV.

[0111] An embodiment of the present application may further provide a drone swarm formation device, and the device includes a processor and a memory:

[0112] The memory is used to store program codes and transmit the program codes to the processor;

[0113] The processor is used to execute the steps of the above-mentioned drone swarm formation method according to the instructions in the program codes.

[0114] As Figure 4 shown, a drone swarm formation device provided by an embodiment of the present application may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all complete mutual communication through the communication bus 13.

[0115] In an embodiment of the present application, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.

[0116] The processor 10 may call the program stored in the memory 11. Specifically, the processor 10 may execute the operations in the embodiment of the drone swarm formation method.

[0117] The memory 11 is used to store one or more programs. The program may include program codes, and the program codes include computer operation instructions. In an embodiment of the present application, the memory 11 stores at least programs for implementing the following functions:

[0118] Establish a drone motion model based on the Markov decision process; the drone motion model includes environmental information, observation information, and an action set;

[0119] Bring the information of the drone motion model into the Actor-Critic network model and adopt a centralized training-distributed decision-making architecture to update the network parameters of the Actor network and the Critic network according to the maximum entropy model-free algorithm;

[0120] Bring the obtained network parameters into the reward function of the reward mechanism to obtain a reward value;

[0121] The Actor network and the Critic network perform cyclic calculations until the reward value of the reward function converges, the model ends the calculation, and the action parameters required for each drone are obtained.

[0122] In a possible implementation, the memory 11 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function (such as a file creation function and a data reading and writing function), etc. The data storage area may store data created during use, such as initialization data, etc.

[0123] In addition, the memory 11 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device or other volatile solid-state storage devices.

[0124] The communication interface 12 may be an interface of a communication module for connecting to other devices or systems.

[0125] Of course, it should be noted that Figure 4 the structure shown does not constitute a limitation on the UAV cluster formation device in the embodiments of the present application. In practical applications, the UAV cluster formation device may include more or fewer components than Figure 4 those shown, or combine some components.

[0126] The embodiments of the present application may also provide a computer-readable storage medium for storing program codes for executing the steps of the above-mentioned UAV cluster formation method.

[0127] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0128] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0129] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0130] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A method for forming a swarm of drones, characterized in that: include: A UAV motion model is established based on a Markov decision process; the UAV motion model includes environmental information, observation information and an action set; the environmental information includes the two-dimensional position of all UAVs, the heading of all UAVs, the speed of all UAVs and the target point position of all UAVs; the observation information includes the two-dimensional position of the current UAV; the heading of the current UAV; the speed of the current UAV; the distance between the current UAV and nearby UAVs; the action set includes left move, right move, forward move, backward move and hold; The information of the UAV motion model is brought into the Actor-Critic network model, and the network parameters of the Actor network and the Critic network are updated according to the maximum entropy model-free algorithm. The obtained network parameters are brought into the reward function of the reward mechanism to obtain a reward value; the design principle of the reward mechanism is that the closer the drone is to the target point, the greater the reward value; the closer the distance between the two drones is to the optimal distance, the greater the reward value; the reward mechanism includes a reward function for guiding the heading of the drone, a reward function for the distance between the target point and the destination, a reward function for guiding the drone to avoid obstacles, a reward function for guiding the drone to maintain a formation with its nearby drones, and a sparse reward function for the drone to arrive near the target point; The reward function for guiding the drone's heading is expressed as follows: Where: θ represents the flight angle of the UAV guided by the target point; The reward function of the distance between the target point and the destination is expressed as follows: Where: d_targe represents the Euclidean distance between the current UAV and the target point, and d_pre represents the Euclidean distance between the UAV and the target point in the previous step; The reward function for guiding the drone to avoid obstacles is expressed as follows: The reward function for guiding the UAV to maintain the formation with its nearby UAVs is expressed as follows: Where: d min Represents the safe distance between drones, d target represents the optimal distance between drones, d max Indicates the maximum distance between drones, d ij Indicates the distance between the drone and its nearby drones, d err Represents the error threshold between the two drones. This reward mechanism is to maintain the formation of the drones. The sparse reward function of the drone arriving near the target point is expressed as follows: In the formula, δd represents the set distance threshold; The overall reward function that the drone relies on is as follows: <h2 style=";text-align:left;direction:ltr">R<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> =k1r1+k2r2+k3r3+k4r4+k5r5 Where k1, k2, k3, k4, k5 are weighted coefficients, r1, r2, r3, r4, r5 are reward functions; The Actor network and the Critic network perform calculations cyclically until the reward value of the reward function converges, and the model ends the calculation to obtain the action parameters required to be executed by each drone.

2. The drone swarm formation method according to claim 1, characterized in that: The parameters of the Actor network and the Critic network are updated according to the maximum entropy model-free algorithm. The Actor network inputs the observation information of the drone and outputs the action set of the drone. The Critic network inputs the observation information and action set of the drone and outputs the evaluation Q value.

3. The UAV swarm formation method according to claim 2, characterized in that: For the maximum entropy model-free algorithm, a centralized training and distributed execution architecture is used: The Critic network inputs the observation information and action set of the drone, and the calculation formula for the output evaluation Q value is as follows: Where i represents the i-th drone, i = 1, 2, 3, ..., N, j represents the number of Critic networks, j = 1, 2, r is the reward value, γ is the discount factor, d is the round termination signal, s ′ Indicates the environmental state of the current drone at the next moment. represents the action set of other drones at the next moment, θ represents the network parameters of Actor, represents the network parameters of Critic, π θ (a|s) is the probability that the strategy takes action a in state s at the next moment, It represents the target Q value calculated by the drone when performing action a in the environment state s; According to the gradient descent method, the parameters are updated and the loss function is calculated as follows: In the formula, a i ={…,a N-1 ,…}, represents the Q value calculated when the drone takes action a in environment s; The observation information of the drone is input to the Actor network, and the action set of the drone is output. The parameters are updated according to the gradient ascent method. The loss function calculation formula is: a other Represents the action set of the drone at the current moment; Use the gradient descent method to update the entropy regularization coefficient α: In the formula, H represents the initial value of entropy; Soft update target network parameters: Where τ represents the soft update parameter.

4. A drone cluster formation device, characterized in that: Used to execute the drone swarm formation method according to any one of claims 1 to 3, the device comprises: A motion model building unit, used to build a UAV motion model based on a Markov decision process; the UAV motion model includes environmental information, observation information and an action set; The parameter input unit is used to bring the information of the UAV motion model into the Actor-Critic network model. The centralized training-distributed decision-making architecture is adopted to update the network parameters of the Actor network and the Critic network according to the maximum entropy model-free algorithm. A reward function calculation unit, used for bringing the obtained network parameters into the reward function of the reward mechanism to obtain a reward value; The action parameter acquisition unit is used for the Actor network and the Critic network to perform cyclic calculations until the reward value of the reward function converges and the model ends the calculation to obtain the actions that each drone needs to perform.

5. A drone cluster formation device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the drone cluster formation method described in any one of claims 1-3 according to the instructions in the program code.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the drone cluster formation method described in any one of claims 1-3.

Citation Information

Patent Citations

  • Unmanned aerial vehicle formation path planning method based on reinforcement learning

    CN116774731A