Unmanned aerial vehicle cluster control method, computer equipment, readable storage medium and program product

By obtaining the flight status information of the drone cluster and using the action detection model to update it in real time, the problem of low accuracy in the drone cluster roundup control is solved, and the precise roundup of the drone cluster is achieved.

CN120540328APending Publication Date: 2025-08-26NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536470.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, the control accuracy of the coordinated roundup of drone clusters is low, and the existing algorithm is not suitable for drone cluster roundup tasks, resulting in insufficient roundup accuracy.

Method used

By obtaining the flight status information of each target drone in the target drone cluster, detecting the current action using the action detection model, and updating the action detection model in real time according to the roundup reward, each target drone is controlled to perform corresponding actions until the roundup target is captured.

Benefits of technology

The control accuracy and coordinated roundup capabilities of the drone cluster have been improved to ensure that the drone cluster can accurately round up the targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540328A_ABST
    Figure CN120540328A_ABST
Patent Text Reader

Abstract

The invention relates to an unmanned aerial vehicle cluster control method, computer equipment, a readable storage medium and a program product, and is applied to the technical field of artificial intelligence, and the method comprises the steps: obtaining the flight state information of each target unmanned aerial vehicle at a current time step; a first action detection step: detecting a current action corresponding to each target unmanned aerial vehicle by an action detection model; each target unmanned aerial vehicle is controlled to execute the corresponding current action, and the flight state information of each target unmanned aerial vehicle in the next time step is obtained after the execution is completed; according to the flight state information of the next time step, respectively evaluating a surrounding award of the current action for the surrounding task, and according to each surrounding award, updating the action detection model; and respectively updating the flight state information of the next time step into the flight state information of the current time step, and returning to execute the first action detection step until the target unmanned aerial vehicle cluster captures the surrounding target. By adopting the method, the control accuracy of the unmanned aerial vehicle cluster can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a drone cluster control method, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] Collaborative drone swarm capture is an important application scenario that has attracted widespread attention. Specifically, it involves a swarm of drones performing coordinated operations and strategic maneuvers to capture a specific target. During this process, the drone swarm must select appropriate actions based on the real-time situation and swarm control strategy, coordinate to form an encirclement, and complete the capture of the target using an appropriate capture formation. To ensure the success of collaborative capture, a method for accurately controlling drone swarms is urgently needed.

[0003] Currently, there is little research on the collaborative capture of drone swarms. Usually, the artificial intelligence model in the robot control scenario is used to solve the target's position in real time, and then the corresponding tracking path is planned to approach and capture the target. However, this algorithm is not suitable for drone swarm capture tasks, resulting in low drone capture accuracy, that is, low control accuracy of drone swarms. Summary of the Invention

[0004] Based on this, it is necessary to provide a drone cluster control method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of drone capture in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for controlling a drone cluster, comprising:

[0006] Obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0007] The first action detection step: Using the action detection model, the current action of each target drone corresponding to the flight state information at the current time step is detected. The current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The action detection model is trained based on training capture targets and training drone clusters in different motion states.

[0008] Control each target UAV to execute the corresponding current action respectively, and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0009] According to the flight status information of each target drone in the next time step, the capture reward of each target drone's current action for the capture task is evaluated respectively, and the action detection model is updated according to the capture reward of each target drone;

[0010] The flight status information of each target UAV in the next time step is updated as the flight status information of each target UAV in the current time step, and the first action detection step is executed again until the target UAV cluster captures the target.

[0011] In a second aspect, the present application also provides a drone cluster control device, comprising:

[0012] The acquisition module is used to obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0013] A detection module is used for the first action detection step: using the action detection model to detect the current action of each target drone corresponding to the flight state information at the current time step, where the current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The action detection model is trained based on training capture targets and training drone clusters in different motion states;

[0014] The control module is used to control each target UAV to execute the corresponding current action and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0015] An evaluation module is used to evaluate the capture reward of each target drone's current action for the capture task based on the flight state information of each target drone in the next time step, and to update the action detection model based on the capture reward of each target drone;

[0016] The updating module is used to update the flight status information of each target UAV in the next time step to the flight status information of each target UAV in the current time step, and return to execute the first action detection step until the target UAV cluster captures the encirclement target.

[0017] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0018] Obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0019] The first action detection step: Using the action detection model, the current action of each target drone corresponding to the flight state information at the current time step is detected. The current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The action detection model is trained based on training capture targets and training drone clusters in different motion states.

[0020] Control each target UAV to execute the corresponding current action respectively, and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0021] According to the flight status information of each target drone in the next time step, the capture reward of each target drone's current action for the capture task is evaluated respectively, and the action detection model is updated according to the capture reward of each target drone;

[0022] The flight status information of each target UAV in the next time step is updated as the flight status information of each target UAV in the current time step, and the first action detection step is executed again until the target UAV cluster captures the target.

[0023] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0024] Obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0025] The first action detection step: Using the action detection model, the current action of each target drone corresponding to the flight state information at the current time step is detected. The current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The action detection model is trained based on training capture targets and training drone clusters in different motion states.

[0026] Control each target UAV to execute the corresponding current action respectively, and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0027] According to the flight status information of each target drone in the next time step, the capture reward of each target drone's current action for the capture task is evaluated respectively, and the action detection model is updated according to the capture reward of each target drone;

[0028] The flight status information of each target UAV in the next time step is updated as the flight status information of each target UAV in the current time step, and the first action detection step is executed again until the target UAV cluster captures the target.

[0029] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0030] Obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0031] The first action detection step: Using the action detection model, the current action of each target drone corresponding to the flight state information at the current time step is detected. The current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The action detection model is trained based on training capture targets and training drone clusters in different motion states.

[0032] Control each target UAV to execute the corresponding current action respectively, and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0033] According to the flight status information of each target drone in the next time step, the capture reward of each target drone's current action for the capture task is evaluated respectively, and the action detection model is updated according to the capture reward of each target drone;

[0034] The flight status information of each target UAV in the next time step is updated as the flight status information of each target UAV in the current time step, and the first action detection step is executed again until the target UAV cluster captures the target.

[0035] The above-mentioned drone cluster control method, device, computer equipment, computer-readable storage medium and computer program product obtain the flight state information of each target drone in the target drone cluster at the current time step; the first action detection step: using the action detection model, detecting the current action corresponding to each target drone under the flight state information of the current time step, wherein the current actions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target, and the action detection model is obtained by training the capture targets and training drone clusters in different motion states; controlling each target drone to perform the corresponding current action, and obtaining the flight state information of each target drone at the next time step after the execution is completed; based on the flight state information of each target drone at the next time step, evaluating the capture reward of the current action corresponding to each target drone for the capture task, and updating the action detection model based on the capture reward of each target drone; updating the flight state information of each target drone at the next time step to the flight state information of each target drone at the current time step, and returning to execute the first action detection step until the target drone cluster captures the capture target.

[0036] In this way, considering that the target may be in different motion states in the actual drone cluster capture scenario, the motion detection model obtained by training the capture targets and the training drone cluster in different motion states is used as the action decision basis for each target drone in the target drone cluster, which can improve the coordinated capture capability of the target drones. In the process of the target drone cluster capturing the capture target, the capture reward is evaluated in real time and the motion detection model is updated in real time, thereby ensuring that the controllable drone cluster can accurately capture the capture target, thereby improving the control accuracy of the drone cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without any creative work.

[0038] Figure 1 This is a diagram showing an application environment of a method for controlling a drone cluster in one embodiment;

[0039] Figure 2 1 is a flow chart of a method for controlling a drone cluster in one embodiment;

[0040] Figure 3 A schematic diagram of a drone swarm capture scenario in one embodiment;

[0041] Figure 4 1. A flow chart of detecting the current action steps corresponding to the flight state information of each target UAV at the current time step by using an action detection model in one embodiment;

[0042] Figure 5 Schematic diagram of the distribution of capture points and capture targets in one embodiment;

[0043] Figure 6 A schematic diagram of a process for evaluating a roundup reward in one embodiment;

[0044] Figure 7 1 is a flow chart of a process for training an action detection model in one embodiment;

[0045] Figure 8 This is a structural block diagram of a drone cluster control device in one embodiment;

[0046] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0048] The drone cluster control method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Each target drone 102 and terminal 104 in the target drone cluster communicates with the server 106 respectively. The data storage system stores the data that the server 106 needs to process. The data storage system can be integrated on the server 106, or it can be placed on the cloud or other network servers. The terminal 104 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 106 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0049] In an exemplary embodiment, Figure 2 As shown, a method for controlling a drone cluster is provided. Figure 1 Taking the server 106 in the example, the description is made in the form of omitting the main body, including the following steps 202 to 210.

[0050] in:

[0051] Step 202: Obtain the flight status information of each target UAV in the target UAV cluster at the current time step.

[0052] The flight state information in step 202 represents at least one of the target drone's location and any environmental obstacles in the target drone's environment. Environmental obstacles can be other drones in the target drone cluster or other obstacles such as birds, without limitation. The current time step defines a discrete time point in the real-time processing of time series data.

[0053] In step 204, the current action corresponding to each target UAV under the flight state information of the current time step is detected through the action detection model, wherein the current actions corresponding to all target UAVs constitute the capture strategy of the target UAV cluster for the capture target. The action detection model is obtained by training the capture targets and the training UAV cluster under different motion states.

[0054] Among them, the current action in step 204 is used to characterize at least one of the movement speed information and movement heading angle information of the target drone; the capture target can be a drone, or other devices with mobile functions, or other devices without mobile functions, and there is no limitation here.

[0055] Exemplarily, step 204 includes: obtaining the capture target state information of the capture target at the current time step; for each target drone, using the action detection model, mapping the capture target state information of the capture target at the current time step and the flight state information of the target drone at the current time step to the current action corresponding to the flight state information of the target drone at the current time step.

[0056] The capture target state information is used to represent the position information of the capture target.

[0057] Furthermore, through the action detection model, the capture target state information of the capture target in the current time step and the flight state information of the target UAV in the current time step are mapped to the current action corresponding to the target UAV under the flight state information of the current time step, including: the action detection model includes a strategy network, and through the strategy network in the action detection model, the capture target state information of the capture target in the current time step and the flight state information of the target UAV in the current time step are mapped to the probability of selecting each preset action under the flight state information of the target UAV at the current time step; through the strategy network in the action detection model, according to the probability of selecting each preset action under the flight state information of the target UAV at the current time step, the current action corresponding to the target UAV under the flight state information of the current time step is selected from each preset action.

[0058] As one embodiment, based on the probability of selecting each preset action corresponding to the flight status information of the target UAV at the current time step, the current action corresponding to the flight status information of the target UAV at the current time step is selected from each preset action, including: selecting the action with the highest corresponding probability from each preset action as the current action corresponding to the flight status information of the target UAV at the current time step.

[0059] As another embodiment, based on the probability of selecting each preset action under the flight status information of the target UAV at the current time step, the current action corresponding to the target UAV under the flight status information of the current time step is selected from each preset action, including: arbitrarily selecting an action whose corresponding probability belongs to the top N from each preset action as the current action corresponding to the target UAV under the flight status information of the current time step, where N is a positive integer.

[0060] In step 206 , each target UAV is controlled to execute the corresponding current action, and after the execution is completed, the flight state information of each target UAV in the next time step is obtained.

[0061] There is a fixed time step interval between the next time step in step 206 and the current time step. The time step interval can be set by the user as needed, or it can be determined by the average distance between each target drone and the capture target. Specifically, the shorter the average distance between each target drone and the capture target, the shorter the time step interval.

[0062] In this way, the capture strategy of the target drone cluster can be set in a more detailed manner to ensure the control accuracy of the drone cluster.

[0063] Exemplarily, obtaining the flight status information of each target UAV in the next time step after the execution is completed includes: obtaining the flight status information of each target UAV in the next time step after the execution is completed and a time step interval has passed.

[0064] In step 208 , based on the flight status information of each target drone in the next time step, the capture reward of the current action corresponding to each target drone for the capture mission is evaluated respectively, and the action detection model is updated according to the capture reward of each target drone.

[0065] Exemplarily, based on the flight status information of each target UAV in the next time step, the capture reward of the current action corresponding to each target UAV for the capture task is evaluated respectively, including: the action detection model includes a value network, through which, based on the flight status information of each target UAV in the next time step and the capture target state information of the capture target in the next time step, the capture reward of the current action corresponding to each target UAV for the capture task is evaluated respectively.

[0066] Exemplarily, updating the action detection model according to the capture reward of each target drone includes updating the policy network and the value network in the action detection model according to the capture reward of each target drone.

[0067] Furthermore, the policy network and the value network in the motion detection model are updated according to the capture reward of each target drone, including: updating the policy network in the motion detection model according to the capture reward of each target drone and the optimization target formula of the policy network; updating the value network in the motion detection model according to the capture reward of each target drone and the optimization target formula of the value network.

[0068] Optionally, the optimization objective formula of the policy network can be expressed as:

[0069]

[0070] Among them, L(θ) is the optimized value of the policy network, B is the batch size, is the importance sampling ratio, The probability of the policy network selecting the capture state information for the current time step, The probability of the policy network selecting the capture state information at the previous time step, is the generalized advantage estimate, ε is the clipping range hyperparameter, σ is the hyperparameter controlling the entropy coefficient, and S is the policy entropy.

[0071] Alternatively, the optimization objective formula of the value network can be expressed as:

[0072]

[0073] Among them, L(φ) is the optimized value of the value network, For discount rewards, is the value estimate of the value network for the current time step for the captured state information, The value network at the previous time step estimates the value of the captured state information.

[0074] In step 210, the flight status information of each target UAV in the next time step is updated to the flight status information of each target UAV in the current time step, and the process returns to step 204 until the target UAV cluster captures the target and the output ends.

[0075] Optionally, when the target of the capture is a drone, the above drone cluster control method can be applied to Figure 3 In the drone cluster encirclement scenario shown, the target drones (U1, U2 and U3 shown in the figure) in the target drone cluster encircle the encirclement target (U shown in the figure) to prevent the encirclement target from reaching the target mission point (T shown in the figure).

[0076] In the above-mentioned drone cluster control method, considering that in the actual drone cluster capture scenario, the capture target may be in different motion states, the motion detection model obtained by training the capture target and the training drone cluster in different motion states is used as the action decision basis for each target drone in the target drone cluster, which can improve the coordinated capture capability of the target drones. In the process of the target drone cluster capturing the capture target, the capture reward is evaluated in real time and the motion detection model is updated in real time, thereby ensuring that the controllable drone cluster can accurately capture the capture target, thereby improving the control accuracy of the drone cluster.

[0077] In an exemplary embodiment, Figure 4 As shown, a method for accurately detecting the current action of a target UAV at the current time step is provided. By using an action detection model, the current action corresponding to the flight state information of each target UAV at the current time step is detected, including steps 302 to 306. Among them:

[0078] For each target UAV, step 302 , the capture phase of the target UAV is classified according to the flight status information of the target UAV at the current time step to obtain the capture phase type of the target UAV, where the capture phase type is an approach phase type or an encirclement phase type.

[0079] Among them, the capture phase type is used to characterize the capture state of the target UAV at each time step.

[0080] Exemplarily, step 302 includes: determining the interval distance between the target UAV and the encirclement target at the current time step based on the flight status information of the target UAV at the current time step and the encirclement target status information of the encirclement target at the current time step; if the interval distance is greater than a preset distance threshold, determining the approach phase type as the encirclement phase type of the target UAV; if the interval distance is not greater than the preset distance threshold, determining the encirclement phase type as the encirclement phase type of the target UAV.

[0081] Among them, the preset distance threshold is a critical value of the interval distance set on demand to determine whether the target drone can be encircled.

[0082] In this way, the capture process of the target UAV on the capture target is divided into the approach phase and the encirclement phase, and the target UAV can be captured and controlled in a targeted manner.

[0083] Furthermore, based on the flight status information of the target UAV at the current time step and the encirclement target status information of the encirclement target at the current time step, the interval distance between the target UAV and the encirclement target at the current time step is determined, including: determining the first position information of the target UAV at the current time step based on the flight status information of the target UAV at the current time step, and determining the second position information of the encirclement target at the current time step based on the encirclement target status information of the encirclement target at the current time step; and determining the interval distance between the first position information and the second position information as the interval distance between the target UAV and the encirclement target at the current time step.

[0084] If the capture phase type of the target UAV is the approach phase type, step 304 is executed to detect the current action of the target UAV corresponding to the flight state information of the current time step based on the capture target state information of the capture target at the current time step and the flight state information of all target UAVs in the target UAV cluster at the current time step through the action detection model.

[0085] As an embodiment, through the action detection model, according to the capture target state information of the capture target at the current time step and the flight state information of all target drones in the target drone cluster at the current time step, the current action corresponding to the target drone under the flight state information of the current time step is detected, including: through the strategy network in the action detection model, the capture target state information of the capture target at the current time step and the flight state information of all target drones in the target drone cluster at the current time step are mapped to the probability of the target drone selecting each preset action under the flight state information of the current time step; through the strategy network in the action detection model, according to the probability of the target drone selecting each preset action under the flight state information of the current time step, the current action corresponding to the target drone under the flight state information of the current time step is selected from each preset action.

[0086] If the capture phase type of the target UAV is the encirclement phase type, step 306 is executed to select the target capture point corresponding to the target UAV from the multiple capture points corresponding to the capture target, and through the action detection model, based on the capture target state information of the capture target at the current time step, the target capture point and the flight state information of all target UAVs in the target UAV cluster at the current time step, the current action corresponding to the target UAV under the flight state information at the current time step is detected.

[0087] As one embodiment, a target capture point corresponding to a target drone is selected from multiple capture points corresponding to the capture target, including: generating multiple capture points whose number is consistent with the total number of drones corresponding to the target drone, wherein the multiple capture points are distributed in a peripheral area corresponding to the capture target; and selecting a target capture point corresponding to the target drone from the multiple capture points corresponding to the capture target based on the flight status information of the target drone at the current time step.

[0088] The outer area may be in a regular shape, such as a circle, a regular polygon, etc. The outer area may also be in an irregular shape, which is not limited here.

[0089] It is understandable that, in the process of encircling the target, each target drone needs to ensure that the target is encircled in all directions in the end to prevent the target from dodging and escaping.

[0090] Among them, multiple capture points can be evenly distributed in the outer area corresponding to the capture target. Specifically, when the outer area is a regular polygon, the multiple capture points are respectively distributed at the vertices of the outer area corresponding to the training capture target; when the outer area is a circle, the multiple capture points are respectively distributed on the arc corresponding to the training capture target, and the area of ​​each fan-shaped area formed by each capture point is equal. Specifically, the angle formed by two adjacent capture points and the center of the circle corresponding to the training capture target is 2π / n.

[0091] In this way, it can be ensured that each capture point is evenly distributed in the peripheral area corresponding to the capture target, so that each capture point can be subsequently assigned to each target drone, so that each target drone can approach the capture point, and the target drone cluster can accurately capture the capture target.

[0092] As an example, refer to Figure 5 When the outer area is a circle and the number of capture points (A, B and C in the figure) is 3, the arc distribution of each capture point corresponding to the capture target (U in the figure) is shown in the figure.

[0093] As one embodiment, according to the flight status information of the target UAV at the current time step, a target capture point corresponding to the target UAV is selected from multiple capture points corresponding to the capture target, including: according to the flight status information of the target UAV at the current time step, using the Hungarian algorithm, selecting the target capture point corresponding to the target UAV from multiple capture points corresponding to the capture target.

[0094] As another embodiment, according to the flight status information of the target UAV at the current time step, a target capturing point corresponding to the target UAV is selected from multiple capturing points corresponding to the capturing target, including: according to the flight status information of the target UAV at the current time step and the capturing target status information of the capturing target at the current time step, the relative orientation between the target UAV and the capturing target is positioned to obtain first relative orientation information; according to the capturing target status information of the capturing target at the current time step, the relative orientation between each capturing point and the capturing target is positioned to obtain second relative orientation information, and the capturing point whose corresponding second relative orientation information is closest to the first relative orientation information is selected from the multiple capturing points as the target capturing point corresponding to the target UAV.

[0095] As another embodiment, according to the flight status information of the target UAV at the current time step, a target capture point corresponding to the target UAV is selected from multiple capture points corresponding to the capture target, including: determining the capture point status information of each capture point at the current time step according to the capture target status information of the capture target at the current time step; generating the capture interval distance between the target UAV and each capture point according to the flight status information of the target UAV at the current time step and the capture point status information of each capture point at the current time step; and selecting the capture point with the smallest capture interval distance from each capture point as the target capture point corresponding to the target UAV.

[0096] As an embodiment, through the action detection model, according to the capture target state information of the capture target at the current time step, the target capture point and the flight state information of all target drones in the target drone cluster at the current time step, the current action corresponding to the target drone under the flight state information of the current time step is detected, including: determining the capture point state information of the target capture point at the current time step according to the capture target state information of the capture target at the current time step; through the value network in the action detection model, mapping the capture point state information of the target capture point at the current time step, the flight state information of the target drone at the current time step and the flight state information of other target drones in the target drone cluster at the current time step to the probability of the target drone selecting each preset action under the flight state information of the current time step; through the strategy network in the action detection model, according to the probability of the target drone selecting each preset action under the flight state information of the current time step, the current action corresponding to the target drone under the flight state information of the current time step is selected from each preset action.

[0097] As an embodiment, the capture point status information of the target capture point at the current time step, the flight status information of the target UAV at the current time step, and the flight status information of other target UAVs in the target UAV cluster at the current time step are mapped to the probability of the target UAV selecting each preset action under the flight status information of the current time step, including: evaluating the global value of each preset action adopted by the target UAV under the flight status information of the current time step according to the capture point status information of the target capture point at the current time step, the flight status information of the target UAV at the current time step, and the flight status information of other target UAVs in the target UAV cluster at the current time step, and obtaining a global value evaluation result; determining the probability of the target UAV selecting each preset action under the flight status information of the current time step according to the global value evaluation result, wherein the global value evaluation method can be a central value function or other methods, which are not limited here.

[0098] Furthermore, based on the capture target status information of the capture target at the current time step, the capture point status information of the target capture point at the current time step is determined, including: determining the distribution information of multiple capture points in the peripheral area corresponding to the capture target, and determining the capture point status information of the target capture point at the current time step based on the distribution information and the capture target status information of the capture target at the current time step. Specifically, according to the distribution information, the geometric relationship between each capture point and the capture target is constructed, and the capture point status information of the target capture point at the current time step is determined based on the geometric relationship between each capture point and the capture target and the capture target status information of the capture target at the current time step.

[0099] In this embodiment, the capture phase of the target drone is divided into an approach phase and an encirclement phase. In the approach phase, the target drone only needs to consider getting as close to the capture target as possible. Therefore, based on the capture target state information of the capture target at the current time step, the motion detection model detects the current motion of the target drone corresponding to the flight state information of the current time step to ensure that the current motion can only make the target drone approach the capture target as much as possible. In the encirclement phase, each target drone needs to encircle the capture target to prevent the capture target from escaping or evading. Therefore, it is necessary to assign a target encirclement point to each target drone based on the capture target. Then, based on the capture target state information of the capture target at the current time step, the target encirclement point, and the flight state information of all target drones in the target drone cluster at the current time step, the motion detection model detects the current motion of the target drone corresponding to the flight state information of the current time step to ensure that each target drone can be coordinated and controlled to get as close as possible to its corresponding target encirclement point, thereby achieving encirclement of the capture target by the target drone cluster.

[0100] In an exemplary embodiment, Figure 6 As shown, a method for accurately evaluating the roundup reward is provided. Before detecting the current action corresponding to the flight state information of each target drone at the current time step through the action detection model, the above method further includes steps 402 to 408. Among them:

[0101] For each target UAV, in step 402 , the capture distance reward of the current action corresponding to the target UAV is evaluated based on the flight state information of the target UAV in the next time step.

[0102] Exemplarily, step 402 includes: locating the interval distance between the target UAV and the capture target in the next time step based on the flight status information of the target UAV in the next time step and the capture target status information of the capture target in the next time step; and evaluating the capture distance reward of the current action corresponding to the target UAV based on the interval distance between the target UAV and the capture target in the next time step, wherein the shorter the interval distance between the target UAV and the capture target in the next time step, the higher the evaluated capture distance reward of the current action corresponding to the target UAV.

[0103] As one embodiment, the capture distance reward of the current action corresponding to the target drone is evaluated based on the interval distance between the target drone and the capture target in the next time step, including: evaluating the capture distance reward of the current action corresponding to the target drone based on the size relationship between the interval distance between the target drone and the capture target in the current time step and the next time step, wherein, the smaller the interval distance between the target drone and the capture target in the next time step is compared with the interval distance between the target drone and the capture target in the previous time step, the higher the evaluated capture distance reward.

[0104] Specifically, if the interval distance between the target drone and the capture target in the next time step is not less than the interval distance between the target drone and the capture target in the previous time step, the first constant is determined as the capture distance reward for the current action corresponding to the target drone; if the interval distance between the target drone and the capture target in the next time step is less than the interval distance between the target drone and the capture target in the previous time step, the second constant is determined as the capture distance reward for the current action corresponding to the target drone, wherein the first constant is less than the second constant, and the first constant and the second constant can both be positive numbers, or the first constant can be 0 and the second constant can be positive, or the first constant can be negative and the second constant can be positive, and there is no restriction here.

[0105] Optionally, the capture distance reward for the current action of the target drone is evaluated based on the relationship between the distance between the target drone and the capture target at the current time step and the next time step. The formula can be expressed as:

[0106]

[0107] Among them, r 1_dis is the capture distance reward of the current action corresponding to the target drone, p1 is the first constant, p2 is the second constant, d t+1 is the distance between the target drone and the captured target in the next time step, d t is the distance between the target UAV and the capture target at the current time step.

[0108] Optionally, when the capture phase type of the target UAV is the approach phase type, a step of evaluating the capture distance reward of the current action corresponding to the target UAV is performed based on the relationship between the interval distance between the target UAV and the capture target in the current time step and the next time step.

[0109] The evaluation method of setting the capture distance reward in this way takes into account that the difference between the corresponding interval distances of the target UAV in the capture stage type of the approaching stage at different time steps may vary greatly. Therefore, only the dimension of the size relationship between the interval distances between the target UAV and the capture target in the current time step and the next time step is used as the basis for setting the capture distance reward, thereby ensuring that the difference between the multiple capture distance rewards obtained by the corresponding evaluation at different time steps for the target UAV in the capture stage type of the approaching stage is small, thereby improving the rationality of the setting of the capture distance reward dimension.

[0110] As another embodiment, the capture distance reward of the current action corresponding to the target UAV is evaluated based on the interval distance between the target UAV and the capture target in the next time step, including: determining the capture distance reward of the current action corresponding to the target UAV by the ratio of a third constant and the interval distance between the target UAV and the capture target in the next time step, wherein the third constant is a positive number.

[0111] Optionally, the ratio of the third constant to the interval distance between the target UAV and the capture target at the next time step is determined as the capture distance reward for the current action corresponding to the target UAV, which can be expressed as:

[0112]

[0113] Wherein, p3 is the third constant.

[0114] Optionally, when the capture phase type of the target UAV is the encirclement phase type, a step of determining the capture distance reward for the current action corresponding to the target UAV by taking the ratio of the third constant to the interval distance between the target UAV and the capture target in the next time step can be performed.

[0115] The evaluation method of the capture distance reward is set in this way. Taking into account that the difference between the corresponding interval distances of the target UAV in the encirclement stage type at different time steps may be small, the interval distance between the target UAV and the capture target in the next time step can be used as the basis for determining the capture distance reward. This makes the interval distance between the target UAV and the capture target in the next time step have a higher proportion in the decision-making process of generating the corresponding capture distance reward under the encirclement stage type, thereby improving the rationality of the setting of the capture distance reward dimension.

[0116] Step 404 : Evaluate the capture and collision avoidance reward of the current action corresponding to the target drone based on the flight status information of all target drones in the target drone cluster at the next time step and the capture target status information of the capture target at the next time step.

[0117] Exemplarily, step 404 includes: evaluating the first collision avoidance reward between the target drone and other target drones in the target drone cluster when the target drone executes the current action according to the flight status information of all target drones in the target drone cluster in the next time step; evaluating the second collision avoidance reward between the target drone and the capture target when the target drone executes the current action according to the flight status information of the target drone in the next time step and the capture target status information of the capture target in the next time step; evaluating the capture collision avoidance reward for the current action corresponding to the target drone according to the first collision avoidance reward and the second collision avoidance reward, wherein the higher the first collision avoidance reward, the higher the evaluated capture collision avoidance reward for the current action corresponding to the target drone; the higher the second collision avoidance reward, the higher the evaluated capture collision avoidance reward for the current action corresponding to the target drone.

[0118] In this way, the first collision avoidance reward and the second collision avoidance reward are used as the basis for determining the generation of the encirclement and collision avoidance reward. Taking into account the collision risks between target drones and between target drones and encirclement targets, the rationality of the setting of the encirclement and collision avoidance reward dimension is ensured.

[0119] As one embodiment, based on the flight status information of all target drones in the target drone cluster in the next time step, the first collision avoidance reward between the target drone and other target drones in the target drone cluster when the target drone executes the current action is evaluated, including: determining the interval distance between the target drone and other drones in the target drone cluster in the next time step based on the flight status information of all target drones in the target drone cluster in the next time step; and evaluating the first collision avoidance reward between the target drone and other target drones in the target drone cluster when the target drone executes the current action based on the interval distance between the target drone and other drones in the target drone cluster in the next time step, wherein the shorter the interval distance between the target drone and other target drones in the next time step, the lower the first collision avoidance reward between the target drone and other target drones when the target drone executes the current action.

[0120] Furthermore, based on the interval distances between the target UAV and other UAVs in the target UAV cluster at the next time step, the first collision avoidance rewards between the target UAV and other target UAVs in the target UAV cluster when the target UAV performs the current action are evaluated, including: the difference between the interval distances between the target UAV and other UAVs in the target UAV cluster at the next time step and a fourth constant is determined as the first collision avoidance rewards between the target UAV and other target UAVs in the target UAV cluster when the target UAV performs the current action, wherein the fourth constant is a positive number.

[0121] As one embodiment, based on the flight status information of the target UAV in the next time step and the capture target status information of the capture target in the next time step, the second collision avoidance reward between the target UAV and the capture target is evaluated when the current action is performed, including: determining the interval distance between the target UAV and the capture target in the next time step based on the flight status information of the target UAV in the next time step and the capture target status information of the capture target in the next time step; based on the interval distance between the target UAV and the capture target in the next time step, evaluating the second collision avoidance reward between the target UAV and the capture target when the current action is performed, the shorter the interval distance between the target UAV and the capture target in the next time step, the lower the second collision avoidance reward between the target UAV and the capture target when the current action is performed.

[0122] Furthermore, based on the interval distance between the target UAV and the encirclement target at the next time step, the second collision avoidance reward between the target UAV and the encirclement target when the current action is performed is evaluated, including: the difference between the interval distance between the target UAV and the encirclement target at the next time step and a fifth constant is determined as the second collision avoidance reward between the target UAV and the encirclement target when the current action is performed, wherein the fifth constant is a positive number.

[0123] As one embodiment, based on the first collision avoidance reward and the second collision avoidance reward, the capture and collision avoidance reward of the current action corresponding to the target UAV is evaluated, including: fusing the first collision avoidance reward between the target UAV and each other target UAV with the second collision avoidance reward to obtain the capture and collision avoidance reward of the current action corresponding to the target UAV, wherein the fusion method includes but is not limited to the sum fusion method and the product fusion method.

[0124] Optionally, the first collision avoidance reward between the target UAV and each other target UAV is combined with the second collision avoidance reward to obtain the capture collision avoidance reward for the current action corresponding to the target UAV, which can be expressed as:

[0125]

[0126] Among them, r 1_ob To round up and avoid collision rewards, d ij is the distance between target UAV i and other target UAV j, p4 is the fourth constant, and p5 is the fifth constant.

[0127] Optionally, the fourth constant can be set by the user as needed, or it can be determined by the size of the drone used for training and capturing drones. Specifically, the larger the size of the drone used for training and capturing drones, the larger the fourth constant is set to. The fifth constant can be set by the user as needed, or it can be determined by the size of the drone used for training and capturing drones and the target size of the training and capturing target. Specifically, the larger the size of the drone used for training and capturing drones, the larger the fifth constant is set to. The larger the target size of the training and capturing target, the larger the fifth constant is set to.

[0128] If the capture phase type of the target UAV is the approach phase type, step 406 is executed to evaluate the step direction reward of the current action of the target UAV according to the current action of the target UAV, and the capture distance reward, capture collision avoidance reward and step direction reward are integrated to obtain the capture reward of the current action of the target UAV for the capture task.

[0129] As one embodiment, based on the current action corresponding to the target UAV, the step direction reward of the current action corresponding to the target UAV is evaluated, including: determining the step direction angle of the target UAV at the current time step based on the current action corresponding to the target UAV; evaluating the step direction reward of the current action corresponding to the target UAV based on the step direction angle of the target UAV at the current time step, wherein the step direction angle of the target UAV at the current time step represents that the closer the step direction of the target UAV is to the capture target, the higher the step direction reward obtained by evaluation.

[0130] Among them, the closer the step direction angle of the target UAV at the current time step is to 0°, the closer the step direction angle of the target UAV at the current time step is to the capture target.

[0131] Furthermore, according to the current action corresponding to the target UAV, the step direction angle of the target UAV at the current time step is determined, including: according to the current action corresponding to the target UAV, the movement direction of the target UAV at the current time step is determined, and the angle between the movement direction of the target UAV at the current time step and the target line is determined as the step direction angle of the target UAV at the current time step, wherein the target line is the line between the target UAV and the encirclement target.

[0132] As one embodiment, the stepping direction reward of the current action corresponding to the target UAV is evaluated based on the stepping direction angle of the target UAV at the current time step, including: multiplying the difference between 90° and the stepping direction angle of the target UAV at the current time step by a sixth constant to determine the stepping direction reward of the current action corresponding to the target UAV, wherein the sixth constant is a positive number.

[0133] Optionally, the product of the difference between 90° and the step direction angle of the target UAV at the current time step and the sixth constant is determined as the step direction reward of the current action corresponding to the target UAV, which can be expressed as:

[0134]

[0135] Among them, r angle is the step direction reward, α is the step direction angle of the target drone at the current time step, and p8 is the sixth constant.

[0136] In this way, considering the geometric relationship between the target UAV and the encirclement target, when the step direction angle is greater than 90°, it is considered that the step direction angle represents that the step direction of the target UAV is away from the encirclement target. At this time, the angle difference is a negative number and is multiplied by the sixth constant, thereby setting the step direction reward to a negative value, and vice versa, ensuring the rationality of the setting of the step direction reward dimension.

[0137] As one embodiment, the capture distance reward, the capture collision avoidance reward and the stepping direction reward are integrated to obtain the capture reward for the current action of the target drone for the capture task, including: the sum of the capture distance reward, the capture collision avoidance reward and the stepping direction reward is determined as the capture reward for the current action of the target drone for the capture task.

[0138] As another embodiment, the capture distance reward, the capture collision avoidance reward and the step direction reward are integrated to obtain the capture reward for the current action of the target UAV for the capture task, including: determining the energy consumption of the current action corresponding to the target UAV, and determining the difference between the sum of the capture distance reward, the capture collision avoidance reward and the step direction reward and the consumed energy as the capture reward for the current action corresponding to the target UAV for the capture task.

[0139] If the capture phase type of the target UAV is the encirclement phase type, step 408 is executed to evaluate the capture state reward of the current action of the target UAV based on the flight state information of the target UAV in the next time step, and the capture distance reward, capture collision avoidance reward and capture state reward are integrated to obtain the capture reward of the current action of the target UAV for the capture task.

[0140] As one embodiment, based on the flight status information of the target UAV in the next time step, the capture state reward of the current action corresponding to the target UAV is evaluated, including: determining the interval distance between the target UAV and the capture target in the next time step and the step direction angle of the target UAV in the next time step based on the flight status information of the target UAV in the next time step and the flight status information of the capture target in the next time step; determining the capture state of the target UAV in the next time step based on the interval distance between the target UAV and the capture target in the next time step and the step direction angle of the target UAV in the next time step; evaluating the capture state reward of the current action corresponding to the target UAV based on the capture state of the target UAV in the next time step, wherein, the more successful the capture state of the target UAV in the next time step, the higher the evaluated capture state reward of the current action corresponding to the target UAV.

[0141] As one embodiment, the capture state of the target drone in the next time step is determined based on the interval distance between the target drone and the capture target in the next time step and the step direction angle of the target drone in the next time step, including: obtaining the energy consumption of the current action corresponding to the target drone; if the energy consumption of the current action corresponding to the target drone is within a preset energy consumption range, the interval distance between the target drone and the capture target in the next time step is within the preset distance range, and the step direction angle of the target drone in the next time step is within the preset angle range, then the capture state of the target drone in the next time step is determined to be a capture success state; if the energy consumption of the current action corresponding to the target drone is not within the preset energy consumption range, the interval distance between the target drone and the capture target in the next time step is not within the preset distance range, or the step direction angle of the target drone in the next time step is not within the preset angle range, then the capture state of the target drone in the next time step is determined to be a capture failure state.

[0142] The preset energy consumption range can be the energy consumption range of the drone, or the range of the number of time steps that the drone has controlled (for example, within 500 steps, etc.). There is no restriction here. The preset angle range can be set by the user as needed or can be an empirical value. For example, The preset distance range can be set by the user as needed, or can be an empirical value, for example, (10m-30m).

[0143] As one embodiment, based on the capture state of the target UAV in the next time step, the capture state reward of the current action corresponding to the target UAV is evaluated, including: if the capture state of the target UAV in the next time step is a capture success state, then the seventh constant is determined as the capture state reward of the current action corresponding to the target UAV; if the capture state of the target UAV in the next time step is a capture failure state, then the eighth constant is determined as the capture state reward of the current action corresponding to the target UAV, wherein the seventh constant is greater than the eighth constant, and the seventh constant and the eighth constant can both be positive numbers; the seventh constant can also be a positive number and the eighth constant can be 0; the seventh constant can also be a positive number and the eighth constant can be negative.

[0144] Optionally, if the capture state of the target drone in the next time step is a capture success state, the seventh constant is determined as the capture state reward of the current action corresponding to the target drone; if the capture state of the target drone in the next time step is a capture failure state, the eighth constant is determined as the capture state reward of the current action corresponding to the target drone, which can be expressed in the formula:

[0145]

[0146] Among them, r task is the capture state reward, p7 is the seventh constant, p8 is the eighth constant, win=true means that the capture state of the target drone in the next time step is a capture success state, win=false means that the capture state of the target drone in the next time step is a capture failure state.

[0147] As one embodiment, the capture distance reward, the capture collision avoidance reward and the capture status reward are integrated to obtain the capture reward for the current action of the target drone for the capture task, including: the sum of the capture distance reward, the capture collision avoidance reward and the capture status reward is determined as the capture reward for the current action of the target drone for the capture task.

[0148] As another embodiment, the capture distance reward, the capture collision avoidance reward and the capture status reward are integrated to obtain the capture reward for the current action of the target drone for the capture task, including: the difference between the sum of the capture distance reward, the capture collision avoidance reward and the capture status reward and the energy consumption of the current action corresponding to the target drone is determined as the capture reward for the current action corresponding to the target drone for the capture task.

[0149] It is understandable that the energy consumption of the current action is set to be negatively correlated with the roundup reward, so as to ensure that the accuracy and safety of the drone roundup are improved while consuming as little energy as possible.

[0150] In this embodiment, taking into account the certain commonalities between the approach phase type and the encirclement phase type of the target UAV, specifically, both need to consider the capture distance and capture collision avoidance, therefore, the capture distance dimension and the capture collision avoidance dimension are used as common factors to participate in the decision-making basis for the capture reward evaluation. In the approach phase type, since the target UAV is usually far away from the capture target, it is necessary to restrict the stepping direction of the target UAV to ensure that the target UAV is as close to the capture target as possible. Therefore, the stepping direction dimension is used as a unique factor of the approach phase type to participate in the decision-making basis for the capture reward evaluation; in the encirclement phase type, since the target UAV is usually close to the capture target, it is possible to start considering whether the capture status of the target UAV is successful, therefore, the capture status dimension is used as a unique factor of the encirclement phase type to participate in the decision-making basis for the capture reward evaluation; in summary, the evaluation accuracy of the capture reward is improved.

[0151] In an exemplary embodiment, Figure 7 As shown, a method for accurately training an action detection model is provided. Before detecting the current action corresponding to the flight state information of each target UAV at the current time step through the action detection model, the above method further includes steps 502 to 510. Among them:

[0152] Step 502: Obtain flight status information of each training UAV in the training UAV cluster at the current time step.

[0153] Among them, the flight status information of each training drone in step 502, as well as the capture target status information of the training tail target described below, can be information generated in a simulated environment, or information collected in a real scene, or synthetic information, for example, information obtained by simulation based on the drone's physical engine, or information obtained after data enhancement of information collected in a real scene, etc., and can also be expert demonstration information, which is not limited here.

[0154] In step 504, the action detection model is used to detect the current action of each training UAV under the flight state information of the current time step according to the capture target state information of the training capture target at the current time step. The current actions corresponding to all training UAVs constitute the capture strategy of the training UAV cluster for the training capture target.

[0155] Optionally, the specific implementation of step 504 may refer to the specific implementation content of step 204 above, and will not be repeated here.

[0156] Among them, the training capture target can be in a static state, that is, the capture target state information of the training capture target is the same at any time step; the training capture target can also move according to a preset avoidance algorithm, and the preset avoidance algorithm includes but is not limited to a rule-based geometric method, a speed obstacle method, and an artificial potential field method; the training capture target can also move according to an avoidance detection model, and the avoidance detection model can be a reinforcement learning driven model, for example, DQN (Deep Q-Network) and the like.

[0157] It is understandable that since the content involved in drone swarm capture tasks is usually more complex, in order to ensure the information richness of the training basis, a combination of multiple information collected from the above-mentioned information sources is usually used as the basis for model training, thereby improving the model generalization of the trained motion detection model.

[0158] It can be understood that when the training capture target is in a stationary state, the training UAV does not actually need to encircle the training capture target when capturing the training capture target. It only needs to approach the training capture target to achieve the capture of the training capture target. Therefore, when determining the current action corresponding to the flight status information of the training UAV at the current time step according to the above step 204, the identification of the capture stage type can be omitted and the operation can be directly performed according to the approach stage type.

[0159] As an embodiment, the capture target state information of the training capture target at the current time step is obtained by an avoidance detection model. The above method also includes: obtaining the capture target state information of the training capture target at the previous time step; an avoidance detection step: through the avoidance detection model, based on the capture target state information of all training drones in the training drone cluster at the previous time step, detecting the avoidance action corresponding to the training capture target under the capture target state information at the previous time step; controlling the training capture target to perform the avoidance action, and obtaining the capture target state information of the training capture target at the current time step after the execution is completed.

[0160] Optionally, after obtaining the capture target state information of the training capture target at the current time step after execution is completed, the above method also includes: evaluating the avoidance reward of the avoidance action corresponding to the training capture target based on the capture target state information of the training capture target at the current time step, and updating the avoidance detection model based on the avoidance reward of the avoidance action corresponding to the training capture target; updating the capture target state information of the training capture target at the current time step to the capture target state information of the training capture target at the previous time step, and returning to execute the avoidance detection step until the training capture target reaches the target task point.

[0161] In this way, the real-time update of the avoidance detection model and the real-time avoidance of the surrounded target can be achieved.

[0162] As one embodiment, based on the capture target state information of the training capture target at the current time step, the avoidance reward of the avoidance action corresponding to the training capture target is evaluated, including: based on the capture target state information of the training capture target at the current time step, the reward for the avoidance action corresponding to the training capture target is evaluated to obtain an immediate avoidance reward; based on the immediate avoidance reward, the Q value of the avoidance action corresponding to the training capture target is evaluated, and the Q value is determined as the avoidance reward for the avoidance action corresponding to the training capture target.

[0163] Furthermore, according to the capture target state information of the training capture target at the current time step, the reward for the corresponding avoidance action of the training capture target is evaluated to obtain an immediate avoidance reward, including: determining the interval distance between the training capture target and each training drone at the current time step according to the capture target state information of the training capture target at the current time step and the flight state information of each training drone in the training drone cluster at the current time step; and evaluating the avoidance distance reward between the training capture target and each training drone according to the interval distance between the training capture target and each training drone at the current time step, wherein the shorter the interval distance between the training capture target and the training drone at the current time step, the better the evaluated avoidance distance reward between the training capture target and the training drone. The lower the avoidance distance reward of the drone; according to the interval distance between the training encirclement target and each training drone at the current time step, the avoidance and collision avoidance reward between the training encirclement target and each training drone is evaluated; according to the encirclement target state information of the training encirclement target at the current time step, the avoidance state of the training encirclement target under the corresponding avoidance action is determined, and according to the avoidance state of the training encirclement target at the current time step, the avoidance state reward of the training drone under the corresponding avoidance action is evaluated, among which, the more successful the avoidance state of the training encirclement target at the current time step, the higher the avoidance state reward of the training drone under the corresponding avoidance action obtained by evaluation; the avoidance distance reward, the avoidance and collision avoidance reward and the avoidance state reward are integrated to obtain the instant avoidance reward.

[0164] As an embodiment, based on the interval distances between the training encirclement target and each training drone in the current time step, the avoidance distance reward between the training encirclement target and each training drone is evaluated, including: if the interval distance between the training encirclement target and the training drone in the current time step is greater than the interval distance between the training encirclement target and the training drone in the previous time step, then the ninth constant is determined as the avoidance distance reward between the training encirclement target and the training drone; if the interval distance between the training encirclement target and the training drone in the current time step is not greater than the interval distance between the training encirclement target and the training drone in the previous time step, then the tenth constant is determined as the avoidance distance reward between the training encirclement target and the training drone, wherein the ninth constant is greater than the tenth constant, and the ninth constant and the tenth constant can both be positive numbers; the ninth constant can also be a positive number and the tenth constant can be a negative number; the ninth constant can also be a positive number and the tenth constant can be 0.

[0165] Optionally, the specific implementation method of evaluating the collision avoidance reward between the training target and each training drone based on the interval distance between the training target and each training drone in the current time step can refer to the above-mentioned specific implementation content of evaluating the second collision avoidance reward between the target drone and the capture target when the target drone performs the current action based on the flight status information of the target drone in the next time step and the capture target status information of the capture target in the next time step, which will not be repeated here.

[0166] As one embodiment, based on the capture target state information of the training capture target at the current time step, the avoidance state of the training capture target under the corresponding avoidance action is determined, including: based on the capture target state information of the training capture target at the current time step, determining the interval distance between the training capture target and the training task point at the current time step; based on the interval distance between the training capture target and the training task point at the current time step, determining the avoidance state of the training capture target under the corresponding avoidance action, wherein the shorter the interval distance between the training capture target and the training task point at the current time step, the more successful the avoidance state of the training capture target under the corresponding avoidance action.

[0167] Optionally, the specific implementation method of evaluating the avoidance state reward of the training drone corresponding to the avoidance action performed according to the avoidance state of the training capture target in the current time step can refer to the specific implementation content of evaluating the capture state reward of the target drone corresponding to the current action according to the capture state of the target drone in the next time step, which will not be repeated here.

[0168] Optionally, all constants involved in the entire text (including the first constant, ..., the tenth constant) are positive constants and can be set as needed, and are not limited here.

[0169] As one embodiment, the avoidance distance reward, the avoidance collision reward and the avoidance status reward are integrated to obtain an immediate avoidance reward, including: integrating the training capture target with the avoidance distance reward between each training drone to obtain a first integrated reward, integrating the training capture target with the avoidance collision reward between each training drone to obtain a second integrated reward, and determining the sum of the first integrated reward, the second integrated reward and the avoidance status reward as the immediate avoidance reward.

[0170] Optionally, the sum of the first fusion reward, the second fusion reward, and the avoidance state reward is determined as the immediate avoidance reward, which can be expressed by the formula:

[0171]

[0172] Among them, reward2 is the immediate avoidance reward, r 2_dis To avoid the distance reward, r 2_ob is the collision avoidance reward, and n is the total number of drones trained to capture drones.

[0173] In this way, the rewards for each training drone's avoidance distance and collision avoidance caused by the training capture target are integrated and integrated with the avoidance state reward to ensure that when controlling the training capture target, the three requirements of maintaining a certain distance from each training drone as much as possible, avoiding collisions with each training drone, and completing the training mission point as much as possible are generated. Therefore, the control accuracy of the training capture target is improved, and the action detection model under adversarial training is made more accurate.

[0174] As one embodiment, based on the immediate avoidance reward, the Q value of the training encirclement target under the corresponding avoidance action is evaluated, including: determining the future avoidance reward when the training encirclement target continues to perform the avoidance action in the future; based on the immediate avoidance reward and the future avoidance reward, determining the Q value of the training encirclement target under the corresponding avoidance action, specifically, the immediate avoidance reward and the future avoidance reward are fused to obtain a fused reward, wherein the fusion method can be addition; and determining the expected value of the fused reward as the Q value of the training encirclement target under the corresponding avoidance action.

[0175] Optionally, the expected value of the fusion reward is determined as the Q value of the training capture target when performing the avoidance action, which can be expressed by the formula:

[0176] Q(s t ,a)=E[r t,2 +γmax a′ Q(s t+1 ,a)]

[0177] Among them, Q(s t,a) is the Q value of the training target when performing the avoidance action at time t, r t,2 is the immediate avoidance reward, γmax a′ Q(s t+1 ,a) is the future avoidance reward, and γ is the discount factor.

[0178] Step 506 , control each training UAV to execute the corresponding current action, and obtain the flight status information of each training UAV in the next time step after the execution is completed.

[0179] Step 508 : Based on the flight status information of each training drone at the next time step, the capture reward of the current action corresponding to each training drone for the capture task is evaluated respectively, and the action detection model is updated according to the capture reward of each training drone.

[0180] Optionally, the specific implementation content of step 508 can refer to the specific implementation of step 208 above, which will not be repeated here.

[0181] In step 510, the flight status information of each training UAV at the next time step is updated to the flight status information of each training UAV at the current time step. The process returns to step 504 until the training UAV cluster captures the training capture target.

[0182] Optionally, the capture target status information of the training capture target at the current time step in step 504 can be updated according to different strategies, including a static strategy, a strategy under a preset avoidance algorithm, and a strategy under the control of an avoidance detection model. Every time the strategy corresponding to the capture target status information of the training capture target at the current time step in step 504 is detected to be updated, step 502 needs to be re-executed.

[0183] In this embodiment, considering that the target may be in different motion states in the actual drone cluster capture scenario, the motion detection model is updated based on the capture target state information of the training capture targets in different motion states and the training drone cluster, thereby improving the generalization of the motion detection model.

[0184] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0185] Based on the same inventive concept, the present application also provides a drone cluster control device for implementing the aforementioned drone cluster control method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more drone cluster control device embodiments provided below can be found in the above-mentioned limitations of the drone cluster control method and will not be repeated here.

[0186] In an exemplary embodiment, Figure 8 As shown, a drone cluster control device 800 is provided, including: an acquisition module 802, a detection module 804, a control module 806, an evaluation module 808 and an update module 810, wherein:

[0187] An acquisition module 802 is used to obtain the flight status information of each target UAV in the target UAV cluster at the current time step;

[0188] Detection module 804 is used for the first motion detection step: using a motion detection model to detect the current motion of each target drone corresponding to the flight state information at the current time step, where the current motions corresponding to all target drones constitute the target drone cluster's capture strategy for the capture target. The motion detection model is trained based on training capture targets and training drone clusters in different motion states;

[0189] The control module 806 is used to control each target UAV to execute the corresponding current action and obtain the flight status information of each target UAV in the next time step after the execution is completed;

[0190] An evaluation module 808 is configured to evaluate the capture reward of the current action corresponding to each target drone for the capture task based on the flight state information of each target drone at the next time step, and update the action detection model based on the capture reward of each target drone;

[0191] The updating module 810 is used to update the flight status information of each target drone in the next time step to the flight status information of each target drone in the current time step, and return to execute the first action detection step until the target drone cluster captures the encirclement target.

[0192] In one embodiment, the detection module 804 is further configured to classify the capture phase of each target drone according to the flight state information of the target drone at the current time step, and obtain the capture phase type of the target drone, wherein the capture phase type is an approach phase type or an encirclement phase type; if the capture phase type of the target drone is an approach phase type, then the motion detection model is used to detect the current motion of the target drone under the flight state information of the current time step according to the capture target state information of the capture target at the current time step and the flight state information of all target drones in the target drone cluster at the current time step; if the capture phase type of the target drone is an encirclement phase type, then the target capture point corresponding to the target drone is selected from the multiple capture points corresponding to the capture target, and the motion detection model is used to detect the current motion of the target drone under the flight state information of the current time step according to the capture target state information of the capture target at the current time step, the target capture point, and the flight state information of all target drones in the target drone cluster at the current time step.

[0193] In one embodiment, the detection module 804 is further configured to generate a plurality of capture points, the number of which is consistent with the total number of drones corresponding to the target drone, wherein the plurality of capture points are distributed in a peripheral area corresponding to the capture target; based on the flight status information of the target drone at the current time step and the capture target status information of the capture target at the current time step, the relative position between the target drone and the capture target is located to obtain relative position information; based on the relative position information, the target capture point corresponding to the target drone is selected from the plurality of capture points.

[0194] In one embodiment, the detection module 804 is further used to determine the interval distance between the target UAV and the encirclement target at the current time step based on the flight status information of the target UAV at the current time step and the encirclement target status information of the encirclement target at the current time step; if the interval distance is greater than a preset distance threshold, the approach phase type is determined as the encirclement phase type of the target UAV; if the interval distance is not greater than the preset distance threshold, the encirclement phase type is determined as the encirclement phase type of the target UAV.

[0195] In one embodiment, the evaluation module 808 is further configured to evaluate, for each target drone, the capture distance reward of the current action corresponding to the target drone based on the flight state information of the target drone in the next time step; evaluate the capture collision avoidance reward of the current action corresponding to the target drone based on the flight state information of all target drones in the target drone cluster in the next time step and the capture target state information of the capture target in the next time step; if the capture phase type of the target drone is the approach phase type, then evaluate the step direction reward of the current action corresponding to the target drone based on the current action corresponding to the target drone, and fuse the capture distance reward, capture collision avoidance reward and step direction reward to obtain the capture reward of the current action corresponding to the target drone for the capture task; if the capture phase type of the target drone is the encirclement phase type, then evaluate the capture state reward of the current action corresponding to the target drone based on the flight state information of the target drone in the next time step, and fuse the capture distance reward, capture collision avoidance reward and capture state reward to obtain the capture reward of the current action corresponding to the target drone for the capture task.

[0196] In one embodiment, before detecting the current action corresponding to each target drone under the flight state information of the current time step through the action detection model, the acquisition module 802 is also used to obtain the flight state information of each training drone in the training drone cluster at the current time step; the detection module 804 is also used for the second action detection step: through the action detection model, according to the capture target state information of the training capture target at the current time step, the current action corresponding to each training drone under the flight state information of the current time step is detected, wherein the current actions corresponding to all training drones constitute the capture strategy of the training drone cluster for the training capture target; the control module 806 is also used to respectively Control each training drone to perform the corresponding current action, and obtain the flight status information of each training drone in the next time step after the execution is completed; the evaluation module 808 is also used to evaluate the capture reward of the current action corresponding to each training drone for the capture task based on the flight status information of each training drone in the next time step, and update the action detection model according to the capture reward of each training drone; the update module 810 is also used to update the flight status information of each training drone in the next time step to the flight status information of each training drone in the current time step, and return to execute the second action detection step until the training drone cluster captures the training capture target.

[0197] In one embodiment, the capture target state information of the training target at the current time step is obtained by the avoidance detection model, and the acquisition module 802 is also used to obtain the capture target state information of the training target at the previous time step; the detection module 804 is also used to detect the corresponding avoidance action of the training target under the capture target state information of the previous time step through the avoidance detection model based on the capture target state information of all training drones in the training drone cluster at the previous time step; the control module 806 is also used to control the training target to perform the avoidance action, and obtain the capture target state information of the training target at the current time step after the execution is completed.

[0198] Each module in the aforementioned drone swarm control device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device's memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0199] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 9 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method for controlling a drone cluster is implemented. The display unit of the computer device is used to form a visually visible image, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0200] Those skilled in the art will understand that Figure 9The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0201] In one exemplary embodiment, a computer device is provided, comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-described method embodiments are implemented. In one embodiment, a computer-readable storage medium is provided, storing the computer program, and when the processor executes the computer program, the steps of the above-described method embodiments are implemented. In one embodiment, a computer program product is provided, comprising the computer program, and when the processor executes the computer program, the steps of the above-described method embodiments are implemented.

[0202] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.

[0203] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0204] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for controlling a drone cluster, characterized in that: The method comprises: Obtain the flight status information of each target UAV in the target UAV cluster at the current time step; A first motion detection step: detecting the current motion of each target drone under the flight state information of the current time step using a motion detection model, wherein the current motions corresponding to all target drones constitute the capture strategy of the target drone cluster for the capture target. The motion detection model is trained based on training capture targets and training drone clusters under different motion states; Controlling each of the target UAVs to execute a corresponding current action, and obtaining flight status information of each of the target UAVs at the next time step after the execution is completed; According to the flight state information of each target drone at the next time step, respectively evaluate the capture reward of the current action corresponding to each target drone for the capture task, and update the action detection model according to the capture reward of each target drone; The flight status information of each target drone in the next time step is updated as the flight status information of each target drone in the current time step, and the first action detection step is executed again until the target drone cluster captures the encirclement target.

2. The method according to claim 1, characterized in that The method of detecting the current action of each target UAV corresponding to the flight state information of the current time step by using the action detection model includes: For each target UAV, classify the capture phase of the target UAV according to the flight state information of the target UAV at the current time step to obtain the capture phase type of the target UAV, wherein the capture phase type is an approach phase type or an encirclement phase type; If the capture phase type of the target UAV is the approach phase type, then using an action detection model, based on the capture target state information of the capture target at the current time step and the flight state information of all target UAVs in the target UAV cluster at the current time step, detect the current action of the target UAV corresponding to the flight state information at the current time step; If the capture phase type of the target UAV is the encirclement phase type, the target capture point corresponding to the target UAV is selected from the multiple capture points corresponding to the capture target, and through the action detection model, according to the capture target state information of the capture target at the current time step, the target capture point and the flight state information of all target UAVs in the target UAV cluster at the current time step, the current action corresponding to the target UAV under the flight state information at the current time step is detected.

3. The method according to claim 2, characterized in that The step of selecting a target capture point corresponding to the target drone from a plurality of capture points corresponding to the capture targets includes: Generating a plurality of capture points whose number is the same as the total number of drones corresponding to the target drones, wherein the plurality of capture points are distributed in a peripheral area corresponding to the capture target; Positioning the relative position between the target UAV and the encircled target according to the flight state information of the target UAV at the current time step and the encircled target state information of the encircled target at the current time step to obtain relative position information; According to the relative position information, a target capture point corresponding to the target UAV is selected from the multiple capture points.

4. The method according to claim 2, characterized in that The step of classifying the capture phase of the target UAV according to the flight state information of the target UAV at the current time step to obtain the capture phase type of the target UAV includes: Determine the interval distance between the target UAV and the encircled target at the current time step according to the flight state information of the target UAV at the current time step and the encircled target state information of the encircled target at the current time step; If the separation distance is greater than a preset distance threshold, the approach phase type is determined as the capture phase type of the target UAV; If the interval distance is not greater than the preset distance threshold, the encirclement phase type is determined as the capture phase type of the target UAV.

5. The method according to claim 2, characterized in that The step of evaluating the capture reward for the capture mission for the current action corresponding to each target drone based on the flight state information of each target drone at the next time step to obtain the capture reward includes: For each target drone, evaluate the capture distance reward for the current action corresponding to the target drone based on the flight state information of the target drone at the next time step; Evaluate the capture and collision avoidance reward of the current action corresponding to the target drone based on the flight state information of all target drones in the target drone cluster at the next time step and the capture target state information of the capture target at the next time step; If the capture phase type of the target drone is the approach phase type, then based on the current action of the target drone, the step direction reward of the current action of the target drone is evaluated, and the capture distance reward, the capture collision avoidance reward and the step direction reward are integrated to obtain the capture reward of the current action of the target drone for the capture task; If the capture phase type of the target UAV is the encirclement phase type, the capture state reward of the current action corresponding to the target UAV is evaluated according to the flight state information of the target UAV in the next time step, and the capture distance reward, the capture collision avoidance reward and the capture state reward are integrated to obtain the capture reward of the current action corresponding to the target UAV for the capture task.

6. The method according to any one of claims 1 to 5, characterized in that Before detecting the current action corresponding to the flight state information of each target UAV at the current time step by the action detection model, the method further includes: Obtain the flight status information of each training UAV in the training UAV cluster at the current time step; A second action detection step: using an action detection model to detect the current action of each training UAV under the flight state information of the training target at the current time step according to the capture target state information of the training target at the current time step, wherein the current actions corresponding to all the training UAVs constitute the capture strategy of the training UAV cluster for the training target; Controlling each of the training drones to execute a corresponding current action, and obtaining flight status information of each of the training drones at the next time step after the execution is completed; According to the flight state information of each training drone at the next time step, respectively evaluate the capture reward of the current action corresponding to each training drone for the capture task, and update the action detection model according to the capture reward of each training drone; The flight status information of each training drone in the next time step is updated to the flight status information of each training drone in the current time step, and the second action detection step is executed again until the training drone cluster captures the training encirclement target.

7. The method according to claim 6, characterized in that The state information of the training capture target at the current time step is obtained by the avoidance detection model, and the method further includes: Obtain the capture target state information of the training capture target in the previous time step; By means of the avoidance detection model, according to the capture target state information of all the training drones in the training drone cluster at the previous time step, the avoidance action corresponding to the training capture target under the capture target state information at the previous time step is detected; The training capture target is controlled to execute the avoidance action, and after the execution is completed, the capture target state information of the training capture target at the current time step is obtained.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.