Control methods, devices, equipment, media, and software products for unmanned aerial vehicle (UAV) swarms
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-14
AI Technical Summary
[0007]本申请提供的技术方案至少带来以下有益效果:由于低层控制器可以提取全局场景特征信息,即可以提取当前场景中的所有信息,从而高层切换器可以根据全局场景特征信息,动态地为每架无人机分配最合适的子任务,使得无人机群可以避免任务重叠并进行补给规划,从而将有限的无人机资源精准地配置到最需要的任务上,实现了对多种类型的子任务的协同权衡,如此提升了整体资源配置效率。
Smart Images

Figure CN122569433A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automatic control technology, and in particular to a control method, device, equipment, medium and program product for an unmanned aerial vehicle (UAV) swarm. Background Technology
[0002] In various emergency rescue scenarios, rapidly acquiring disaster situation information and ensuring communication connectivity are crucial for improving rescue efficiency. Drones, with their mobility and flexibility and lack of ground-based limitations, have become important tools for performing disaster coordination and communication relay tasks.
[0003] Currently, drone swarms mainly employ a distributed collaborative approach to perform the aforementioned tasks. This means that drones make autonomous decisions based on local perception information, such as the detection range of their own sensors and the communication status with neighboring drones, in order to reduce dependence on a central command node and adapt to the complex communication conditions in mountainous and forested terrain.
[0004] However, because drones only have local environmental conditions and limited collaborative information, it is difficult to uniformly optimize perception coverage and communication assurance at the global level, resulting in low overall resource allocation efficiency. Summary of the Invention
[0005] This application provides a method, apparatus, equipment, medium, and program product for controlling unmanned aerial vehicle (UAV) swarms, which can improve the overall resource allocation efficiency.
[0006] In a first aspect, embodiments of this application provide a method for controlling a drone swarm. The method includes: acquiring scene information and state information of the drone swarm's mission scenario, wherein the drone swarm includes at least one drone; the scene information includes scene map information and personnel location information; and the state information includes drone location information and drone battery level information; inputting the scene information and state information into a drone control model, which includes a low-level task controller and a high-level task switcher; extracting global feature information from the scene information and state information through the low-level task controller, the global feature information including scene feature information and drone feature information; assigning at least one sub-task to each drone in the drone swarm based on the global feature information through the high-level task switcher; and generating and outputting control strategy information for the drone swarm for each of the at least one sub-task through the low-level action controller.
[0007] The technical solution provided in this application brings at least the following beneficial effects: Since the low-level controller can extract global scene feature information, that is, it can extract all the information in the current scene, the high-level switcher can dynamically allocate the most suitable sub-task to each UAV according to the global scene feature information, so that the UAV swarm can avoid task overlap and carry out replenishment planning, thereby accurately allocating limited UAV resources to the most needed tasks, realizing the collaborative trade-off of multiple types of sub-tasks, thus improving the overall resource allocation efficiency.
[0008] One possible implementation is that the above-mentioned acquisition of the task scenario information and the state information of the drone swarm includes: acquiring the task scenario information and the state information of the drone swarm at each time step.
[0009] Another possible implementation is that the above-mentioned at least one sub-task includes at least one type of perception, communication coverage, and resupply; the above-mentioned allocation of at least one sub-task to each UAV in the UAV swarm based on global feature information by the high-level task switcher includes: inputting global feature information into the high-level policy network in the high-level task switcher, outputting the task score of each type of sub-task corresponding to each UAV in the UAV swarm; and assigning the sub-task of the type with the highest task score to the corresponding UAV.
[0010] Another possible implementation is that, after generating and outputting control strategy information for each of the at least one subtasks of the UAV swarm through a low-level motion controller, the method further includes: when at least one subtask includes perception, generating a first control command based on the control strategy information, which is used to control the UAV to collect scene information and circumnavigate along the scene.
[0011] Another possible implementation is that, after generating and outputting control strategy information for each of the at least one subtasks of the UAV swarm through a low-level motion controller, the method further includes: generating a second control command based on the control strategy information when at least one subtask includes communication coverage, the second control command being used to control the UAV to move to the target location.
[0012] Another possible implementation involves generating and outputting control strategy information for the UAV swarm for each subtask in at least one subtask via a low-level motion controller. The method further includes generating a third control command based on the control strategy information when at least one subtask includes resupply. This third control command is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
[0013] Secondly, embodiments of this application provide a control device for a drone swarm, comprising: an acquisition module, an input module, an extraction module, an allocation module, and a generation module. The acquisition module is used to acquire scene information of the task scenario of the drone swarm and state information of the drone swarm, wherein the drone swarm includes at least one drone, the scene information includes scene map information and personnel location information, and the state information includes drone location information and drone battery information; the input module is used to input the scene information and state information into a drone control model, the drone control model including a low-level task controller and a high-level task switcher; the extraction module is used to extract global feature information from the scene information and state information through the low-level task controller, the global feature information including scene feature information and drone feature information; the allocation module is used to allocate at least one sub-task to each drone in the drone swarm based on the global feature information through the high-level task switcher; the generation module is used to generate and output control strategy information for the drone swarm for each of the at least one sub-task through the low-level action controller.
[0014] One possible implementation is that the aforementioned acquisition module is specifically used to acquire scene information of the drone swarm's mission scenario and state information of the drone swarm at each time step.
[0015] Another possible implementation is that the above-mentioned at least one sub-task includes at least one type of perception, communication coverage, and resupply; the above-mentioned allocation module is specifically used to: input global feature information into the high-level policy network in the high-level task switcher, output the task score of each type of sub-task corresponding to each UAV in the UAV swarm; and allocate the sub-task of the type with the highest task score corresponding to each UAV to the corresponding UAV.
[0016] In another possible implementation, the aforementioned generation module is further configured to, after generating and outputting control strategy information for each of the at least one subtasks of the UAV swarm through the low-level motion controller, generate a first control command based on the control strategy information, provided that at least one subtask includes perception. This first control command is used to control the UAV to collect scene information and circumnavigate along the scene.
[0017] In another possible implementation, the aforementioned generation module is further configured to, after generating and outputting control strategy information for each of the at least one subtasks through the low-level motion controller, generate a second control command based on the control strategy information, provided that at least one subtask includes communication coverage. The second control command is used to control the UAVs to move to the target location.
[0018] In another possible implementation, the above-mentioned generation module is further configured to generate a third control command based on the control strategy information after generating and outputting control strategy information for each subtask in at least one subtask through the low-level motion controller. If at least one subtask includes resupply, the third control command is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
[0019] Thirdly, this application provides an electronic device comprising: a processor and a memory; the memory stores a program or instructions executable on the processor, wherein the program or instructions, when executed by the processor, implement the method of the first aspect described above.
[0020] Fourthly, this application provides a readable storage medium on which a program or instructions are stored, which, when executed by a computer, implement the method of the first aspect described above.
[0021] Fifthly, this application provides a computer program product stored in a storage medium, which, when executed by a computer, implements the method described in the first aspect.
[0022] In a sixth aspect, embodiments of this application provide a chip including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0023] The beneficial effects of the second to sixth aspects mentioned above are described in the corresponding description of the first aspect and will not be repeated here. Attached Figure Description
[0024] Figure 1 A schematic diagram of the network architecture for an application of a control method for an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0025] Figure 2 A flowchart illustrating a method for controlling an unmanned aerial vehicle (UAV) swarm, as provided in an embodiment of this application;
[0026] Figure 3 This is a schematic diagram illustrating a scenario of a swarm of drones collaboratively performing a synesthetic task during a forest fire, as provided in an embodiment of this application.
[0027] Figure 4 A schematic diagram of a Moore-type neighborhood provided in an embodiment of this application;
[0028] Figure 5 A schematic diagram of a fire scene structure division model provided in an embodiment of this application;
[0029] Figure 6 A schematic diagram illustrating the movement direction of the unmanned aerial vehicle (UAV) provided in an embodiment of this application;
[0030] Figure 7 A flowchart illustrating another method for controlling an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0031] Figure 8 A flowchart illustrating another method for controlling an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0032] Figure 9 The architecture diagram of the framework for solving the problem of Hierarchical Markov Decision Process (H-MDP) collaborative communication in a forest fire scenario provided in this application embodiment;
[0033] Figure 10 A flowchart illustrating another method for controlling an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0034] Figure 11 A flowchart illustrating another method for controlling an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0035] Figure 12 A flowchart illustrating another method for controlling an unmanned aerial vehicle (UAV) swarm provided in this application embodiment;
[0036] Figure 13 Fire exploration rate - personnel coverage over time for the hierarchical reinforcement learning method (CS-HRL, Communication–SensingCooperative Hierarchical Reinforcement Learning) and the fixed task assignment method provided in the embodiments of this application;
[0037] Figure 14 A schematic diagram comparing the CS-HRL method and the fixed strategy provided in the embodiments of this application;
[0038] Figure 15 This is a schematic diagram of the structure of a control device for an unmanned aerial vehicle (UAV) swarm provided in an embodiment of this application;
[0039] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0040] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0041] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0042] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0043] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0044] The present application provides a method, apparatus, equipment, medium, and program product for controlling a drone swarm, which can be applied in emergency scenarios where drone swarms collaboratively perform disaster awareness and communication support tasks.
[0045] In existing technologies, rapidly acquiring disaster situation information and ensuring communication connectivity between rescue personnel and the command system are key issues in various emergency rescue scenarios. Due to their advantages such as high mobility, flexible deployment, and lack of ground condition limitations, drones have been widely used to simultaneously undertake tasks such as disaster awareness and communication relay. This application model has strong universality in various emergency scenarios such as forest fires, earthquakes, and landslides.
[0046] In existing emergency applications, traditional methods for using drones to perform disaster perception and communication support tasks mainly fall into two categories. One method separates or sequentially executes these tasks; the drone first acquires environmental information and then plans communication relay locations or coverage areas based on the perception results. The other method relies primarily on manual planning to organize the drone to simultaneously perform perception and communication tasks. Specifically, operators pre-plan the perception flight path, perception area, and communication relay point locations for the drone based on terrain conditions, disaster experience, or mission objectives, and make adjustments via manual commands during mission execution.
[0047] In addition to traditional methods, recent research has begun to explore algorithm-driven approaches for the joint scheduling of disaster perception and communication support tasks by UAV swarms. Existing implementations employ centralized decision-making structures, where ground stations or central nodes collect UAV status, perception results, and environmental information, uniformly calculate decision results such as perception coverage, communication relay locations, or flight paths, and then distribute these decisions to individual UAVs for execution. Other methods utilize distributed decision-making, allowing UAVs to make autonomous decisions and coordinate based on their acquired local perception information and limited communication conditions, thus reducing dependence on central nodes. In practical implementation, related research typically guides UAV swarms to coordinate deployments in terms of perception coverage and communication connectivity by predefined task assignments, coordination rules, or optimization objectives, thereby supporting the simultaneous execution of disaster perception and communication support tasks in complex emergency scenarios.
[0048] Currently, there are significant shortcomings in deploying drone swarms for perception and communication support in mountainous emergency scenarios. Existing collaborative methods for using drone swarms for mountainous emergency perception and communication support often model and plan disaster perception tasks and communication support tasks separately or sequentially, lacking a mechanism for coordinating and balancing the two types of tasks. Under this separate implementation approach, the drone's perception path planning, communication coverage deployment, and task allocation relationships are often determined based on static rules or preset parameters, making it difficult to dynamically update as the disaster evolves, the location and number of rescue personnel change, and the drone scale adjusts. Furthermore, this type of method has limited adaptability to environmental and task dynamics, making it difficult to achieve continuous and effective task adjustments in complex emergency scenarios.
[0049] Meanwhile, in complex terrain environments such as mountains and forests, the terrain makes it difficult for drone swarms to maintain stable and reliable global communication conditions over long periods, limiting the applicability of centralized scheduling methods. Therefore, existing technologies are increasingly adopting distributed collaboration methods, where drones make autonomous decisions based on local perception information. However, in distributed decision-making models, drones typically only possess local environmental conditions and limited collaborative information, making it difficult to uniformly optimize perception coverage and communication assurance at the global level, easily leading to a decline in overall resource allocation efficiency. Without an effective collaboration mechanism, locally optimal decisions are difficult to translate into globally optimal results.
[0050] Finally, existing research does not adequately consider issues such as power consumption, return-to-base resupply, and redeployment of drones during long-duration emergency missions. It lacks effective methods to incorporate energy constraints into a unified decision-making process, making it difficult to maintain stable perception and communication capabilities during multiple rounds of long-duration rescue operations.
[0051] To address the aforementioned issues, the purpose of this application is to propose a collaborative decision-making method for UAV swarms in mountainous emergency scenarios. This method enables UAV swarms to coordinate disaster awareness and communication support tasks under unified decision-making guidance, adapting to the evolution of the disaster and changes in rescue needs. Under conditions of limited communication, a reasonable decision evaluation and guidance mechanism supports distributed collaboration of UAVs based on local perception, using local information to optimize global objectives. Simultaneously, the decision-making process comprehensively considers the energy status of UAVs and the requirements for mission continuity to meet the requirements for stable collaboration in multi-round, long-term emergency rescue operations.
[0052] To address the aforementioned technical issues, this application provides a control method, apparatus, device, medium, and program product for a drone swarm. Since the lower-level controller can extract global scene feature information, that is, it can extract all information in the current scene, the higher-level switcher can dynamically allocate the most suitable sub-task to each drone based on the global scene feature information. This allows the drone swarm to avoid task overlap and perform replenishment planning, thereby accurately allocating limited drone resources to the most needed tasks and achieving collaborative trade-offs for multiple types of sub-tasks, thus improving the overall resource allocation efficiency.
[0053] The following description, in conjunction with the accompanying drawings, details the control method, apparatus, equipment, medium, and program products for unmanned aerial vehicle (UAV) swarms provided in the embodiments of this application.
[0054] Figure 1 The network architecture of a drone swarm control method provided in an embodiment of this application is illustrated. For example... Figure 1 As shown, the network architecture includes a control device 101 for the drone swarm and a terminal device 102. The control device 101 and the terminal device 102 are interconnected.
[0055] In some embodiments, the control device 101 for the drone swarm can be a server, a computer, or a processor or processing unit within a server or computer. The server can be a single server or a server cluster consisting of multiple servers. It should be noted that this application does not limit the specific device form of the control device 101 for the drone swarm. Figure 1 The control device 101 for the Chinese and Israeli drone swarm is shown as a single server as an example.
[0056] In some embodiments, the terminal device may be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and the embodiments of this application do not specifically limit it. Figure 1 The example shown is a mobile phone, with terminal device 102 as an example.
[0057] In some embodiments, the control device 101 of the drone swarm receives high-level mission instructions and initial personnel location information from the terminal device 102, and then autonomously executes the perception, decision-making and control process, and reports the generated task allocation results, drone real-time status and key perception data to the terminal device 102 for visualization display in real time; at the same time, the terminal device 102 serves as a monitoring and intervention interface, allowing operators to view the overall situation and send manual intervention instructions to the control device 101.
[0058] It should be noted that the network architecture described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As network architectures evolve, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0059] See Figure 2 This is a flowchart illustrating a method for controlling an unmanned aerial vehicle (UAV) swarm, as provided in an embodiment of this application. Figure 2 As shown, the control method for a drone swarm provided in this application embodiment can be implemented by the aforementioned drone swarm control device, specifically including the following steps 201 to 205.
[0060] Step 201: The control device of the drone swarm acquires the scene information of the mission scenario and the status information of the drone swarm.
[0061] In some embodiments, the drone swarm includes at least one drone, the scene information includes scene map information and personnel location information, and the status information includes drone location information and drone battery level information.
[0062] In some embodiments, the control device of the aforementioned drone swarm refers to a physical or logical entity that integrates computing, communication and decision-making functions. It may be a ground control station or a leading drone with strong computing power in the swarm.
[0063] In some embodiments, the scenario information of the above-mentioned task scenario refers to the objective environment in which the drone needs to operate.
[0064] In some embodiments, the above-mentioned scene map information refers to the geospatial information of the task area.
[0065] In some embodiments, the aforementioned personnel location information refers to the real-time location of personnel involved in a task who need to be protected or assisted.
[0066] In some embodiments, the status information of the aforementioned drone swarm refers to the current status of the drones performing the mission, which is the basis for mission allocation and energy management.
[0067] In some embodiments, the above-mentioned drone location information is the real-time coordinates of each drone in three-dimensional space.
[0068] In some embodiments, the aforementioned drone battery information is the current remaining battery power or estimated flight time of each drone, used to determine whether the drone can continue to perform its mission or whether it needs to return to base immediately for resupply.
[0069] In some embodiments, such as Figure 3 The image shows a scenario where a swarm of drones collaboratively performs a sensing task during a forest fire. This embodiment of the application defines the task area... Discretization into cells Assuming a cell The state at time t is represented as -1 represents unexplored, 0 represents healthy, 1 represents burning, 2 represents complete combustion, and 3 represents extinguished. The set of burning units at time t is defined as follows: The set of edge fire points at time t is defined as follows: That is, in the four neighboring regions There is at least one healthy (unburned) burning cell. At the start of the mission, M firefighters depart from the command center, carrying portable fire extinguishing equipment, and advance along the fire line to conduct ground firefighting operations at the edge of the fire. N drones take off from the charging station, equipped with communication relay equipment and multimodal sensing payloads, to perform aerial support tasks such as fire information perception and communication coverage for ground rescue personnel. During the mission, the fire continues to evolve according to a pre-set forest fire spread model, and the burning status of each cell in the fire is dynamically updated over time. At any time step, when rescue personnel move to a certain edge set... When a fire occurs in a cell adjacent to a burning cell, it is considered that the rescuer has carried out a fire extinguishing operation on the edge of the fire. The state of the corresponding cell changes from burning to extinguished to reflect that the fire has been effectively controlled.
[0070] In some embodiments, the objective of this application is to extinguish all fire spots at the edge of the fire scene, i.e. At this point, the weighted sum of the average fire exploration rate and the average communication coverage rate of rescue personnel throughout the entire mission cycle reaches its maximum value, thus achieving effective communication support for ground firefighting operations. To objectively reflect the fire development pattern and quantify the operational constraints of the drone swarm, the following sections will introduce the forest fire spread model, the fire field structure division model, the drone model, the rescue personnel model, and the communication model in sequence.
[0071] In some embodiments, the forest fire spread model can be established using a cellular automaton model, a fire spread model, and a dynamic evolution mechanism for forest fire spread. The cellular automaton discretizes the fire field using a regular grid, and synchronously updates the combustion state at each time step based on the neighborhood state and preset transition rules. Its local-parallel evolution mechanism is consistent with the spatiotemporal coupling characteristics of forest fire spread and can embed wind speed and slope correction coefficients. Accordingly, this embodiment uses a cellular automaton as a framework, loading the velocity calculation results of the fire spread model into the state transition rules to achieve dynamic simulation of forest fire spread. Specifically:
[0072] (1) Cellular automata model
[0073] Cellular automata are discrete-time and discrete-space grid dynamic models that can evolve automatically under given conditions. Their characteristic is that the system is divided into a regular cellular network, where each grid cell exists in a finite number of states and is influenced by the states of its neighboring cells. The model mainly consists of four parts: cellular space, cellular neighborhood, cellular states, and transition rules. Specifically, this paper discretizes the fire area into a two-dimensional regular grid, with each grid cell denoted as . The neighborhood of a cell adopts the mole neighborhood. That is, its eight neighboring cells, such as Figure 4 As shown, a schematic diagram of a Moore-type neighborhood is presented.
[0074] In some embodiments, the following cell states and transition rules can be defined.
[0075] A cell state of -1 indicates that the cell has not yet been explored and its state is unknown.
[0076] Cell state = 0 indicates that the cell is healthy and has not yet burned.
[0077] Cell state = 1 indicates that the cell is burning, but its heat radiation is absorbed by the unburned material inside the cell, so it does not have the ability to radiate heat to the surroundings.
[0078] Cell state = 2 indicates that the cell is fully burned and radiates heat to the surrounding molar neighborhood.
[0079] Cell state = 3 indicates that the cell has either burned completely or been extinguished.
[0080] (2) Fire spread model
[0081] The fire spread model is a semi-empirical forest fire spread rate model, widely used for forest fire prediction in the primary coniferous forests and mixed coniferous and broad-leaved forests of Northeast China. The fire spread model has four parameters: initial spread rate, combustible material correction coefficient, wind speed correction coefficient, and topographic correction coefficient. The calculation method is shown in the following formula (1):
[0082] Formula (1)
[0083] Where R is the forest fire spread rate (m / min). It is the initial spread rate, calculated as shown in the following formula (2):
[0084] Formula (2)
[0085] Where T is the daily maximum temperature (°C), W is the average wind speed at noon, and h is the daily minimum humidity (%).
[0086] It is a combustible material correction factor, and the value varies for different fuel types (e.g., pine needles 0.8, dead branches and fallen leaves 1.2, grassland 2.0, etc.). It is the wind speed correction factor, defined as , where v is the average wind speed (m / s). Terrain correction factor, defined as ,in The slope angle is positive for uphill and negative for downhill.
[0087] (3) Dynamic evolution mechanism of forest fire spread
[0088] At the start of each time step, the velocity contribution of each fully burned cell to its surrounding moiré-shaped neighborhood is calculated based on the fire spread model. This contribution is multiplied by the time step length to calculate the burning distance, and then divided by the edge length to obtain the contribution rate. The state of a cell at time t+1 is determined by the states of its neighboring cells and its own state, expressed as:
[0089] Where (i, j) is the cell row and column number, t is the time step, A is the cell state, and R is the rate at which the cell contributes to the central cell. is the time interval between every two time steps; 'a' is the cell side length.
[0090] In some embodiments, the fire scene structure partitioning model is as follows: Figure 5 As shown. In order to better reflect the dynamic evolution characteristics of the fire scene, this invention introduces the definition of fire scene structure based on various professional fire extinguishing strategies, and divides the combustion area into three types of areas according to its relative position with respect to the wind direction: the fire head, the flanks, and the tail.
[0091] In some embodiments, such as Figure 5 As shown, let the set of combustion cells at a certain high-level decision-making moment in the fire situation be . Then the geometric center of the fire is ,in Let c represent the two-dimensional coordinates of cell c. Let the environmental wind direction vector be... Its unit vector is To describe the angular relationship with the wind direction, a relative vector is introduced for any combustion cell c. (i.e., the vector from the center of mass of the fire to the cell), and the cosine of the angle between this vector and the wind direction. .like Figure 6 The diagram shown illustrates the direction of movement of the UAV. This paper uses the angle threshold method to analyze the combustion set. The area is divided into three zones: head, flanks, and rear. Let the head angle threshold be... and the flank detection angle threshold Then it is defined as follows.
[0092] (1) The firefighters gathered
[0093]
[0094] The combustion unit located downwind is usually the leading edge of the fire that spreads rapidly, reflecting the main direction of the fire's advance. Its length and spread rate can characterize the development trend of the fire.
[0095] (2) Tail gathering
[0096]
[0097] The combustion unit located upwind is usually on the upwind side of the ignition point or in the burn-out zone. Its shape can reflect whether the fire has been effectively contained.
[0098] (3) Flanking assembly
[0099]
[0100] The flanks are defined as combustion units that do not belong to the head or tail of the fire, and can reflect the degree of lateral spread of the fire.
[0101] In some embodiments, when modeling the drone model, the drone serves as the main body for performing forest fire perception and communication support, and its modeling is described from three aspects: motion, sensors, and power.
[0102] (1) Exercise
[0103] The drones perform actions in a discrete gridded environment. At each time step, any drone can move from its current position to any grid point in its eight neighborhoods, as shown in the diagram. Direction, such as Figure 4 As shown. Drones can choose to remain stationary. Unlike the continuous spread of a fire, drones update their actions more frequently. Before the fire evolves, drones can perform multiple actions consecutively, thus completing multiple steps within a single fire evolution cycle.
[0104] (2) Sensor
[0105] Each drone is equipped with multimodal cameras and sensors to perceive local fire conditions and environmental information. Key sensors include: a thermal imaging camera for real-time fire size identification. Meteorological sensors are used to measure wind speed v. Additionally, the UAV carries a radio transceiver for inter-drone communication and to acquire approximate location information of the initial ignition point at the start of the mission. After completing its maneuvers, the UAV's thermal imaging camera identifies the states of cells with different radii r centered on itself, based on the mission status, thus forming a local perception result. If the drone is performing a perception task, then r=2; otherwise, r=1.
[0106] (3) Electricity
[0107] Each drone has a limited battery capacity; the battery level of drone i at time t is denoted as . Movement and sensing consume power and cause... .when When the resource level falls below a threshold, the drone must resupply at a designated resupply station to ensure the continued execution of subsequent missions. Resource constraints are also incorporated into high-level decision-making and reward design as a crucial factor influencing mission switching.
[0108] In some embodiments, the rescuer model includes a movement model, a target point selection model, a path planning model, and a firefighting model, wherein:
[0109] (1) Moving model
[0110] Time is discretized using unit step sizes. To reflect the limited movement speed of personnel, it is assumed that a person is allowed to move only once after taking two steps, i.e., when... Only then can personnel k move from their current position. Move to a neighboring cell in its mole neighborhood Otherwise, remain stationary. Personnel are prohibited from entering burning cells. Let the set of burning cells at a certain high-level decision-making moment be denoted as . Therefore, the feasible move set is When movement is permitted, person k moves from... We have taken another step forward along the planned route.
[0111] (2) Target point selection model
[0112] Rescuers selected the current time step as the set of burning cells at the edge of the fire scene. This serves as a set of candidate target points. To avoid target conflicts among personnel and ensure fairness in allocation, this invention employs a target point allocation mechanism based on random order and comprehensive scoring. The specific steps are as follows:
[0113] ① At the start of each round of target allocation, the personnel index set is... Perform a random permutation to obtain the sequence. .
[0114] ② in sequence Target selection for personnel Select the target point.
[0115] ③ The set of candidate points that have not yet been selected by other people Calculate the number of rescuers separately For candidate points Overall rating:
[0116]
[0117] Where dist(,) is the Euclidean distance; This is a structural item used to describe the type of fire scene structure at the target point. ; This is an exclusion term used to prevent multiple people from clustering near adjacent or the same target. Let the set of people who have already selected target points in the current round of target allocation be . The target point selected by the kth assigned person in this round is The repulsion radius is defined as Then for any candidate boundary point The exclusion term is defined as
[0118] Among them, for personnel ,choose As the target point.
[0119] To ensure that the same target point is selected by at most one person, in the personnel Define the goal Then, remove that point from the candidate set of subsequent personnel.
[0120] (3) Path planning model
[0121] To enable rescuers to reach their respective target points as quickly as possible while avoiding entering the burning area, this invention employs a grid-based breadth-first search (BFS) path planning algorithm. After each round of target point allocation, the algorithm generates an unobstructed shortest path from the current position to the target boundary cell for each rescuer, as shown in the pseudocode below.
[0122] Algorithm 1: Breadth-First Search Path Planning Algorithm for Rescue Personnel Input: Current personnel position p_start, target point p_goal, set of passable grid cells G_free at the current time, neighboring cell function N(·), returns the moiré neighborhood of a given cell in G_free. Output: The path (raster sequence) from p_start to p_goal. 1: If p_start == p_goal then return [p_start] 2: Create an empty queue Q; create an empty set Visited; create a dictionary Prev 3: Enqueue p_start in Q; Visited ← {p_start}; Prev[p_start] ← NULL 4: While Q is not empty do: 5: g ← Q (Depart from team) 6: If g == p_goal then break 7: For each g' in N(g) do: 8: If g'∈G_free and g' has not been visited: 9: Enqueue g' to Q; Add g' to Visited; Prev[g'] ← g (records the predecessor) 10: End If 11: End For 12: End While 13: If p_goal is not in Prev, then return ∅. 14: Create an empty path list P; cur ← p_goal 15: While cur ≠ NULL do 16: Insert cur at the beginning of P. 17: cur ← Prev[cur] 18: End While
[0123] (4) Fire extinguishing model
[0124] When personnel k approach the target edge cell and reach a certain distance threshold, the fire extinguishing operation for that fire point is considered complete. Specifically, if at time step t, the following conditions are met... It is then assumed that personnel k has extinguished the target cell. The fire spreads, and the corresponding cell state changes from the burning state to the extinguished state.
[0125] In some embodiments, the communication model includes an inter-machine communication model and a human-machine communication model. Communication between UAV swarms is limited by spatial distribution and communication equipment performance, resulting in a finite communication radius. Let the effective communication distance between UAVs be... ,but:
[0126] (1) Inter-machine communication model
[0127] The information that can be exchanged between drones falls into two categories:
[0128] ① Small-scale status information: including location Battery Task status This type of information can be transmitted stably via low-bandwidth intranet links, without being limited by distance.
[0129] ②Large-scale sensing information: refers to local fire scene observations collected by multimodal cameras. At any time step t, if the Euclidean distance between UAV i and UAV j satisfies Then the two can exchange large amounts of sensory information. Assuming... Let be the fire map matrix of drone i at time t, then the fire map update formula is: ,in This indicates taking the maximum value element by element at the cell level.
[0130] (2) Human-computer communication model
[0131] The communication radius between ground firefighters' communication terminals and drones is also limited. Let the distance between personnel k and drone i satisfy... This is considered as personnel being able to access the aerial communication link via drones to obtain information about the fire scene.
[0132] In some embodiments, combined with Figure 2 ,like Figure 7 As shown, step 201 above can be implemented through step 201a.
[0133] Step 201a: In each time step, the control device of the drone swarm acquires the scene information of the mission scenario and the status information of the drone swarm.
[0134] In some embodiments, the aforementioned time step is the basic working cycle or decision rhythm of the control system.
[0135] In some embodiments, at each time step, the control device of the drone swarm acquires scene information of the mission scenario and state information of the drone swarm to complete a comprehensive data update.
[0136] In this way, by acquiring data at each time step, the control device of the drone swarm can continuously capture changes such as fire spread, personnel movement, and drone wear and tear. Based on the new state after the control strategy output at the previous time step is executed, it can make decisions for the next round, forming a closed loop of "perception-decision-execution-re-perception".
[0137] Step 202: The control device of the drone swarm inputs scene information and status information into the drone control model.
[0138] In some embodiments, the above-described drone control model includes a low-level task controller and a high-level task switcher.
[0139] In some embodiments, the above-described UAV control model is a core algorithm module for achieving intelligent decision-making, employing a layered architecture and comprising two key sub-modules: a low-level task controller and a high-level task switcher.
[0140] 1. The low-level task controller is used to receive the original scene and state information, and its core responsibility is to perform feature extraction and fusion.
[0141] 2. The high-level task switcher receives global feature information processed by the low-level task controller and, based on these features, performs macro-level task coordination and allocation for the entire UAV swarm, determining whether each UAV should perform the "perception", "communication coverage" or "resupply" sub-task at the current moment.
[0142] The objective of this application's embodiments is to extinguish all fire spots at the edge of the fire scene, i.e. At the same time, maximize the weighted sum of the average fire detection rate of the drone swarm and the average communication coverage rate of rescue personnel throughout the entire mission cycle. Define the performance index of the communication-sensing fusion as follows: Where E(t) is the exploration rate at time t, defined as , C(t) represents the population coverage rate at time t, defined as... , This problem can be modeled as follows:
[0143]
[0144] in, All drones must have a remaining battery level of at least 0. All drones must have a remaining payload of at least 0. All edge fire points must be extinguished by the end of the mission.
[0145] In some embodiments, the above optimization problem The aim is to maximize the performance index J of sensor-communication fusion while meeting the power constraints of the drone. and terminal fire scene status constraints This problem is a time-series decision-making problem: at each discrete time step t, the drone selects an action based on the current environmental state. The decision not only affects the immediate reward but also determines the fire distribution and resource status at the next moment. Variables such as fire evolution, drone position, and battery power change dynamically over time, and the state transition exhibits the Markov property, meaning the current state... With action The next state can then be determined. The distribution of the distribution is such that the optimization problem can be formalized as a Markov Decision Process (MDP).
[0146] In some embodiments, since the low-level strategy in this invention has been determined through heuristic algorithms, only task-level decision learning is required at the high level. The low level includes three sub-tasks: perception, communication coverage, and resupply, which respectively employ heuristic rules such as flight along the boundary, approaching the optimal coverage location, and approaching the nearest resupply point for action decisions. Therefore, this dynamic optimization problem can be further transformed into a hierarchical Markov decision process. In this model, the high-level decision-maker selects the task subspace (Option) based on the global fire situation to achieve task-level scheduling; the low-level controller generates a continuous sequence of actions under the selected task and executes specific maneuvers and firefighting operations.
[0147] In some embodiments, the entire hierarchical MDP is denoted as ⟨S,A,P,R,γ>, where S, A, P, and R represent state, action, transition, and reward elements, respectively, and γ is a discount factor, the value of which will be given in the subsequent parameter configuration section. S, A, P, and R are defined below.
[0148] 1. Status Input S
[0149] Low-level motion controller at atomic step The state input is defined as
[0150]
[0151] in The location of the UAV (discrete grid coordinates). Map showing the fire situation. For the location aggregation of other drones, For electricity, Let be the set of personnel locations, representing the two-dimensional position coordinates of all ground rescue personnel at time t, i.e. .
[0152] The state of the high-level task selector is obtained from the low-level state through feature extraction, and is defined as follows: The detailed definition is given below:
[0153] (1) Status of the unmanned aerial vehicle
[0154]
[0155] Where (x, y) represents the drone's position, e represents the battery level, and q represents the Euclidean distance from the drone to the centroid of all personnel positions, where the centroid of all personnel positions is... Defined as The distance from drone i to this average position .
[0156] (2) Environmental conditions
[0157]
[0158] The following are defined as: fire scene exploration rate (E), personnel coverage rate (C), and wind speed. Fire growth Fire length characteristics , , ,in This indicates the number of elements in the set.
[0159] (3) Team Status
[0160]
[0161] in This is the proportion vector of each task in the previous high-level time step. The shortest distance between the center of mass of the drone swarm and the crowd: ; This represents the average battery level of the drone swarm.
[0162] 2. Motion space
[0163] High-level action space These correspond to three subtasks: perception, communication coverage, and resupply; the low-level action space. These correspond to moving in 8 directions and staying in place, respectively.
[0164] The following section will detail the execution logic of the three subtasks, as well as the mapping mechanism from high-level options to specific low-level actions. The option selected for drone i at time t.
[0165] (1) Perception
[0166] When high-level options During this process, the drones perform a perception sub-task to quickly establish fire scene awareness, providing information for firefighters to divide their firefighting tasks. When performing this sub-task, each drone will appropriately increase its flight altitude to expand the coverage of visual and thermal imaging, and at each time step, it will collect local fire information of a 5×5 grid centered on itself. At the same time, it will fuse its own observations with those of neighboring drones through group communication to update its internal fire map.
[0167] When the drone performs a perception task, its movement follows a clockwise circle around the fire boundary. Let the drone's current position be... Select the nearest fire boundary point within its perception range: To maintain clockwise rotation, the UAV calculates the local tangent direction based on its current relative position to the boundary. And take it as the target direction: The drone selects the action closest to the target direction: ,in Let be the unit direction vector corresponding to action a.
[0168] (2) Communication coverage
[0169] When high-level options At this time, the UAV performs the communication coverage sub-task. All UAVs that have selected the communication task undergo a centralized planning process to determine their communication station locations. Let the set of UAVs currently participating in the communication task be denoted as . The ground rescue personnel are assembled at the following locations: Candidate site set First, Random rearrangement yields a sequence and initialize the set of covered personnel as For the m-th communication UAV in the sequence. At candidate site locations The above calculation covers the set of people.
[0170]
[0171] in Define the communication radius of the drone and define the station location score as follows.
[0172]
[0173] in To cover the basic income of an individual, To punish overlapping personnel already covered by preceding drones, For the movement cost weight, This shows the current location of the drone.
[0174] Communication drones Select the station with the highest score as the communication target point: And add the people covered by it to the already covered set: The target direction vector of UAV i is defined as follows: The drone selects the action closest to the target direction: If the drone has reached the vicinity of the target (distance less than or equal to 1 grid), it will remain hovering.
[0175] (3) Supply
[0176] When high-level options At this time, the drone performs a resupply subtask. When the drone selects this task, its low-level control strategy will calculate the range of all resupply points from its current position. Calculate the Euclidean distance and select the nearest supply point. Subsequently, the drone defined the target direction as The drone selects the direction closest to the target. When the drone arrives at the resupply point, it will automatically perform the resupply operation. This means restoring the battery to its maximum capacity.
[0177] 3. State transition function P
[0178] The high-level strategy selects a subtask at time t. And controlled by low-level strategies Perform K consecutive steps in the environment until the task's termination condition is met or the maximum number of execution steps is reached. The lower-level state evolves over time as follows: When the lower layer completes its task, its final state is... via function The inputs are aggregated into higher-level states, forming higher-level state transitions: ,in It represents a comprehensive mapping of the results of multiple interactions between higher and lower layers in the environment, including the cumulative feedback from lower-level actions, changes in fire intensity, and other factors.
[0179] 4. Reward R
[0180] The various weighting parameters mentioned in this section All values are adjustable, and recommended value ranges and default values will be provided at the end of this section to facilitate calibration and reproduction experiments under different fire conditions and resource requirements. The cumulative high-altitude reward for the i-th drone during the high-altitude option execution period is:
[0181]
[0182] in For drone i, the low-level reward at step t, This is the discount factor.
[0183] Let the high-level option of drone i at time t be... Low-level single-step reward Defined as:
[0184]
[0185] The definitions of each part are as follows:
[0186] (1) Perceptual reward
[0187] The perceptual reward for a single step is defined as:
[0188]
[0189] in For the current step, the drone i has never previously discovered the set of fire grids. This is a set of lattices that were previously discovered by drone i, but the observed fire size has changed.
[0190] (2) Communication coverage bonus
[0191] Single-step communication coverage bonus depends on the marginal coverage population, defined as:
[0192]
[0193] in For the group of people covered by drone i, M represents the total number of people covered by "all other drones".
[0194] (3) Charging rewards
[0195] The single-step charging reward is defined as:
[0196]
[0197] (4) Parameter settings
[0198] Weights in the reward function The settings are optimized through simulation to balance the importance of different tasks. Typical value ranges and meanings are shown in Table 1.
[0199]
[0200] Table 1
[0201] Step 203: The control device of the UAV swarm extracts global feature information from scene information and status information through the low-level task controller.
[0202] In some embodiments, the global feature information mentioned above includes scene feature information and drone feature information.
[0203] In some embodiments, the aforementioned global feature information is a set of quantitative or structured descriptions that can summarize the current overall situation of the entire system, including scene feature information and UAV feature information, wherein:
[0204] 1. Scene feature information refers to the global environmental state related to the task objective extracted from scene information.
[0205] 2. UAV characteristic information refers to the features extracted from status information that reflect the overall capability and resource allocation of the UAV fleet.
[0206] Step 204: The control device of the drone swarm assigns at least one sub-task to each drone in the drone swarm based on global feature information through a high-level task switcher.
[0207] In some embodiments, the at least one sub-task mentioned above includes three basic types: "perception," "communication coverage," and "resupply." Each drone can be assigned at least one sub-task at any given time to achieve a clear division of responsibilities.
[0208] In some embodiments, the at least one sub-task described above includes at least one type of sensing, communication coverage, and resupply. (Combined) Figure 2 ,like Figure 8 As shown, step 204 above can be implemented through steps 204a and 204b.
[0209] Step 204a: The control device of the UAV swarm inputs global feature information into the high-level policy network in the high-level task switcher and outputs the task score of each type of sub-task corresponding to at least one type for each UAV in the UAV swarm.
[0210] In some embodiments, the high-level policy network described above is the core algorithm implementation of the high-level task switcher, and can be a trained deep learning model, such as a multi-agent double dueling deep Q-network (MADDQN).
[0211] In some embodiments, the task score is a quantitative evaluation value that represents the expected long-term benefit that the high-level policy network believes, based on its training experience, can bring by assigning "drone i" to perform "subtask type j" in the current global situation. The higher the score, the more effective and appropriate the assignment is in the current state.
[0212] Step 204b: The control device of the drone swarm assigns the sub-task with the highest task score for each drone to the corresponding drone.
[0213] In some embodiments, for each drone in the list, the drone swarm's controller can independently select the subtask with the highest task rating for that drone.
[0214] In this way, the control device of the drone swarm can select the sub-task with the highest task score for each drone, that is, select the task type that the model judges to be the most suitable and expected to yield the greatest benefit for each drone in the current state, thus maximizing the overall long-term return.
[0215] Step 205: The control device of the UAV swarm generates and outputs control strategy information for the UAV swarm for each subtask in at least one subtask through the low-level motion controller.
[0216] In some embodiments, the control strategy information described above is a set of underlying instructions that can directly drive the drone's actuators.
[0217] In some embodiments, the control unit of the drone swarm can calculate the specific control sequence required to implement each subtask, i.e., control strategy information, based on the subtask type received from each drone, using a low-level motion controller. The control unit can then send these control sequences to the corresponding drones for execution.
[0218] In some embodiments, the framework for solving the UAV swarm collaborative communication-sensing H-MDP problem in forest fire scenarios, as described in this application, is as follows: Figure 9 As shown.
[0219] 1. A hierarchical reinforcement learning solution method based on MADDQN:
[0220] In some embodiments, to solve the H-MDP model described above, this invention employs the multi-agent distributed reinforcement learning framework MADDQN to train the high-level task switching policy. Each agent corresponds to a drone, and its high-level policy network structure consists of two fully connected layers, each containing 256 neurons, with ReLU activation functions used between layers. The network input is the high-level observation. The output is the corresponding options. The Q-value vector.
[0221] In some embodiments, MADDQN maintains an online network in each agent. ( ) and target network ( The parameters of both are periodically synchronized, and their update rules are as follows: ,in These are soft update coefficients. In each training iteration, a batch of samples is randomly sampled from the empirical replay pool D. Calculate the TD objective of Double DQN:
[0222] In some embodiments, to further enhance the expression of the importance of different task options, the present invention employs a Dueling structure to represent state values. With the option advantage function Separate modeling:
[0223]
[0224] Where O is the set of all possible options.
[0225] Network parameter updates follow the mean squared error loss:
[0226]
[0227] In some embodiments, the loss is minimized using the Adam optimizer with a learning rate of 0.001. During training, all drones share network parameters to improve sample efficiency. The -greedy strategy is used for exploration, with an exploration rate of... The decay is linear within the range [1, 0.05]. At each time step, the agent selects an option based on the current high-level state. And execute a K-step low-level action sequence in the environment. At the end of the execution period, the cumulative reward is calculated. The high-level rewards are written into the replay pool. The experience replay mechanism is used to break sample correlations, and its capacity is set to... Each time, 128 experiences are sampled and updated. The discount factor is set to γ=0.99.
[0228] 2. Algorithm Pseudocode
[0229] In some embodiments, the MADDQN-based algorithm is as shown in Algorithm 2:
[0230] Algorithm 2: High-level training process (SMDP, Dueling Double DQN) Input: Number of training episodes (num_episodes), maximum number of steps (max_steps), number of options (n_options), length of K steps (K). Output: The trained high-level policy , 1: Initialize the environment (env); initialize high-level policies. Target Strategy Playback pool D 2: For episode = 1 … num_episodes do 3: obs_low ← env.reset(); t ← 0 4: While t < max_steps do 5: For each drone i do 6: Construct high-level observations based on low-level observations. 7: If the random number b < then = Random options 8: else = ; o.append( ) 9: End If 10: End For 11: For k=1....K do 12: Map the heuristic low-level policy to the action set a based on the option set o. 13: Interacting with the environment: next_obs_low, r, done = env.step(a) 14: obs_low ← next_obs_low; t←t + 1 15: If then break 16: else ← + disc·r; disc←disc·γ 17: End If 18: End For 19: Based on the updated environment, construct the next high-level observation. 20: For each drone i do 21: Write an SMDP migration to the replay pool D.push( , , , (done) 22: End for 23: Sample from the playback pool, update the online network according to the Double DQN target, and periodically update the target network. 24: If done, then break. 25: End While 26: End for
[0231] In the UAV swarm control method provided in this application, since the lower-level controller can extract global scene feature information, that is, it can extract all information in the current scene, the higher-level switcher can dynamically allocate the most suitable sub-task to each UAV based on the global scene feature information. This allows the UAV swarm to avoid task overlap and perform replenishment planning, thereby accurately allocating limited UAV resources to the most needed tasks and realizing the collaborative trade-off of multiple types of sub-tasks, thus improving the overall resource allocation efficiency.
[0232] In some embodiments, combined with Figure 2 ,like Figure 10 As shown, after step 205 above, the control method for a drone swarm provided in this application embodiment may further include the following step 301.
[0233] Step 301: When at least one subtask includes perception, the control device of the UAV swarm generates a first control command based on the control strategy information.
[0234] In some embodiments, the first control command described above is used to control the UAV to collect scene information and circumnavigate along the scene.
[0235] In some embodiments, the above-mentioned at least one sub-task including perception means that when the control device of the drone swarm assigns tasks to the drone swarm through a high-level task switcher, at least one drone is assigned the "perception" task.
[0236] In some embodiments, the first control command is an instruction to control the UAV to perform a "perception" task, mainly including a data acquisition instruction and a motion trajectory instruction, wherein:
[0237] 1. Data acquisition commands can control the drone to turn on and configure its onboard sensors, such as visible light cameras and infrared thermal imagers, to collect scene information at specific frequencies and modes.
[0238] 2. Motion Trajectory Command: Control the drone to navigate around the scene. In emergency scenarios, control the drone to cruise along the boundaries or front lines of fire or disaster areas according to preset rules, such as clockwise, to achieve continuous and efficient scanning coverage of key areas.
[0239] In this way, the control device of the drone swarm can instantiate the "sensing" task into a specific sequence of actions that can drive a single drone to complete the "sensing" task, so as to achieve real-time acquisition of disaster information.
[0240] In some embodiments, combined with Figure 2 ,like Figure 11 As shown, after step 205 above, the control method for a drone swarm provided in this application embodiment may further include the following step 401.
[0241] Step 401: In cases where at least one subtask includes communication coverage, the control unit of the UAV swarm generates a second control command based on the control strategy information.
[0242] In some embodiments, the second control command described above is used to control the drone to move to the target location.
[0243] In some embodiments, the inclusion of communication coverage in at least one subtask means that when the control device of the drone swarm assigns tasks to the drone swarm through a high-level task switcher, at least one drone is assigned the "communication coverage" task.
[0244] In some embodiments, the second control command described above is a command to control the UAV to perform a "communication coverage" task, mainly including waypoint navigation commands and position-keeping commands, wherein:
[0245] 1. Waypoint navigation instructions: Guide the drone to the optimal relay point that can effectively cover the communication needs of rescue personnel.
[0246] 2. Hovering command: When the drone arrives near the target location, the hovering command is used to stabilize it in that position and make it act as a fixed communication relay node.
[0247] In this way, the control device of the drone swarm can transform the "communication coverage" task into precise point deployment instructions, control the drone swarm to build a dynamic aerial communication network, and ensure the stable transmission of rescue instructions and disaster information.
[0248] In some embodiments, combined with Figure 2 ,like Figure 12 As shown, after step 205 above, the control method for a drone swarm provided in this application embodiment may further include the following step 501.
[0249] Step 501: In cases where at least one sub-task includes resupply, the control unit of the UAV swarm generates a third control command based on the control strategy information.
[0250] In some embodiments, the third control command described above is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
[0251] In some embodiments, the above-mentioned at least one sub-task including resupply means that when the control device of the drone swarm assigns tasks to the drone swarm through a high-level task switcher, at least one drone is assigned the "resupply" task.
[0252] In some embodiments, the aforementioned third control command is an instruction to control the UAV to perform a "resupply" mission, mainly including a return-to-home command and a resupply operation command, wherein:
[0253] 1. Return-to-Home Command: Controls the drone to "move along the optimal path to the nearest resupply location". For example, it can fly along a straight line or a safe path to a preset charging station or mobile charging platform.
[0254] 2. Replenishment Operation Command: When the drone detects that it has arrived at the replenishment location, such as through visual tags or precise positioning, it will automatically trigger "execute replenishment operation", that is, control the drone to land, dock and start charging until the power is restored to the set threshold, such as full power.
[0255] In this way, the control device of the drone swarm can actively schedule low-battery drones to leave the mission sequence and autonomously complete energy replenishment through resupply commands, thereby supporting the drone swarm to maintain the sustainability of its overall combat capability in long-term, multi-round emergency missions and ensuring the long-term stable operation of the drone swarm.
[0256] In some embodiments, to verify the effectiveness of the CS-HRL method, this paper constructs a realistic forest fire scenario, with parameter settings shown in Table 2 below. The terrain selected for the experiment is derived from real mountain and forest landform data of the Greater Khingan Mountains in China to ensure the authenticity of the spatial distribution and propagation path of the fire. Meteorological conditions are set as a typical dry summer environment, with the wind direction set to due east and the wind force randomly ranging from 4 to 6.
[0257]
[0258] Table 2
[0259] In some embodiments, regarding fire source initialization, in order to simulate the process in reality where small-scale fires caused by lightning strikes, open flames, etc., gradually evolve into large-scale forest fires, a 4×4 grid is lit before the experiment begins, and the fire spreads randomly for several time steps according to the fire spread model, so as to approximate the time delay between the fire department observing the fire and actually arriving at the scene.
[0260] In some embodiments, regarding drone initialization, referencing common drone parameters, the maximum battery capacity of a single drone can be set to 75 watt-hours (Wh), corresponding to approximately 30 minutes of flight time. Each time step in the simulation is 15 seconds, therefore a single drone can run a maximum of 120 steps. To improve the generalization ability of the strategy, the initial position of the drone is randomly set during the training phase.
[0261] In some embodiments, a comparative experiment between the fixed task allocation strategy and the CS-HRL dynamic task switching strategy is described below.
[0262] The following describes the control method for a drone swarm according to specific embodiments of this application.
[0263] In some embodiments, to verify the performance of the CS-HRL dynamic task switching method in terms of communication-sensory collaboration, this embodiment constructs a comparative experiment under uniform fire conditions and drone quantity configuration. The number of drones in the experimental environment is fixed at 3, numbered 1, 2, and 3. The initial state of the fire is set as a 4×4 point fire source, which spreads for 40 steps under wind conditions of level 6 before the task process is initiated.
[0264] In some embodiments, a fixed task allocation strategy is set as the baseline scheme under this condition: UAV 1 performs communication coverage tasks, and UAVs 2 and 3 perform fire exploration tasks; when the UAV battery level is below 10Wh, it switches to a resupply task. Under the condition that the above fixed strategy is completely consistent with the conditions of the CS-HRL method, the fire exploration rate and rescue personnel coverage rate at each time step during task execution are recorded, and the results are as follows: Figure 13 As shown, the evolution of fire exploration rate-personnel coverage over time is illustrated for the CS-HRL method and the fixed task assignment method.
[0265] In some embodiments, from Figure 13 It can be observed that in the early stages of the mission, the exploration rate of the CS-HRL method increases significantly faster than that of the fixed strategy. Because the ratio of communication to exploration tasks is fixed, the fixed strategy has difficulty concentrating its efforts on improving situational awareness during the phase when the communication load is relatively light.
[0266] In some embodiments, as the mission progresses, the CS-HRL method gradually increases the priority of communication coverage tasks after the exploration rate approaches saturation, enabling the coverage rate to recover quickly and remain at a high level. In contrast, the fixed strategy, by maintaining a fixed task ratio, cannot dynamically adjust communication resource allocation based on the location of rescue personnel, resulting in a significant drop in coverage rate in the middle stage, and a much slower recovery speed than the CS-HRL method. In the later stages of the mission, the CS-HRL method's exploration rate steadily increases, while the coverage rate shows a clear upward trend, demonstrating that it achieves a dynamic balance between communication and exploration capabilities through task switching and resupply alternation mechanisms. In contrast, the fixed strategy, due to the inability to switch roles, experiences coverage gaps once the communication drone returns for resupply, causing the coverage rate curve to remain persistently low.
[0267] In some embodiments, Figure 14This paper compares the CS-HRL method and a fixed strategy, visually demonstrating the differences in UAV collaborative behavior at different time stages. At t=0, rescue personnel are generally within the command center's communication coverage area, with low communication needs. The CS-HRL method assigns all UAVs to perception tasks to maximize the efficiency of fire scene information acquisition. In contrast, in the fixed strategy, some UAVs continuously perform communication tasks, resulting in relatively limited utilization of exploration resources. At t=30, rescue personnel begin to gradually leave the command center's communication coverage area, but their spatial distribution remains relatively concentrated. The CS-HRL method only schedules one UAV to perform communication coverage tasks, while the remaining UAVs continue fire scene perception, thus maintaining high exploration efficiency while ensuring basic communication. By t=60, rescue personnel are further dispersed, and communication needs increase significantly. The CS-HRL method automatically increases the weight of communication tasks, assigning two UAVs to perform communication coverage tasks, while the remaining UAV continues fire scene perception, achieving parallel collaboration between communication support and information acquisition. In the later stages of the mission (t=90), some communication drones entered the resupply phase due to insufficient power. The CS-HRL method can dynamically adjust task allocation, allowing drones originally performing perception tasks to switch to communication tasks to take their place, thereby maintaining the continuity of personnel communication coverage. In contrast, the fixed strategy resulted in a significant decrease in personnel coverage due to the communication drones returning for resupply and the inability to switch roles. This comparison shows that the CS-HRL method has stronger cooperative scheduling and system robustness under resource-constrained conditions.
[0268] In some embodiments, the exploration rate and coverage experiments under different numbers of drones and initial fire scales are described below.
[0269] In some embodiments, Table 3 shows the comparison results of different numbers of drones and initial fire sizes. In the experiments, the initial fire areas were 141, 195, and 261, corresponding to the fire sizes after 30, 40, and 50 time steps of natural spread from a 4×4 initial fire point under wind force 6 conditions, simulating the arrival of firefighting forces in the early, middle, and late stages of a fire. Table 3 is shown below:
[0270]
[0271] Table 3
[0272] The results show that under the three initial fire scale scenarios, as the number of drones increases, the exploration rate, coverage rate, and their combined indicators all show a stable upward trend, indicating that increasing drone resources can effectively enhance the system's overall capabilities in fire scene perception and communication support. However, the performance improvement varies significantly across different number ranges. When the number of drones increases from 2 to 4, the combined indicators show a significant improvement under all three fire scale scenarios. This indicates that under relatively limited resources, adding drones can effectively alleviate resource competition between fire scene perception and communication coverage. The performance improvement at this stage mainly stems from the method's dynamic allocation capability for multiple tasks, meaning that additional drones can be promptly dispatched to more urgent sub-tasks based on changes in the fire scene situation and personnel distribution, thereby amplifying the overall collaborative benefits. When the number of drones further increases from 4 to 6, the combined indicators still maintain an upward trend, but the increase is relatively smaller, indicating that under given fire scene scale and personnel configuration, the system gradually approaches the saturation state of the need for sensory collaboration, and the marginal benefits brought by adding drones begin to decline. This phenomenon demonstrates that the proposed method does not rely on simply increasing the number of resources to achieve performance improvements. Instead, it avoids redundant investment through reasonable scheduling when resources are sufficient, thus maintaining the stability and efficiency of system operation. Furthermore, although the overall indicator level decreased when the initial fire size increased from 141 to 261, the relative change trend remained consistent under different numbers of drones, indicating that the proposed method can reliably function under various fire conditions, including the early, middle, and late stages of a fire.
[0273] The results above show that the proposed method has good adaptability and robustness under different fire scales and UAV resource configurations, and can achieve efficient coordination between fire scene perception and communication support.
[0274] This application proposes a CS-HRL based on communication-sensing collaboration and provides an application example in a fire scenario. The example aims to maximize the fire scene exploration rate and the communication coverage rate for rescue personnel, and explicitly considers the power constraints of the UAVs in the model. The proposed method adopts a two-level structure of high-level and low-level: the high-level task switcher dynamically allocates sub-tasks such as perception, communication coverage, and resupply based on the global situation and the structural characteristics of the fire scene; the low-level action controller generates specific execution actions based on the high-level decision and fine-grained observation using heuristic strategies, thereby realizing the autonomous collaboration of the UAV swarm.
[0275] The first protection point of this application is the proposal of a communication-sensing collaborative multi-task modeling method for emergency scenarios in mountainous forests. Addressing the characteristics of limited communication conditions, dynamically changing emergency task requirements, and limited UAV resources in mountainous forest environments, this method unifies and abstracts the different functions of UAVs in emergency rescue processes—such as disaster awareness, personnel communication coverage, and energy replenishment—into switchable task states, and then collaboratively models them within the same decision space. This modeling method breaks through the existing approach of separating or sequentially executing perception and communication tasks, enabling the task organization of UAV swarms to directly reflect the coupling relationship between communication and sensing needs in emergency scenarios, providing a unified problem description basis for subsequent collaborative decision-making.
[0276] The second protection point of this application embodiment lies in the construction of a hierarchical UAV swarm collaborative decision-making framework based on the aforementioned communication-sensory collaborative task modeling, used to solve the problems of task selection and task execution respectively. At the high-level decision-making level, decisions are made on the selection and switching of UAVs among tasks such as perception, communication coverage, and resupply based on the communication-sensory task status and changes in the emergency situation. At the low-level execution level, under conditions of limited communication, specific execution behaviors that meet task constraints are generated based on the high-level decision-making results and combined with the UAVs' local observations, motion constraints, and energy states. By decoupling the functions of "task selection" and "task execution," this hierarchical framework enables the UAV swarm to maintain good adaptability and scalability when the scale changes, the task expands, and the scene migrates, forming a communication-sensory collaborative multi-task decision-making protection structure for mountain and forest emergency scenarios.
[0277] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there is no conflict, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0278] As can be seen, the above mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0279] This application embodiment can divide the control device of the UAV swarm into functional modules according to the above method example. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. Optionally, the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0280] In some embodiments, this application also provides a control device for a drone swarm. The control device may include one or more functional modules for implementing the drone swarm control method described in the above embodiments.
[0281] For example, Figure 15 This is a schematic diagram of a control device for an unmanned aerial vehicle (UAV) swarm provided in an embodiment of this application. Figure 15 As shown, the control device 900 for the UAV swarm includes: an acquisition module, an input module, an extraction module, an allocation module, and a generation module.
[0282] The aforementioned acquisition module is used to acquire scene information and status information of the drone swarm's mission scenario. The drone swarm includes at least one drone. The scene information includes scene map information and personnel location information. The status information includes drone location information and drone battery level information.
[0283] The aforementioned input module is used to input scene information and status information into the UAV control model, which includes a low-level task controller and a high-level task switcher.
[0284] The extraction module described above is used to extract global feature information from scene information and state information through a low-level task controller. This global feature information includes scene feature information and UAV feature information.
[0285] The aforementioned allocation module is used to allocate at least one sub-task to each UAV in the UAV swarm based on global feature information through a high-level task switcher.
[0286] The aforementioned generation module is used to generate and output control strategy information for the UAV swarm for each subtask in at least one subtask through a low-level action controller.
[0287] In some embodiments, the acquisition module is specifically used to acquire scene information of the task scenario and state information of the drone swarm at each time step.
[0288] In other embodiments, the at least one subtask includes at least one type of perception, communication coverage, and resupply. The allocation module is specifically configured to: input global feature information into the high-level policy network of the high-level task switcher, output a task score for each type of subtask corresponding to each UAV in the UAV swarm; and allocate the subtask of the type with the highest task score for each UAV to the corresponding UAV.
[0289] In some other embodiments, the generation module is further configured to generate a first control command based on the control strategy information after the generation module generates and outputs control strategy information for each subtask in at least one subtask through a low-level motion controller, wherein at least one subtask includes perception. The first control command is used to control the drone to collect scene information and circumnavigate along the scene.
[0290] In some other embodiments, the generation module is further configured to generate a second control command based on the control strategy information after the generation module generates and outputs control strategy information for each subtask in at least one subtask through the low-level motion controller, wherein at least one subtask includes communication coverage. The second control command is used to control the drone to move to the target location.
[0291] In some other embodiments, the generation module is further configured to generate a third control instruction based on the control strategy information after the generation module generates and outputs control strategy information for each subtask in at least one subtask through the low-level motion controller. The third control instruction is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
[0292] It should be noted that the control device for the drone swarm can implement all the processes implemented in the above method embodiments and achieve the same beneficial effects. To avoid repetition, it will not be described again here.
[0293] In the case where the functions of the integrated modules described above are implemented in hardware, this application provides a possible structural schematic diagram of the electronic device involved in the above embodiments. For example... Figure 16 As shown, the electronic device 90 includes: a processor 92, a communication interface 93, and a bus 94. Optionally, the electronic device 90 may also include a memory 91.
[0294] Processor 92 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0295] Communication interface 93 is used to connect with other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0296] The memory 91 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0297] As one possible implementation, the memory 91 can exist independently of the processor 92. The memory 91 can be connected to the processor 92 via a bus 94 and is used to store instructions or program code. When the processor 92 calls and executes the instructions or program code stored in the memory 91, it can implement the UAV swarm control method provided in the embodiments of this application.
[0298] In another possible implementation, memory 91 can also be integrated with processor 92.
[0299] Bus 94 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 94 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 16 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0300] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the service calling device can be divided into different functional modules to complete all or part of the functions described above.
[0301] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described UAV swarm control method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0302] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0303] This application also provides a readable storage medium storing a program or instructions. When executed by a computer, the program or instructions implement the UAV swarm control method provided in the above embodiments. It is understood that all or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware. The readable storage medium can be any of the foregoing embodiments or memory. The readable storage medium can also be an external storage device of the service invocation device, such as a pluggable hard drive, SmartMedia Card (SMC), Secure Digital (SD) card, flash card, etc., equipped on the service invocation device. Further, the readable storage medium can include both internal storage units of the service invocation device and external storage devices. The readable storage medium is used to store the computer program and other programs and data required by the service invocation device. The readable storage medium can also be used to temporarily store data that has been output or will be output.
[0304] This application also provides a computer program product, which is stored in a storage medium and, when executed by a computer, implements the drone swarm control method provided in the above embodiments.
[0305] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0306] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0307] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for controlling a swarm of unmanned aerial vehicles (UAVs), characterized in that, include: Acquire scene information of the mission scenario of the drone swarm and status information of the drone swarm, wherein the drone swarm includes at least one drone, the scene information includes scene map information and personnel location information, and the status information includes drone location information and drone battery information; The scene information and the state information are input into the UAV control model, which includes a low-level task controller and a high-level task switcher. The low-level task controller extracts global feature information from the scene information and the state information. The global feature information includes scene feature information and UAV feature information. The high-level task switcher assigns at least one sub-task to each drone in the drone swarm based on the global feature information; The low-level motion controller generates and outputs control strategy information for the UAV swarm for each of the at least one subtask.
2. The control method for a swarm of unmanned aerial vehicles according to claim 1, characterized in that, The acquisition of scene information of the task scenario and status information of the drone swarm includes: Within each time step, acquire scene information of the mission scenario of the drone swarm and the status information of the drone swarm.
3. The control method for a swarm of unmanned aerial vehicles according to claim 1, characterized in that, The at least one subtask includes at least one type of sensing, communication coverage, and resupply; The process of assigning at least one sub-task to each drone in the drone swarm based on the global feature information via a high-level task switcher includes: Global feature information is input into the high-level policy network in the high-level task switcher, and task scores for each type of sub-task of at least one type are output for each UAV in the UAV swarm. The subtask with the highest task score for each drone is assigned to the corresponding drone.
4. The control method for a drone swarm according to claim 3, characterized in that, After generating and outputting control strategy information for the UAV swarm for each of the at least one subtask through the low-level action controller, the method further includes: If at least one subtask includes perception, a first control command is generated based on the control strategy information. The first control command is used to control the UAV to collect scene information and circumnavigate along the scene.
5. The control method for a swarm of unmanned aerial vehicles according to claim 3, characterized in that, After generating and outputting control strategy information for the UAV swarm for each of the at least one subtask through the low-level action controller, the method further includes: If at least one subtask includes communication coverage, a second control command is generated based on the control strategy information. The second control command is used to control the UAV to move to the target location.
6. The control method for a drone swarm according to claim 3, characterized in that, After generating and outputting control strategy information for the UAV swarm for each of the at least one subtask through the low-level action controller, the method further includes: If at least one subtask includes resupply, a third control command is generated based on the control strategy information. The third control command is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
7. A control device for a swarm of unmanned aerial vehicles (UAVs), characterized in that, include: The module includes: acquisition module, input module, extraction module, allocation module, and generation module. The acquisition module is used to acquire scene information of the mission scenario of the drone swarm and status information of the drone swarm, wherein the drone swarm includes at least one drone, the scene information includes scene map information and personnel location information, and the status information includes drone location information and drone battery information. The input module is used to input the scene information and the state information into the UAV control model, which includes a low-level task controller and a high-level task switcher. The extraction module is used to extract global feature information from the scene information and the state information through the low-level task controller. The global feature information includes scene feature information and UAV feature information. The allocation module is used to allocate at least one sub-task to each drone in the drone swarm based on the global feature information through the high-level task switcher. The generation module is used to generate and output the control strategy information of the UAV swarm for each of the at least one subtask through the low-level action controller.
8. The control device for a swarm of unmanned aerial vehicles according to claim 7, characterized in that, The acquisition module is specifically used to acquire scene information of the mission scenario of the drone swarm and the status information of the drone swarm at each time step.
9. The control device for a drone swarm according to claim 7, characterized in that, The at least one subtask includes at least one type of sensing, communication coverage, and resupply; The allocation module is specifically used for: Global feature information is input into the high-level policy network in the high-level task switcher, and task scores for each type of sub-task of at least one type are output for each UAV in the UAV swarm. The subtask with the highest task score for each drone is assigned to the corresponding drone.
10. The control device for a drone swarm according to claim 9, characterized in that, The generation module is further configured to, after generating and outputting control strategy information of the UAV swarm for each of the at least one subtask through the low-level motion controller, generate a first control command based on the control strategy information when the at least one subtask includes perception. The first control command is used to control the UAV to collect scene information and circumnavigate along the scene.
11. The control device for a drone swarm according to claim 9, characterized in that, The generation module is further configured to, after generating and outputting control strategy information for the UAV swarm for each of the at least one subtask through the low-level motion controller, generate a second control command based on the control strategy information, provided that the at least one subtask includes communication coverage, the second control command being used to control the UAV to move to the target location.
12. The control device for a drone swarm according to claim 9, characterized in that, The generation module is further configured to, after generating and outputting control strategy information for the UAV swarm for each of the at least one subtask through the low-level motion controller, generate a third control command based on the control strategy information if the at least one subtask includes resupply. The third control command is used to control the UAV to move to the nearest resupply location and perform a resupply operation.
13. An electronic device, characterized in that, It includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the control method for a drone swarm as described in any one of claims 1-6.
14. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a computer, implement the control method for a drone swarm as described in any one of claims 1-6.
15. A computer program product, characterized in that, The computer program product is stored in a storage medium, and when executed by a computer, the computer program product implements the control method for a drone swarm as described in any one of claims 1-6.