Communication-limited multi-unmanned aerial vehicle sustainable task planning method guided by edge system

Through the communication-constrained multi-UAV task planning method guided by edge systems, guided learning and heterogeneous reinforcement learning are used to optimize global sustainability indicators, solving the problem of task collaborative optimization of multi-UAV in dynamic environments and communication-constrained scenarios, and achieving balance and sustainability of global tasks.

CN120337974APending Publication Date: 2025-07-18TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324331.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the dynamic environment and communication restricted scenarios in the existing technology, it is difficult for multi-UAV task planning to achieve real-time collaborative optimization of global tasks, resulting in poor geographical fairness and weak global balance of task completion.

Method used

The edge system guidance method is adopted to optimize global sustainability indicators through guided learning and heterogeneous reinforcement learning drone decision model, combined with feedback mechanisms, and realize indirect information sharing and value guidance among drones.

Benefits of technology

The global spatio-temporal balance and adaptability of multi-UAV systems have been enhanced, and sustainable task planning under communication restricted conditions has been realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337974A_ABST
    Figure CN120337974A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an edge system-guided communication-limited multi-unmanned aerial vehicle sustainable task planning method, which comprises the steps of setting a working scene comprising a plurality of distributed heterogeneous unmanned aerial vehicles, a plurality of PoI and a plurality of edge nodes, enabling the distributed heterogeneous unmanned aerial vehicles to go to the PoI to execute data acquisition tasks, enabling the plurality of edge nodes to form an edge system, and enabling the distributed heterogeneous unmanned aerial vehicles to go to the PoI to execute data acquisition tasks; the method is used for indirect information sharing between unmanned aerial vehicles. And in the working scene, taking the global sustainability index as an optimization target of a communication-limited multi-unmanned aerial vehicle sustainable task planning problem, taking the unmanned aerial vehicle attribute as a constraint condition of the problem, and carrying out unmanned aerial vehicle decision making by using an unmanned aerial vehicle decision making model based on guide learning and heterogeneous reinforcement learning. And an edge system decision model based on a feedback mechanism is used for edge system decision making, so that the edge system carries out value guidance on the distributed heterogeneous unmanned aerial vehicle, and a global sustainability index is optimized. In this way, sustainable task planning of multiple unmanned aerial vehicles under the condition that communication is limited can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a sustainable task planning method for communication-constrained multi-unmanned aerial vehicles guided by an edge system. Background Art

[0002] In recent years, multi-unmanned aerial vehicle task planning, as a technology for efficiently solving information collection and task execution problems in complex scenarios, has received extensive attention. However, existing technologies face significant challenges in global task planning for scenarios such as dynamic environments and communication constraints. Due to limited computing power, unmanned aerial vehicles usually can only make autonomous decisions based on local perception information. Such short-sighted planning often ignores high-value tasks in remote areas, resulting in poor geographical fairness and weak global balance in task completion. In addition, communication limitations between unmanned aerial vehicles and between unmanned aerial vehicles and the edge system further exacerbate this problem, making it difficult to achieve real-time collaborative optimization of global tasks. A large number of existing studies assume that unmanned aerial vehicles can communicate with a central server in real time, or assume that unmanned aerial vehicles can obtain global environmental information, which is not applicable to real scenarios. In real scenarios, unmanned aerial vehicles often face problems such as communication constraints.

[0003] At the same time, the introduction of edge computing technology provides a new idea for multi-unmanned aerial vehicle collaborative planning. The edge system has stronger computing power and global information integration capabilities, and can perform global optimization planning by analyzing the task status information shared by unmanned aerial vehicles. However, traditional centralized task planning methods are inefficient in communication-constrained scenarios, and the planning results are difficult to be fed back to unmanned aerial vehicles in a timely manner, limiting their practicality. Therefore, how to achieve sustainable task planning for multi-unmanned aerial vehicles under communication constraints has become an urgent technical problem to be solved. Summary of the Invention

[0004] In a first aspect, an embodiment of the present invention provides a sustainable task planning method for communication-constrained multi-unmanned aerial vehicles guided by an edge system, the method including:

[0005] Set a working scenario, where the working scenario includes multiple distributed heterogeneous unmanned aerial vehicles, multiple data points to be collected (Points of Interest, PoIs), and multiple edge nodes. The distributed heterogeneous unmanned aerial vehicles need to plan routes to the PoIs to perform data collection tasks, and the multiple edge nodes form an edge system for indirect information sharing between the distributed heterogeneous unmanned aerial vehicles;

[0006] In this working scenario, the global sustainability indicator is used as the optimization objective of the communication-constrained multi-UAV sustainable mission planning problem, and the distributed heterogeneous UAV attributes are used as the constraints of the communication-constrained multi-UAV sustainable mission planning problem. Based on this, a UAV decision-making model based on guided learning and heterogeneous reinforcement learning is used to make decisions for UAVs, and an edge system decision-making model based on a feedback mechanism is used to make decisions for the edge system, so that the edge system can provide value guidance for distributed heterogeneous UAVs and optimize the global sustainability indicator.

[0007] In some realizable ways of the first aspect, the global sustainability indicator is composed of a data freshness indicator, a geographical fairness indicator, and a data volume indicator.

[0008] In some realizable ways of the first aspect, distributed heterogeneous UAVs form different levels of cognition of the environment through a task status information model, and combine their own task objectives and ability constraints to use their corresponding local evaluation functions to score the cost and benefit of the task, and then obtain a local evaluation vector.

[0009] In some realizable ways of the first aspect, starting from the task status information from its own perspective, the edge system constructs a global evaluation function for the PoI, and uses this to score the cost and benefit of the task from the global perspective of the edge system, and then obtains a global evaluation vector.

[0010] In some realizable ways of the first aspect, in the UAV decision-making model based on guided learning and heterogeneous reinforcement learning, distributed heterogeneous UAVs further introduce the guidance vector given by the edge system as value guidance on the basis of local task status information and local evaluation vectors to construct their own observation vectors, and then make decisions through the Actor network to output the actions of the current time slot; among them, the guidance vector given by the edge system is calculated by the edge system according to the global evaluation vector.

[0011] In some realizable ways of the first aspect, the training of the UAV decision-making model based on guided learning and heterogeneous reinforcement learning is implemented based on the HADDPG algorithm, in which a drift function and a mirror operation for heterogeneous agents are introduced.

[0012] In some realizable ways of the first aspect, in the edge system decision-making model based on a feedback mechanism, the edge system constructs a UAV task planning feedback model based on the historical decision data of the UAV, and calculates the guidance vector for distributed heterogeneous UAVs based on the UAV task planning feedback model and the global evaluation vector.

[0013] In the second aspect, an embodiment of the present invention provides a device for edge system-guided communication-constrained multi-UAV sustainable mission planning, and the device includes:

[0014] A setting module for setting a working scenario, where the working scenario includes multiple distributed heterogeneous drones, multiple points of interest (PoIs) to be collected with data, and multiple edge nodes. The distributed heterogeneous drones need to plan routes to the PoIs to perform data collection tasks, and the multiple edge nodes form an edge system for indirect information sharing among the distributed heterogeneous drones;

[0015] A planning module for, in this working scenario, taking the global sustainability metric as the optimization objective of the communication-constrained multi-drone sustainable mission planning problem, taking the attributes of the distributed heterogeneous drones as the constraints of the communication-constrained multi-drone sustainable mission planning problem, and based on this, using a drone decision-making model based on guided learning and heterogeneous reinforcement learning to make decisions for the drones, and using an edge system decision-making model based on a feedback mechanism to make decisions for the edge system, so that the edge system can provide value guidance for the distributed heterogeneous drones and optimize the global sustainability metric.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0017] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above.

[0018] In the embodiments of the present invention, for the communication-constrained multi-drone sustainable mission planning problem, a drone decision-making method based on guided learning and multi-agent heterogeneous reinforcement learning is proposed, and the edge system is used to guide the drones to dynamically understand different task values, so as to balance short-term local optimization and long-term global goals, enhance the global spatio-temporal balance and adaptability of the multi-drone system, and further achieve the communication-constrained multi-drone sustainable mission planning.

[0019] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the present invention and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0021] Figure 1Flow chart of a sustainable mission planning method for communication - limited multi - UAVs guided by an edge system provided by an embodiment of the present invention;

[0022] Figure 2 Schematic diagram of task status information synchronization between UAVs provided by an embodiment of the present invention;

[0023] Figure 3 Schematic diagram of UAV decision - making provided by an embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the HADDPG algorithm architecture provided by an embodiment of the present invention;

[0025] Figure 5 Structure diagram of a sustainable mission planning device for communication - limited multi - UAVs guided by an edge system provided by an embodiment of the present invention;

[0026] Figure 6 Structure diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] In addition, the term "and / or" in the present invention is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the front and rear associated objects.

[0029] To solve the technical problems in the background art, embodiments of the present invention provide a sustainable mission planning method, device, equipment, and storage medium for communication - limited multi - UAVs guided by an edge system. The following will, with reference to the accompanying drawings, specifically describe the sustainable mission planning method, device, equipment, and storage medium for communication - limited multi - UAVs guided by an edge system provided by embodiments of the present invention in detail.

[0030] Figure 1 Flow chart of a sustainable mission planning method for communication - limited multi - UAVs guided by an edge system provided by an embodiment of the present invention. As Figure 1 shown, the sustainable mission planning method 100 for communication - limited multi - UAVs may include:

[0031] S110. Set the working scenario, which includes multiple distributed heterogeneous drones, multiple points of interest (PoIs) for data collection, and multiple edge nodes. The distributed heterogeneous drones need to plan routes to the PoIs to perform data collection tasks. The multiple edge nodes form an edge system for indirect information sharing among the distributed heterogeneous drones.

[0032] S120. In this working scenario, use the global sustainability metric as the optimization objective for the communication-constrained multi-UAV sustainable mission planning problem, and use the attributes of the distributed heterogeneous drones as the constraints for the communication-constrained multi-UAV sustainable mission planning problem. Based on this, use a UAV decision-making model based on guided learning and heterogeneous reinforcement learning for UAV decision-making, and use an edge system decision-making model based on a feedback mechanism for edge system decision-making, so that the edge system can provide value guidance to the distributed heterogeneous drones and optimize the global sustainability metric.

[0033] For further understanding, the above steps are described in detail below with specific embodiments:

[0034] (1) Working scenario

[0035] Consider U distributed heterogeneous drones (denoted by the numbers u = 1, 2, …, U) performing data collection on P PoIs (denoted by the numbers p = 1, 2, …, P) evenly distributed in a rectangular working area. There is no direct communication between the drones, and the drones cannot communicate with the central server in real time. There are N edge nodes (denoted by the numbers n = 1, 2, …, N) set in the working area. These edge nodes are connected through a backbone network to form an edge system, denoted by N. When the Euclidean distance between a drone and any edge node n is within the threshold range ρ U~N , the drone can access the edge node to synchronize the task status information. Here, consider the above-mentioned edge system-assisted UAV data collection scenario. The data collection process of the UAV is regarded as a number of discrete time slots, denoted by t i (i = 1, 2, …, T).

[0036] (2) UAV-edge system evaluation function model

[0037] The distributed heterogeneous drones need to perform long-term data collection on the data in different PoIs. The distributed heterogeneous drones have different capabilities (here only consider two capabilities of the UAV: the data collection radius ρ u and the data collection rate λ u ). During the data collection process, the UAV needs to make decisions starting from the optimization of the global sustainability metric, where the global sustainability metric is composed of indicators such as data freshness, geographical fairness, and data volume. The decision variables are The unmanned aerial vehicle (UAV) forms different levels of cognition of the environment through the task status information model, and combines its own task objectives and ability constraints. Using the local evaluation function (represented by the following formula ), it scores the cost and benefit of the task, realizes the spatio-temporal measurement and understanding of the task value, and then forms an action locally (i.e., the decision variable ) to execute various data collection tasks in the environment.

[0038]

[0039] Among them, K s and L s are the value evaluation functions for the spatio-temporal attributes of the task and the dynamic factors of environmental interaction respectively; is the s-th element in the task status information vector of PoI p from the perspective of UAV u (the task status information vector is embedded with attributes such as data freshness, geographical fairness, and data volume); t i is the time slot; Q is the time-varying benefit value evaluation function. Note that the output result here is a scalar. For all P PoIs, constitutes a vector, denoted as

[0040] Exemplarily, the spatio-temporal attributes of the task include the task geographical location and the task data growth rate; the dynamic factors of environmental interaction include data freshness, historical access times, and the most recent access time.

[0041] In this embodiment, the edge system has the following two functions:

[0042] First, synchronize the task status information between UAVs. This process is as shown in Figure 2 , which shows the information exchange process when two UAVs (A, B) visit the edge system successively. In Figure 2 , are the task status information of the two UAVs at times t1 and t2. Through interaction with the edge system and by means of the edge information fusion method of the edge system, UAVs A and B synchronize their local task status information with the edge at times t1 and t2 respectively. In particular, UAV B is fed back the task status information uploaded by UAV A through the edge system at time t2.

[0043] Second, guide the UAV to maximize the optimization objective. To guide the UAV, the edge system starts from the task status information from its own perspective, constructs a global evaluation function (similar to the local evaluation function of the UAV) for each PoI, and scores the cost and benefit of the task from the global perspective of the edge system.

[0044]

[0045] Note that the output result here is also a scalar. Similarly, the global evaluation functions of the edge system for each PoI can be concatenated into a global evaluation vector

[0046] (3) UAV Decision-making Model Based on Guided Learning and Heterogeneous Reinforcement Learning

[0047] The UAV makes decisions using a UAV decision-making model based on guided learning and heterogeneous reinforcement learning, and the process is as Figure 3 shown.

[0048] Specifically, based on the local task status information and the local evaluation vector of the UAV, the guidance vector given by the edge system is further introduced as "value guidance" to construct its own observation vector In the formula represents the observation space of the UAV. That is to say, is composed of three parts concatenated together. Then, decisions are made through the Actor network from to output the action

[0049] of the current time slot to calculate the guidance vector Specifically, the edge system regards as the observation In the formula represents the observation space of the edge system. The edge system makes decisions through the feedback model to calculate its guidance vector and sends it to the UAV u.

[0050] The training process of the UAV decision-making model is implemented based on the HADDPG algorithm. Similar to the MADDPG algorithm, it equips heterogeneous agents with Actor networks and Critic networks, and introduces a drift function and mirror operation for heterogeneous agents, enabling each heterogeneous agent to update its own local policy on the basis of more fully considering the strategies of other agents. The framework of this algorithm is as Figure 4 shown.

[0051] (4) Edge System Decision-making Model Based on Feedback Mechanism

[0052] The edge system constructs a UAV planning feedback model based on the historical decision data of the UAV, and based on this model and the global value understanding Calculate the guidance vector for the UAV u That is

[0053] To enable the edge system to guide multiple distributed heterogeneous UAVs for global mission planning, the edge system utilizes the historical decision data of the UAVs Construct the mission planning feedback model for each UAV (action network), and collect several groups of the planning action feedback of the UAV after guidance to form a set To reduce the number of networks, the UAV number is used as the network input here, and only one network is used to calculate the guidance vectors of different UAVs and achieve personalized value guidance.

[0054] Next, the edge system evaluates according to the Critic network performance, and trains the network parameters in the way of minimizing the loss function

[0055]

[0056] In the formula is the edge system reward function, which is weighted and defined by the global sustainability index, that is:

[0057]

[0058] In the formula:

[0059]

[0060] Thus, the edge system continuously collects the decision results of heterogeneous UAVs under different feedback models. Through the above process, considering the global sustainability index, by continuously updating to optimize the guidance vector Thereby improving the long-term sustainability of mission planning.

[0061] (5) Experimental demonstration

[0062] Set α = 0.1, β = 1, γ = 5. The ablation experiment and comparison results of the present invention are listed in Table 1 below:

[0063] Table 1 Sustainability indicators under different numbers of UAVs

[0064] Number of drones The method of the present invention (Method 1) Method 2 Method 3 1 52.38 40.38 27.38 2 85.72 68.91 41.17 3 107.63 83.45 52.20 4 125.31 105.32 72.63

[0065] Method 2: Modify HADDPG to MADDPG (without introducing the heterogeneous reinforcement learning mechanism).

[0066] Method 3: Do not consider the guidance of the edge system (that is, do not establish a feedback mechanism and do not use guidance learning).

[0067] Through experiments, it can be seen that in different scenarios (different numbers of heterogeneous UAVs), the heterogeneous reinforcement learning mechanism and the guided learning mechanism introduced by the method of the present invention can effectively improve the global sustainability index.

[0068] In summary, according to the embodiments of the present invention, at least the following technical effects are achieved:

[0069] 1. Expand the uses of the edge system. In addition to sharing task status information, it can also provide value guidance for UAVs, balance short-term task efficiency and long-term task goals, and enhance the global spatio-temporal balance of task planning.

[0070] 2. Propose a decision-making method for guided learning and heterogeneous multi-agent reinforcement learning. In view of the communication-limited conditions, establish a two-way guidance and feedback mechanism between the edge system and UAVs to improve the long-term sustainability of task planning.

[0071] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0072] The above is the introduction of the method embodiments. The following further illustrates the solution of the present invention through device embodiments.

[0073] Figure 5 The following is a structural diagram of a communication-limited multi-UAV sustainable task planning device guided by an edge system provided for an embodiment of the present invention. As Figure 5 shown, the communication-limited multi-UAV sustainable task planning device 500 may include:

[0074] A setting module 510, configured to set a working scenario, where the working scenario includes a plurality of distributed heterogeneous UAVs, a plurality of data collection points PoI to be collected, and a plurality of edge nodes. The distributed heterogeneous UAVs need to plan routes to the PoI to perform data collection tasks, and the plurality of edge nodes form an edge system for indirect information sharing between the distributed heterogeneous UAVs.

[0075] The planning module 520 is configured to use the global sustainability metric as the optimization objective of the communication-constrained multi-UAV sustainable mission planning problem and the distributed heterogeneous UAV attributes as the constraints of the communication-constrained multi-UAV sustainable mission planning problem in this working scenario. Based on this, a UAV decision-making model based on guided learning and heterogeneous reinforcement learning is used to make UAV decisions, and an edge system decision-making model based on a feedback mechanism is used to make edge system decisions, so that the edge system can provide value guidance for the distributed heterogeneous UAVs and optimize the global sustainability metric.

[0076] It can be understood that Figure 5 each module / unit in the communication-constrained multi-UAV sustainable mission planning device 500 shown has the function of implementing Figure 1 each step in the communication-constrained multi-UAV sustainable mission planning method 100 shown and can achieve its corresponding technical effects. For the sake of brevity, they will not be described in detail here.

[0077] Figure 6 It is a structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. The electronic device 600 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 600 can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.

[0078] As Figure 6 shown, the electronic device 600 may include a computing unit 601, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0079] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0080] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 can be implemented as a computer program product, including a computer program, which is tangibly contained in a computer-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 100 described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute method 100 in any other suitable manner (e.g., by means of firmware).

[0081] The various embodiments described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0082] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when the program codes are executed by the processor or the controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0083] In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer-readable storage media would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0084] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute Method 100 and achieve the corresponding technical effects achieved by the method of the embodiments of the present invention. For the sake of brief description, details are not repeated herein.

[0085] In addition, the present invention also provides a computer program product, which includes a computer program that implements Method 100 when executed by a processor.

[0086] It should be understood that various forms of the flow shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. The present invention is not limited herein.

[0087] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A communication-constrained multi-UAV sustainable mission planning method guided by the limbic system, characterized in that The method includes: Setting up a working scenario, which includes multiple distributed heterogeneous drones, multiple Points of Interest (PoIs) for data collection, and multiple edge nodes. The distributed heterogeneous drones need to plan routes to the PoIs to perform data collection tasks. The multiple edge nodes form an edge system for indirect information sharing among the distributed heterogeneous drones; In this working scenario, taking the global sustainability metric as the optimization objective of the communication-constrained multi-drone sustainable mission planning problem, and the attributes of the distributed heterogeneous drones as the constraints of the communication-constrained multi-drone sustainable mission planning problem. Based on this, a drone decision-making model based on guided learning and heterogeneous reinforcement learning is used for drone decision-making, and an edge system decision-making model based on a feedback mechanism is used for edge system decision-making, so that the edge system can provide value guidance to the distributed heterogeneous drones and optimize the global sustainability metric.

2. The method according to claim 1, wherein The global sustainability metric is composed of a data freshness metric, a geographical fairness metric, and a data volume metric in combination.

3. The method according to claim 2, wherein The distributed heterogeneous drones form different levels of cognition of the environment through a task status information model, and combine their own task objectives and ability constraints, and use their corresponding local evaluation functions to score the cost and benefit of the task, and then obtain a local evaluation vector.

4. The method according to claim 3, characterized in that Starting from the task status information from its own perspective, the edge system constructs a global evaluation function for the PoIs, and uses this to score the cost and benefit of the task from the global perspective of the edge system, and then obtains a global evaluation vector.

5. The method according to claim 4, wherein In the drone decision-making model based on guided learning and heterogeneous reinforcement learning, based on the local task status information and the local evaluation vector, the distributed heterogeneous drones further introduce the guidance vector given by the edge system as value guidance to construct their own observation vectors, and then make decisions through the Actor network to output the actions in the current time slot; among them, the guidance vector given by the edge system is calculated by the edge system based on the global evaluation vector.

6. The method according to claim 5, characterized in that, The training of the drone decision-making model based on guided learning and heterogeneous reinforcement learning is implemented based on the HADDPG algorithm, in which a drift function and a mirror operation for heterogeneous agents are introduced.

7. The method according to claim 6, wherein In the edge system decision-making model based on a feedback mechanism, the edge system constructs a drone task planning feedback model based on the historical decision data of the drones, and calculates the guidance vector for the distributed heterogeneous drones based on the drone task planning feedback model and the global evaluation vector.

8. An edge system-guided communication-constrained multi-UAV sustainable mission planning device, characterized in that, The device includes: A setting module for setting up a working scenario, which includes multiple distributed heterogeneous drones, multiple Points of Interest (PoIs) for data collection, and multiple edge nodes. The distributed heterogeneous drones need to plan routes to the PoIs to perform data collection tasks. The multiple edge nodes form an edge system for indirect information sharing among the distributed heterogeneous drones; A planning module, which, in this working scenario, uses the global sustainability indicator as the optimization objective of the communication-constrained multi-UAV sustainable mission planning problem, and the distributed heterogeneous UAV attributes as the constraints of the communication-constrained multi-UAV sustainable mission planning problem. Based on this, a UAV decision-making model based on guided learning and heterogeneous reinforcement learning is used for UAV decision-making, and an edge system decision-making model based on a feedback mechanism is used for edge system decision-making, so that the edge system can conduct value guidance for the distributed heterogeneous UAVs and optimize the global sustainability indicator.

9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-7.