Multi-unmanned aerial vehicle dynamic task allocation method based on edge system and data flow age
By introducing the concept of edge systems and data flow age in the multi-drone collaborative data acquisition system, and using multi-agent layered deep reinforcement learning for intelligent assignment of tasks, the problem of data flow information freshness in dynamic environments is solved, and efficient data acquisition and feedback are achieved.
Patent Information
- Application Number
- CN202510040851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The prior art is difficult to effectively deal with the freshness of data flow information in a dynamic environment for multi-UAV collaborative data acquisition, especially in the case of dynamic changes in data volume and scene dynamics, making it difficult to achieve timely data acquisition and feedback.
A multi-drone dynamic task intelligent allocation method based on edge system and data flow age is proposed. By setting a partially observable Markov decision-making process model for multiple agents, and using multi-agent hierarchical deep reinforcement learning to make macro option decisions and micro action decisions, optimizing the data flow age mean of the drone under energy constraints.
It improves the dynamic response performance of the drone in dynamic scenarios, ensures the freshness and timeliness of the data stream, reduces the age average and energy consumption of the data stream, and performs superiorly compared with the baseline algorithm.
Smart Images

Figure CN120013136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and in particular to a method for intelligently allocating dynamic tasks of multiple unmanned aerial vehicles based on edge systems and data stream age. Background Art
[0002] Since the concept of mobile crowd sensing was proposed, task allocation and task planning in multi-agent collaborative data collection have been regarded as important issues in mobile crowd sensing. Currently, a lot of work has been done to study this issue, and many frameworks and algorithms for task allocation in mobile crowd sensing have been proposed.
[0003] Some studies use dynamic programming and heuristic optimization algorithms to find the optimal or suboptimal solution to the task allocation problem under complex constraints. However, these methods have weak transferability to different environments and cannot handle dynamic environments. In the problem of crowd sensing, most existing studies assume that the total amount of data to be collected in the Point of Interest (PoI) remains unchanged, and the scene dynamics is weak. However, in real mobile crowd sensing tasks, the total amount of data in each PoI does not remain constant. For example, when a camera continues to take photos or record videos, the amount of data stored internally will increase dynamically over time.
[0004] In a dynamic environment, the timeliness of data collection is more important. Generally speaking, it is hoped that the data in the PoI will be collected as early as possible, rather than being stored inside the PoI for a long time and unable to be sensed by external drones, so that timely feedback and decisions cannot be made for emergencies (such as abnormal temperature, suspicious persons, etc.). Among them, the age of information (AoI) is an important indicator for measuring the freshness of information. This indicator is defined as the stagnation time of data at a certain place. A large number of existing studies use information age as a target or constraint to optimize the dynamic performance of mobile crowd intelligence perception systems, such as transmission delay. However, information age can only be used to measure the information freshness of a single data packet, and cannot handle the information freshness of a data stream consisting of multiple data packets with different generation times.
[0005] In recent years, methods based on deep reinforcement learning have attracted widespread attention from researchers because they can solve joint task allocation and path planning problems in more complex scenarios. In general, judging from the high-level research results in recent years, a large amount of work on task allocation and path planning problems in the field of mobile crowd sensing data collection, including the most advanced research work, is almost all based on deep reinforcement learning or multi-agent deep reinforcement learning. Such methods currently achieve the best performance in the data collection decision-making process of multiple drones in mobile crowd sensing compared to other methods. However, the decision variables of these methods are often the movement direction and distance of the drone in each time slot, etc., which makes it difficult to explain the reasons why the drone makes decisions from both the macro and micro levels, so the flexibility is poor. Summary of the invention
[0006] In a first aspect, an embodiment of the present invention provides a method for intelligently allocating dynamic tasks of multiple UAVs based on edge systems and data stream age, the method comprising:
[0007] Set up a working scenario, which includes multiple drones, multiple PoIs, and multiple edge nodes. Multiple edge nodes form an edge system for indirect information sharing between drones.
[0008] In this working scenario, the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information are regarded as options, the micro-behavior of the drone in the time slot is regarded as an action, and the optimization goal of the drone's mobile crowd perception task is determined as minimizing the mean age of data stream (AoDS) of all PoIs in the data collection process under energy constraints. Each drone is regarded as an intelligent agent, and a multi-agent partially observable Markov decision process model based on options is established, so that macro-option decisions and micro-action decisions are made through multi-agent hierarchical deep reinforcement learning.
[0009] In some implementations of the first aspect, the data stream age is defined as:
[0010]
[0011] Among them, A i (p) indicates PoI p at t i The age of the data stream at time, Refers to t i The set of data packets contained in PoI p at the time, δ m is the amount of data in packet m, ι m The time when the data packet is generated.
[0012] In some implementations of the first aspect, the process of adding new PoI data is regarded as a Poisson process, specifically:
[0013] The new data in PoI p are regarded as random events, and these events are independent of each other and have equal probability of occurrence. i The number of new data in the time slot follows the Poisson distribution, and the time slot t i The number of times new data appears in PoI p is K i (p), the probability of each event is λ(p), then K i (p)~Poisson(λ(p));
[0014] Assuming that for the same PoI p, the amount of new data added each time is a fixed value D(p), then time slot t i The amount of new data added is δ i (p) = K i (p)D(p), these data are packaged into data packets m i , and added to the data stream of PoI p, the data stream will be placed in the drop box of the PoI, waiting for collection by drones that visit in the future;
[0015] The UAV obtains the data flow information of PoI p, regards the data volume of the data packets generated in these Poisson processes as samples, and estimates the Poisson distribution parameter λ(p) and the proportion D(p) of PoI p.
[0016] In some possible implementations of the first aspect, the drone collecting data on the PoI includes:
[0017] In the same time slot, the drone can transmit data with at most C PoIs. i In the process, the drone first visits its perception radius ρ s The drop box of all PoIs within the range is used to obtain the data flow age of all PoIs. Here, the perception radius ρ s All PoIs in the like Then the drone u will be Select the C PoIs with the largest data stream age for data collection; if You will choose Collect data from all PoIs within the time slot t i Select all PoIs collected from data as a set
[0018] In time slot t i in Approximate data transmission rate between UAV u at position and PoI p at position (x(p), y(p)) It is defined by the Shannon capacity taking into account the path loss exponent as shown below:
[0019]
[0020] in, represents the approximate data transmission rate, B represents the total effective bandwidth, q0 represents the average transmission power of the UAV, and z 2 is the white noise power of Gaussian distribution, g0 represents the channel gain at the reference distance, α represents the path loss index, and dist(u,p,i) represents the distance between UAV u and PoI p in time slot t i At the discrete grid scale l grid The Euclidean distance under
[0021] For each data to be collected Drones will be Within a certain time, each data packet in the data stream is collected from its throwing box in a first-in-first-out order, and the data volume of each data packet is recorded. After that, the data stream age of the PoI will be automatically updated.
[0022] In some implementations of the first aspect, the state space of the partially observable Markov decision process model includes the location of the PoI, data generation parameters, the amount of remaining data of the PoI in the current time slot, the age of the data stream in the current time slot, and the location and energy of all drones in the current time slot;
[0023] The observation space of the partially observable Markov decision process model includes the location of the PoI, the estimated value of the data generation parameter, the estimated value of the remaining data volume of the PoI in the current time slot, the estimated value of the age of the data stream in the current time slot, and the location and energy of the drone in the current time slot;
[0024] The option space of the partially observable Markov decision process model represents the macroscopic active behavior of the agent, and the action space represents the microscopic behavior of the agent, specifically the movement vector of the agent in the time slot;
[0025] The reward function of the partially observable Markov decision process model includes the reward for the agent to collect data, the penalty for energy consumption, the reward given by the edge system, and the penalty for the agent to move out of the working area.
[0026] In some implementations of the first aspect, the training of multi-agent hierarchical deep reinforcement learning includes:
[0027] Before training begins, each agent is equipped with an option-value network for macro option decision making And configure a pair of executor-critic network inside each option Each network also has a corresponding target network; in addition, each intelligent body is also equipped with a corresponding option experience replay pool and option experience replay pool
[0028] The training consists of EP rounds. In each round, the environment is set to the initial state s0, and each agent obtains the initial observation of the environment The option termination identifier f of each agent u is set to True;
[0029] Each round has M time slots. In each time slot, if the option termination identifier f u is True, the agent chooses the macro option according to the following formula:
[0030]
[0031] Afterwards, f u Set to False, for f u If it is False, the agent maintains the original option and does not need to choose the option. Next, each agent generates actions according to the strategy function inside the option. And add Gaussian noise to the action. Next, each agent interacts with the environment according to the following steps:
[0032] The drone is following the noise to move;
[0033] The drone collects data from the PoI according to the data collection node model;
[0034] If the drone is within the communication range of any edge node, it will exchange, synchronize, and update mission status information with the edge system;
[0035] The drone estimates the parameters of the PoI according to the parameter estimation model;
[0036] PoI generates new data packets according to the data flow model and updates the data flow age;
[0037] Then, the environment state shifts to s i+1 , each agent obtains a local observation of the new state Next, the agent decides whether to stop the current option. If the option is stopped, it sets f u ←True, the stopping conditions for different options are as follows:
[0038] ω=1, that is, when the agent's option is to collect data, the stopping condition is i―i ω ≥M c , that is, the agent continuously collects M cAfter a time slot has been filled, the options need to be re-determined;
[0039] ω=2, that is, when the agent's option is to visit the edge node, the stopping condition is that the agent successfully visits the edge node, that is When you need to re-determine the options;
[0040] Afterwards, each agent will experience Stored in the experience replay pool of each corresponding option In the above example, the value network parameters in the option are updated by the gradient descent method, the policy network parameters in the option are updated by the gradient ascent method, and then the target network in the option is updated by the soft update factor τ;
[0041] If agent u is in the current time slot t i If an option is terminated, the environment will give the agent feedback on the reward corresponding to the option. After that, the agent will experience Deposit Option Experience Replay Next, the agent Sampling a batch of experience And use the loss function shown in the following formula to update the network parameters according to the gradient descent method
[0042] Finally, according to the soft update factor τ Ω Update target options - value target network parameters.
[0043] In a second aspect, an embodiment of the present invention provides a multi-UAV dynamic task intelligent allocation device based on edge system and data stream age, the device comprising:
[0044] A setting module is used to set a working scenario. The working scenario includes multiple drones, multiple PoIs of data points to be collected, and multiple edge nodes. Multiple edge nodes form an edge system for indirect information sharing between drones.
[0045] The allocation module is used to regard the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information as options in this working scenario, regard the micro-behavior of the drone in the time slot as an action, and determine the optimization goal of the drone's mobile crowd perception task as minimizing the mean age of the data stream of all PoIs in the data collection process under energy constraints. Each drone is regarded as an intelligent agent, and a multi-agent partially observable Markov decision process model based on options is established, so as to make macro-option decisions and micro-action decisions through multi-agent hierarchical deep reinforcement learning.
[0046] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.
[0047] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described above.
[0048] According to the embodiments of the present invention, at least the following technical effects are achieved:
[0049] First, the data stream age in dynamic scenarios is established and defined as a data freshness indicator to evaluate the dynamic response performance of UAV collaborative data collection. The data stream age considers both the amount of data in the data packet and the residence time of the data packet. Compared with the traditional information age, it can better reflect the data freshness of a data stream composed of multiple data packets;
[0050] Second, based on the edge system-assisted indirect task state information sharing architecture in the existing technology, a parameter estimation based on the Poisson distribution random process and a sample exchange mechanism based on the edge system are further proposed, which improves the accuracy of the UAV mobile crowd perception system in estimating the dynamic parameters of PoI;
[0051] Third, to solve the problem of drone data collection in dynamic scenarios, an algorithm based on multi-agent hierarchical deep reinforcement learning (DRL-DCEA) is proposed. The algorithm regards the macro and micro behaviors of drones as options and actions respectively. The data collection and access to edge nodes of drones under timeliness requirements are regarded as two macro options. By introducing option decision-making, the dynamic responsiveness of drones in distributed autonomous decision-making is improved;
[0052] Fourth, simulation experiments verified that in a typical dynamic mobile crowd-sensing scenario, the proposed DRL-DCEA algorithm is superior to baseline algorithms such as DRL-ASPT, Edics, MADDPG, greedy strategy, and random strategy in terms of data stream mean age, energy consumption, and other evaluation indicators. For example, compared with the DRL-ASPT algorithm, the DRL-DCEA algorithm reduces the data stream mean age by 29.74% and the energy consumption by 24.33%.
[0053] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present invention. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0055] Figure 1 A flowchart of a method for intelligently allocating dynamic tasks of multiple UAVs based on edge systems and data stream age provided by an embodiment of the present invention;
[0056] Figure 2 It is a schematic diagram of the change of data stream age over time;
[0057] Figure 3 Schematic diagram of Poisson process parameter estimation based on task state information;
[0058] Figure 4 Box plot for PoI parameter estimates;
[0059] Figure 5 Schematic diagram of the DRL-DCEA algorithm architecture;
[0060] Figure 6 It is a schematic diagram of the DRL-DCEA algorithm flow;
[0061] Figure 7 A schematic diagram showing the effect of different numbers of drones on performance;
[0062] Figure 8 This is a schematic diagram showing the impact of different perception radii on performance;
[0063] Fig. 9 A structural diagram of a multi-UAV dynamic task intelligent allocation device based on edge system and data stream age provided by an embodiment of the present invention;
[0064] Fig.10 The figure is a structural diagram of an exemplary electronic device capable of implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] In addition, the term "and / or" in the present invention is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects before and after are in an "or" relationship.
[0067] In order to solve the technical problems in the background technology, the embodiments of the present invention provide a method, device, equipment and storage medium for intelligent allocation of dynamic tasks of multiple unmanned aerial vehicles based on edge system and data stream age. The following is a detailed description of a method, device, equipment and storage medium for intelligent allocation of dynamic tasks of multiple unmanned aerial vehicles based on edge system and data stream age provided by the embodiments of the present invention in conjunction with the accompanying drawings through specific embodiments.
[0068] Figure 1 A flowchart of a method for intelligently allocating dynamic tasks of multiple UAVs based on edge systems and data stream age is provided in an embodiment of the present invention, such as Figure 1 As shown, the multi-UAV dynamic task intelligent allocation method 100 may include the following steps:
[0069] S110, setting a working scene, where the working scene includes multiple drones, multiple PoIs, and multiple edge nodes. The multiple edge nodes form an edge system for indirect information sharing between drones.
[0070] S120, in this working scenario, the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information are regarded as options, the micro-behavior of the drone in the time slot is regarded as an action, and the optimization goal of the drone's mobile crowd perception task is determined as minimizing the mean age of the data stream of all PoIs in the data collection process under energy constraints. Each drone is regarded as an intelligent agent, and a multi-agent partially observable Markov decision process model based on options is established, so as to make macro-option decisions and micro-action decisions through multi-agent hierarchical deep reinforcement learning.
[0071] For further understanding, the above steps are described below in conjunction with specific embodiments:
[0072] (1) Set up the work scene
[0073] Consider a rectangular working area with P PoIs evenly distributed in the rectangular working area according to their geographical locations. U drones are used to collect data from these PoIs, and the drones do not have direct communication capabilities. N edge nodes are set to form an edge system for indirect information sharing between drones.
[0074] (2) Data flow model of PoI
[0075] Based on AoI, the Age of Data Stream (AoDS) is proposed to measure the information freshness of multiple data packets generated at different times in a PoI. This indicator takes into account the data volume of each data packet and the time it is stored in the PoI, and is defined as:
[0076]
[0077] Among them, A i (p) indicates PoI p at t i The age of the data stream at time, Refers to t i The set of data packets contained in PoI p at the time, δ m is the amount of data in packet m, ι m The time when the data packet is generated. Figure 2 Describes the change of data stream age over time, where m a 、m a 、m c Respectively in t a ,t b ,t c The time generated, t d Time m a and m b Collected, m e In t c The moment is generated.
[0078] In the case of time slot discretization, the data stream age can also be calculated in an incremental manner, namely:
[0079]
[0080] in, is the set of data packets collected by the drone. In the case of discrete time slots, for the convenience of expression and calculation, the time when each PoI generates a new data packet is the moment when the time slot ends.
[0081] (3) Parameter estimation model of UAV for PoI
[0082] The process of adding new data to PoI is considered as a Poisson process. Specifically, the new data in PoI p (such as taking photos) is considered as a random event, and these events are independent of each other and have equal probability of occurrence. Then the time slot t i The number of new data in the time slot follows the Poisson distribution, and the time slot t i The number of times new data appears in PoI p is K i (p), the probability of each event is λ(p), then K i(p)~Poisson(λ(p)).
[0083] Assuming that for the same PoI p, the amount of new data added each time is a fixed value D(p) (for example, each photo taken occupies the same storage space), then the time slot t i The amount of new data added is δ i (p) = K i (p)D(p), these data are packaged into data packets m i , and added to the data stream of PoI p, the data stream will be placed in the drop box of the PoI, waiting to be collected by drones that visit in the future.
[0084] Drones from Figure 3 The data flow information of PoI p is obtained from the fused TSI shown in the figure. The data volume of the data packets generated in these Poisson processes is regarded as samples, and the Poisson distribution parameter λ(p) and the proportion D(p) of p are estimated. Assume that the data volume sample of the drone for PoI p is. According to the first-order moment and second-order moment of the Poisson process, we have:
[0085]
[0086] Then the estimated data growth ratio of the PoI can be calculated Poisson distribution parameter estimator and an estimate of the data growth rate Figure 4 The parameter estimation results of PoI in an embodiment scenario under specific parameters are shown. It can be considered that the more samples the drone observes of the data volume of a PoI data packet, the more accurate the parameter estimation value of the PoI is.
[0087] (4) UAV data collection model
[0088] Assuming that each drone can use multiple antennas and frequency division multiplexing mechanism to transmit data to multiple PoIs at the same time, in order to ensure the transmission rate between drones and PoIs, it is assumed that in the same time slot, the drone can transmit data to at most C PoIs. i In the process, the drone first visits its perception radius ρ s The drop boxes of all PoIs in the collection are used to obtain all PoIs (recorded as a set )’s data stream age. If >C, then the drone u will Select the C PoIs with the largest data stream age for data collection; if UAV u will choose Collect data from all PoIs within the time slot t iSelect all PoIs collected from data as a set According to the above definition, there is Since the amount of data required for the drone to access the PoI drop box is small, for simplicity, the transmission delay problem is not considered here.
[0089] In time slot t i in Approximate data transmission rate between UAV u at position and PoI p at position (x(p), y(p)) It is defined by the Shannon capacity taking into account the path loss exponent as shown below:
[0090]
[0091] Where B represents the total effective bandwidth (when u transmits data to multiple PoIs simultaneously, it is assumed that the total bandwidth is evenly distributed), q0 represents the average transmission power of the UAV, and z 2 is the white noise power of Gaussian distribution, g0 represents the channel gain at a reference distance (e.g., a unit distance, 1 m in this paper), α represents the path loss exponent, and dist(u,p,i) represents the distance between UAV u and PoI p in time slot t i At the discrete grid scale l grid The Euclidean distance under (corresponding to the number of grids), that is:
[0092]
[0093] For each PoI to collect data Drones will be Each data packet in the data stream is collected from its throwing box in the first-in-first-out (FIFO) order within the time, and the data volume of each data packet is recorded, and then the data stream age of the PoI will be automatically updated.
[0094] (5) Multi-agent hierarchical deep reinforcement learning
[0095] In dynamic scenes, the timeliness of drone data collection becomes more important. In order to make the drone mission planning more flexible under the timeliness requirement, two macro-active behaviors of drones are considered, including:
[0096] The drone goes to the data point to be collected to perform the task;
[0097] The drone goes to the edge node to share mission status information.
[0098] Therefore, it is necessary to design the action selection method of the drone at the macro and micro levels. The drone regards the above two high-level macro behaviors as action options, reflecting its overall goal within a period of time (the next several time slots); the micro behavior represents the specific action of the drone in each time slot, including the selection of detailed behaviors such as moving direction, distance, speed, etc.
[0099] The execution process of the UAV's mobile crowd sensing task is divided into M time slots, and the time span of each time slot is Δt. The optimization goal of the mobile crowd sensing task is to minimize the mean age of the data stream of all PoIs during the collection process under energy constraints, that is:
[0100]
[0101] Among them, A i (p) represents the data stream age of PoI. i (p) is related to the data packet generation time and data volume contained in the PoI data stream. Each data packet is generated by the aforementioned process and transferred to the drone as the drone collects data. Since this problem introduces time dynamics, its complexity is no less than that of the multi-traveling salesman problem, so the above optimization problem is very difficult to solve. Therefore, it is necessary to find a drone path planning and task allocation strategy with better performance to deal with the above time dynamics. Since the multi-agent reinforcement learning method shows performance superiority in sequential decision problems, and considering the hierarchical decomposition ability and high interpretability of hierarchical reinforcement learning for complex behavioral decision problems, the present invention proposes to use multi-agent hierarchical reinforcement learning to solve the above complex optimization problem.
[0102] For the above-mentioned sequential decision-making problem, each UAV is regarded as an intelligent agent, and a multi-agent option-based partially observable Markov decision process (game) model is established. The meaning and composition of each element will be explained below.
[0103] 1. The state space S of the environment consists of the following parts:
[0104] The location of the PoI, data generation parameters, the amount of remaining data of the PoI in the current time slot, the age of the data stream in the current time slot, and the location and energy of all drones in the current time slot.
[0105] 2. The observation space O of the drone is the observable part of the environment, which consists of the following parts:
[0106] The location of the PoI, the estimated values of the data generation parameters, the estimated value of the remaining data volume of the PoI in the current time slot, the estimated value of the age of the data stream in the current time slot, the location and energy of the drone in the current time slot.
[0107] 3. The option space Ω and action space A are defined as follows:
[0108] The option space of the agent is Ω = {1, 2}, which represents the macro-level behavior of the agent. Option ω = 1 means that the agent goes to the PoI to collect data, and option ω = 2 means that the agent goes to the edge system to exchange task status information.
[0109] The action space of the agent is the moving vector of the intelligence in a time slot, that is and Conditional constraints. The length of the action vector of a single agent is 2.
[0110] 4. The reward function is as follows:
[0111] At the end of each time slot, the environment gives each agent a reward:
[0112]
[0113] The four items in the above formula are the reward for the agent to collect data, the penalty for energy consumption, the reward given by the edge system, and the penalty for the agent moving out of the working area (i.e., out-of-bounds behavior).
[0114] If option ω is in time slot t i At the end, the environment will give each agent an option reward to learn the parameters of the option-value network. The option reward is set as follows:
[0115]
[0116] Among them, i ω Indicates the time slot when option ω starts, i―i ω Indicates the number of time slots that this option lasts. It is defined as the average value of the agent’s data collection reward during the option duration, defined as:
[0117]
[0118] Similarly, It is defined as the average of the marginal rewards over the option duration, defined as:
[0119]
[0120] also, They represent the proportional coefficients of data collection reward and edge reward in option reward respectively.
[0121] The present invention proposes a "DRL-DCEA" algorithm, which is an edge-assisted mobile crowd-sensing dynamic data collection algorithm based on multi-agent hierarchical deep reinforcement learning, where DCEA is the abbreviation of Data Collection or Edge Accessing, which represents the two macro actions of "collecting data or accessing edge nodes". Figure 5 is the architecture of the DRL-DCEA algorithm. In this architecture, each option of each agent u adopts an actor-critic network architecture, where Used to estimate the Q value of micro-action, Used to generate micro-actions, where ω∈{1,2} is used to represent two options; each agent also has an additional Used to estimate the value of options and for option decision making.
[0122] The DRL-DCEA algorithm process is as follows Figure 6 As shown, the specific steps include:
[0123] Consider each drone as an intelligent agent. Before training begins, equip each agent with an option-value network for macro-option decision making. And configure a pair of executor-critic network inside each option Each network also has a corresponding target network. In addition, each intelligent body is also equipped with a corresponding option experience replay pool. and option experience replay pool
[0124] The training of the algorithm consists of EP rounds. In each round, the environment is set to the initial state s0, and each agent obtains the initial observation of the environment The option termination identifier f of each agent u (Boolean variable) is set to True.
[0125] Each round has M time slots. In each time slot, if the option termination identifier f u is True, the agent needs to select the macro option according to the following formula:
[0126]
[0127] Afterwards, f u Set to False. u If it is False, the agent maintains the original option and does not need to choose the option. Next, each agent generates actions according to the strategy function inside the option. And add Gaussian noise to the action. Next, each agent interacts with the environment in the following steps:
[0128] (1) The drone is to move;
[0129] (2) The drone collects data from the PoI according to the data collection model;
[0130] (3) If the UAV is within the communication range of any edge node, it exchanges, synchronizes, and updates the mission status information with the edge system;
[0131] (4) The UAV estimates the parameters of the PoI according to the parameter estimation model;
[0132] (5) PoI generates new data packets according to the data flow model and updates the data flow age.
[0133] Then, the environment state shifts to s i+1 , each agent obtains a local observation of the new state Next, the agent calculates ω Determines whether to stop the current option. If the option is stopped, set f u ←True. Here the β of different options ω The stopping conditions are as follows:
[0134] ω=1, that is, when the agent's option is to collect data, the stopping condition is i―i ω ≥M c That is, the agent continuously collects M c After a time slot has been filled, the options need to be re-determined;
[0135] ω=2, that is, when the agent's option is to visit the edge node, the stopping condition is that the agent successfully visits the edge node, that is You need to redefine your options.
[0136] Afterwards, each agent will experience Stored in the experience replay pool of each corresponding option In the above, the parameters of the option value (critic) network are updated by the gradient descent method, the parameters of the option strategy (executor) network are updated by the gradient ascent method, and then the option target network is updated by the soft update factor τ.
[0137] If agent u is in the current time slot t i If an option is terminated, the environment will give the agent feedback on the reward corresponding to the option. After that, the agent will experience Deposit Option Experience Replay Next, the agent Sampling a batch of experience And use the loss function shown in the following formula to update the network parameters according to the gradient descent method Right now:
[0138]
[0139] Finally, according to the soft update factor τ Ω Update target options - value target network parameters.
[0140] Furthermore, the proposed DRL-DCEA algorithm is compared with the DRL-ASPT algorithm, MADDPG algorithm, Edics algorithm, greedy strategy and random strategy in the test scenario. Figure 7 , Figure 8 The changes of two performance indicators are shown when the number of drones and the perception radius of drones are different.
[0141] First, the impact of the number of drones on the algorithm performance when it is 1, 2, 3, and 4 is evaluated. Figure 7 As shown. The results show that, except for the MADDPG algorithm, the mean age of data streams at each PoI decreases as the number of drones increases. This is because as the number of drones increases, the chances of data packets in the PoI being collected increase. Overall, the greedy strategy and the random strategy perform the worst. Among the remaining algorithms, the performance of MADDPG, Edics, DRL-ASPT, and DRL-DCEA shows a trend from weak to strong. In addition, DRL-DCEA performs better than other algorithms in terms of the mean age of data streams. For example, when the number of drones is 4, the mean age of data streams of DRL-DCEA is 34.43% lower than that of the second-best DRL-ASPT. In addition, from the perspective of energy consumption, such as Figure 7 As shown, it also presents a similar pattern to the mean age of the data stream. It should be pointed out that the DRL-DCEA proposed in the present invention reduces the energy consumption by 11.40% when the number of drones is 4 compared with the second-ranked DRL-ASPT algorithm. This shows that the DRL-DCEA algorithm has achieved better response performance in dynamic scenarios by introducing a multi-agent hierarchical reinforcement learning architecture, improved the data freshness index, and made the decision of multi-agents more flexible and efficient.
[0142] Next, the impact of perception radius of 8, 10, 12, and 14 on algorithm performance was evaluated. Figure 8As shown. The results show that as the perception radius of the drone gradually increases, the four algorithms based on reinforcement learning (DRL-ECDA, DRL-ASPT, Edics and MADDPG) all show a trend of gradually decreasing mean data stream age. When the perception radius changes from 12 to 14, the downward trend of the mean data stream age weakens, and even slightly increases. This is because by increasing the perception radius, the added PoI is relatively far away from the drone, which reduces the data collection speed and makes it difficult for the drone to collect more data. On the whole, when the perception radius is 10, DRL-ECDA reduces the mean data stream age by 29.74% compared with DRL-ASPT. On the other hand, with the change of the perception radius, the DRL-ECDA proposed in the present invention also shows the best performance in the energy consumption index among the comparison algorithms. For example, when the perception radius is 8, DRL-ECDA reduces the energy consumption by 24.33% compared with the second-best DRL-ASPT.
[0143] In summary, according to the embodiments of the present invention, at least the following technical effects are achieved:
[0144] First, the data stream age in dynamic scenarios is established and defined as a data freshness indicator to evaluate the dynamic response performance of UAV collaborative data collection. The data stream age considers both the amount of data in the data packet and the residence time of the data packet. Compared with the traditional information age, it can better reflect the data freshness of a data stream composed of multiple data packets.
[0145] Second, based on the architecture of edge system-assisted indirect task state information sharing in the existing technology, we further proposed parameter estimation based on Poisson distribution random process and sample exchange mechanism based on edge system, which improved the accuracy of PoI dynamic parameter estimation by drone mobile crowd intelligence perception system.
[0146] Third, to solve the problem of UAV data collection in dynamic scenarios, an algorithm based on multi-agent hierarchical deep reinforcement learning (DRL-DCEA) is proposed. The algorithm regards the macro and micro behaviors of UAVs as options and actions respectively. The data collection and access to edge nodes of UAVs under timeliness requirements are regarded as two macro options. By introducing option decision-making, the dynamic responsiveness of UAVs in distributed autonomous decision-making is improved.
[0147] Fourth, simulation experiments verified that in a typical dynamic mobile crowd-sensing scenario, the proposed DRL-DCEA algorithm is superior to baseline algorithms such as DRL-ASPT, Edics, MADDPG, greedy strategy, and random strategy in terms of data stream mean age, energy consumption, and other evaluation indicators. For example, compared with the DRL-ASPT algorithm, the DRL-DCEA algorithm reduces the data stream mean age by 29.74% and the energy consumption by 24.33%.
[0148] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0149] The above is an introduction to a method embodiment. The following is a further explanation of the solution of the present invention through an apparatus embodiment.
[0150] Fig. 9 A structural diagram of a multi-UAV dynamic task intelligent allocation device based on edge system and data stream age provided by an embodiment of the present invention, such as Fig. 9 As shown, the multi-UAV dynamic task intelligent allocation device 900 may include:
[0151] The setting module 910 is used to set a working scene, which includes multiple drones, multiple PoIs for data points to be collected, and multiple edge nodes. The multiple edge nodes constitute an edge system for indirect information sharing between drones.
[0152] The allocation module 920 is used to, in this working scenario, regard the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information as options, regard the micro-behavior of the drone in the time slot as an action, determine the optimization goal of the drone's mobile crowd perception task as minimizing the mean age of the data stream of all PoIs in the data collection process under energy constraints, regard each drone as an intelligent agent, and establish a multi-agent option-based partially observable Markov decision process model, so as to make macro-option decisions and micro-action decisions through multi-agent hierarchical deep reinforcement learning.
[0153] Understandably, Fig. 9 Each module / unit in the multi-UAV dynamic task intelligent allocation device 900 has the function of realizing Figure 1The functions of each step in the multi-UAV dynamic task intelligent allocation method 100 shown and its corresponding technical effects are not repeated here for the sake of brevity.
[0154] Fig.10 FIG. 1 is a structural diagram of an exemplary electronic device capable of implementing an embodiment of the present invention. Fig.10 As shown, the electronic device 1000 may include a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0155] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0156] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as method 100. For example, in some embodiments, the method 100 may be implemented as a computer program product, including a computer program, which is tangibly contained in a computer-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the method 100 in any other appropriate manner (eg, by means of firmware).
[0157] The various embodiments described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0158] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0159] In the context of the present invention, a computer-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a computer-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0160] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute method 100 and achieve the corresponding technical effect achieved by executing the method in an embodiment of the present invention. For the sake of concise description, they will not be repeated here.
[0161] In addition, the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method 100 is implemented.
[0162] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and the present invention is not limited here.
[0163] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for intelligent allocation of multi-UAV dynamic tasks based on edge system and data stream age, characterized in that: The method comprises: Set up a working scenario, which includes multiple drones, multiple PoIs of data points to be collected, and multiple edge nodes. Multiple edge nodes form an edge system for indirect information sharing between drones; In this working scenario, the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information are regarded as options, the micro-behavior of the drone in the time slot is regarded as an action, and the optimization goal of the drone's mobile crowd perception task is determined as minimizing the mean age of the data stream of all PoIs in the data collection process under energy constraints. Each drone is regarded as an intelligent agent, and a multi-agent partially observable Markov decision process model based on options is established, so that macro-option decisions and micro-action decisions are made through multi-agent hierarchical deep reinforcement learning.
2. The method according to claim 1, characterized in that The age of a data stream is defined as: Among them, A i (p) indicates PoI p at t i The age of the data stream at time, Refers to t i The set of data packets contained in PoI p at the time, δ m is the amount of data in data packet m, l m The time when the data packet is generated.
3. The method according to claim 1, characterized in that The process of adding new data to PoI is considered as a Poisson process, specifically: The new data in PoI p are regarded as random events, and these events are independent of each other and have equal probability of occurrence. i The number of new data in the time slot follows the Poisson distribution, and the time slot t i The number of times new data appears in PoI p is K i (p), the probability of each event is λ(p), then K i (p)~Poisson(λ(p)); Assuming that for the same PoIp, the amount of new data added each time is a fixed value D(p), then the time slot t i The amount of new data added is δ i (p) = K i (p)D(p), these data are packaged into data packets m i , and added to the data stream of PoI p, the data stream will be placed in the drop box of the PoI, waiting for collection by drones that visit in the future; The UAV obtains the data flow information of PoI p, regards the data volume of the data packets generated in these Poisson processes as samples, and estimates the Poisson distribution parameter λ(p) and the proportion D(p) of PoI p.
4. The method according to claim 1, characterized in that: The drone collects data about PoI including: In the same time slot, the drone can transmit data with at most C PoIs. i In the process, the drone first visits its perception radius ρ s The drop box of all PoIs within the range is used to obtain the data flow age of all PoIs. Here, the perception radius ρ s All PoIs in the like Then the drone u will be Select the C PoIs with the largest data stream age for data collection; if You will choose Collect data from all PoIs within the time slot t i Select all PoIs collected from data as a set In time slot t i in Approximate data transmission rate between UAV u at position and PoIp at position (x(p), y(p)) It is defined by the Shannon capacity taking into account the path loss exponent as shown below: in, represents the approximate data transmission rate, B represents the total effective bandwidth, q0 represents the average transmission power of the UAV, and z 2 is the white noise power of Gaussian distribution, g0 represents the channel gain at the reference distance, α represents the path loss index, and dist(u,p,i) represents the distance between UAV u and PoI p in time slot t i At the discrete grid scale l grid The Euclidean distance under For each PoI to collect data Drones will be Within a certain time, each data packet in the data stream is collected from its throwing box in a first-in-first-out order, and the data volume of each data packet is recorded. After that, the data stream age of the PoI will be automatically updated.
5. The method according to claim 1, characterized in that The state space of the partially observable Markov decision process model includes the location of the PoI, the data generation parameters, the amount of remaining data of the PoI in the current time slot, the age of the data stream in the current time slot, the location and energy of all UAVs in the current time slot; The observation space of the partially observable Markov decision process model includes the location of the PoI, the estimated value of the data generation parameter, the estimated value of the remaining data volume of the PoI in the current time slot, the estimated value of the age of the data stream in the current time slot, and the location and energy of the drone in the current time slot; The option space of the partially observable Markov decision process model represents the macroscopic active behavior of the agent, and the action space represents the microscopic behavior of the agent, specifically the movement vector of the agent in the time slot; The reward function of the partially observable Markov decision process model includes the reward for the agent to collect data, the penalty for energy consumption, the reward given by the edge system, and the penalty for the agent to move out of the working area.
6. The method according to claim 5, characterized in that The training of multi-agent hierarchical deep reinforcement learning includes: Before training begins, each agent is equipped with an option-value network for macro option decision making And configure a pair of executor-critic network inside each option Each network also has a corresponding target network; in addition, each intelligent body is also equipped with a corresponding option experience replay pool and option experience replay pool The training consists of EP rounds. In each round, the environment is set to the initial state s0, and each agent obtains the initial observation of the environment The option termination identifier f of each agent u is set to True; Each round has M time slots. In each time slot, if the option termination identifier f u is True, the agent chooses the macro option according to the following formula: Afterwards, f u Set to False, for f u If it is False, the agent maintains the original option and does not need to choose the option. Next, each agent generates actions according to the strategy function inside the option. And add Gaussian noise to the action. Next, each agent interacts with the environment according to the following steps: The drone is following the noise to move; The drone collects data from the PoI according to the data collection node model; If the drone is within the communication range of any edge node, it will exchange, synchronize, and update mission status information with the edge system; The drone estimates the parameters of the PoI according to the parameter estimation model; PoI generates new data packets according to the data flow model and updates the data flow age; Then, the environment state shifts to s i+1 , each agent obtains a local observation of the new state Next, the agent decides whether to stop the current option. If the option is stopped, it sets f u ←True, the stopping conditions for different options are as follows: ω=1, that is, when the agent's option is to collect data, the stopping condition is i―i ω ≥M c , that is, the agent continuously collects M c After a time slot has been filled, the options need to be re-determined; ω=2, that is, when the agent's option is to visit the edge node, the stopping condition is that the agent successfully visits the edge node, that is When you need to re-determine the options; Afterwards, each agent will experience Stored in the experience replay pool of each corresponding option In the above example, the value network parameters in the option are updated by the gradient descent method, the policy network parameters in the option are updated by the gradient ascent method, and then the target network in the option is updated by the soft update factor τ; If agent u is in the current time slot t i If an option is terminated, the environment will give the agent feedback on the reward corresponding to the option. After that, the agent will experience Deposit Option Experience Replay Next, the agent Sampling a batch of experience And use the loss function shown in the following formula to update the network parameters according to the gradient descent method Finally, according to the soft update factor τ Ω Update target options - values of target network parameters.
7. A multi-UAV dynamic task intelligent allocation device based on edge system and data stream age, characterized in that: The device comprises: A setting module is used to set a working scenario. The working scenario includes multiple drones, multiple PoIs of data points to be collected, and multiple edge nodes. Multiple edge nodes form an edge system for indirect information sharing between drones. The allocation module is used to regard the two macro-active behaviors of the drone going to the data point to be collected to perform the task and the drone going to the edge node to share the task status information as options in this working scenario, regard the micro-behavior of the drone in the time slot as an action, and determine the optimization goal of the drone's mobile crowd perception task as minimizing the mean age of the data stream of all PoIs in the data collection process under energy constraints. Each drone is regarded as an intelligent agent, and a multi-agent partially observable Markov decision process model based on options is established, so as to make macro-option decisions and micro-action decisions through multi-agent hierarchical deep reinforcement learning.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Unmanned aerial vehicle task unloading method and system based on reinforcement learning in edge calculation
CN111787509A
Multi-unmanned aerial vehicle auxiliary edge computing resource allocation method based on task prediction
CN112351503A
Unmanned aerial vehicle task allocation method based on edge system and reinforcement learning
CN117707198A
Task system of unmanned cluster and physical test method
CN118170115A
Cited By
Unmanned aerial vehicle assisted Internet of Things data acquisition method for multi-mechanism overlapping deployment
CN120475437A