Method for dynamic task allocation of multiple unmanned aerial vehicles based on edge system and data flow age

By introducing edge system and data stream age indicators into multi-UAV systems and combining them with multi-agent hierarchical deep reinforcement learning, the flexibility and efficiency issues of multi-UAV data acquisition in dynamic environments are solved, and the average age of data streams and energy consumption are optimized.

CN120013136BActive Publication Date: 2025-11-07TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510040851.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-11-07
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle collaborative data acquisition from multiple drones in dynamic environments, particularly in their inability to effectively measure the freshness of data streams composed of multiple data packets. Furthermore, existing methods exhibit poor flexibility and efficiency in dynamic environments.

Method used

A dynamic task intelligent allocation method for multiple UAVs based on edge systems and data stream age is adopted. Through multi-agent hierarchical deep reinforcement learning, a data stream age index is defined. Combined with Poisson distributed stochastic process and edge system parameter estimation, the macroscopic and microscopic behaviors of UAVs are optimized to minimize the mean data stream age.

Benefits of technology

It improves the efficiency and flexibility of UAV data acquisition in dynamic environments, reduces the average age of data streams and energy consumption, and enhances data freshness and the accuracy of parameter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013136B_ABST
    Figure CN120013136B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a multi-unmanned aerial vehicle dynamic task intelligent allocation method based on an edge system and data stream age. The method comprises: setting a working scene, including a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected, and a plurality of edge nodes, the edge nodes forming an edge system; in the working scene, the unmanned aerial vehicles go to the data points PoI to be collected to perform tasks, and the two macro-initiative behaviors of the unmanned aerial vehicles going to the edge nodes to share task state information are regarded as options, the micro-behaviors of the unmanned aerial vehicles in a time slot are regarded as actions, and the optimization target of the mobile crowd sensing task of the unmanned aerial vehicles is determined as minimizing the average data stream age of all PoI in the data collection process under the energy constraint, each unmanned aerial vehicle is regarded as an intelligent agent, a multi-agent option-based partially observable Markov decision process model is established, macro-option decision and micro-action decision are made through multi-agent hierarchical deep reinforcement learning, and the dynamic task intelligent allocation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned aerial vehicles, and particularly relates to a multi-unmanned aerial vehicle dynamic task intelligent allocation method based on an edge system and data flow age. BACKGROUND

[0002] Since the concept of mobile crowd sensing was proposed, task allocation and task planning in multi-agent cooperative data collection have been regarded as important problems in mobile crowd sensing, and a large number of studies have been carried out on this problem, and many frameworks and algorithms for task allocation in mobile crowd sensing have been proposed.

[0003] Some studies use dynamic programming and heuristic optimization algorithms to find optimal or suboptimal solutions to task allocation problems under complex constraints. However, these methods have weak portability for different environments and cannot handle dynamic environments. In the problem of crowd sensing, most existing studies assume that the total amount of data to be collected in the Point of Interest (PoI) is constant, and the scene dynamics are weak. However, in real-world mobile crowd sensing tasks, the total amount of data in each PoI will not remain constant, for example, when a camera continuously takes photos or records videos, the amount of data stored inside it will dynamically increase over time.

[0004] In a dynamic environment, the timeliness of data collection is even more important. In general, it is hoped that the data in the PoI will be collected as soon as possible, rather than being stored in the PoI for a long time and being unable to be perceived by external unmanned aerial vehicles, so as to be unable to make timely feedback and decisions on sudden situations (such as abnormal temperature, suspicious personnel, etc.). Among them, the Age of Information (AoI) is an important indicator to measure the freshness of information, which is defined as the stagnation time of data at a certain place. A large number of existing studies take the information age as a target or constraint to optimize the transmission delay and other dynamic performance of the mobile crowd sensing system. However, the information age can only be used to measure the freshness of a single data packet, and cannot handle the freshness of a data flow composed of multiple data packets with different generation times.

[0005] In recent years, methods based on deep reinforcement learning have attracted extensive attention of researchers because they can solve the joint task allocation and path planning problem in more complex scenarios. In general, from the high-level research results in recent years, a large amount of work on task allocation and path planning problems in the field of mobile crowd sensing data collection, including the most advanced research work, is almost based on deep reinforcement learning or multi-agent deep reinforcement learning. Such methods currently achieve the optimal performance in the data collection decision-making process of multiple unmanned aerial vehicles in mobile crowd sensing. However, the decision variables of these methods are often the moving direction and distance of unmanned aerial vehicles in each time slot, which is difficult to explain the reasons for the decision-making of unmanned aerial vehicles from the macro and micro levels, and therefore has poor flexibility. SUMMARY

[0006] In a first aspect, embodiments of the present application provide a multi-unmanned aerial vehicle dynamic task intelligent allocation method based on an edge system and an age of data stream, which comprises:

[0007] Setting a working scenario, the working scenario comprising a plurality of unmanned aerial vehicles, a plurality of points of interest (PoIs) and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles;

[0008] In the working scenario, the unmanned aerial vehicles go to the data points to be collected to perform tasks, and the two macro-initiative behaviors of the unmanned aerial vehicles going to the edge nodes to share task state information are regarded as options, the micro-behaviors of the unmanned aerial vehicles in a time slot are regarded as actions, and the optimization goal of the mobile crowd sensing task of the unmanned aerial vehicles is determined as minimizing the average of the age of data stream (AoDS) of all PoIs in the data collection process under the energy constraint. Each unmanned aerial vehicle is regarded as an agent, a multi-agent partially observable Markov decision process model based on options is established, and macro-option decision and micro-action decision are made through multi-agent hierarchical deep reinforcement learning.

[0009] In some implementable manners of the first aspect, the age of data stream is defined as:

[0010]

[0011] wherein, A i (p) represents the age of data stream of the PoI p at the t i moment, which refers to a set composed of data packets contained in the PoI p at the t i moment, δ m is the data volume in the data packet m, and ι m is the moment when the data packet is generated.

[0012] In some possible implementation manners of the first aspect, the process of the newly added data of the PoI is regarded as a Poisson process, and specifically:

[0013] The newly added data in the PoI p is regarded as a random event, and the events are independent of each other and have equal occurrence probability, and the number of times of occurrence of the newly added data in the time slot t i is subject to a Poisson distribution, and the number of times of occurrence of the newly added data in the time slot t i is K i (p), and the probability of each event is λ(p), then K i (p) ~ Poisson(λ(p));

[0014] It is assumed that, for the same PoI p, the data amount of each newly added data is a fixed value D(p), and the data amount of the newly added data in the time slot t i is δ i (p) = K i (p)D(p), and the data is packaged as a data packet m i , and is added to the data stream of the PoI p, and the data stream is placed in the throwing box of the PoI, and is waiting to be collected by a future visiting unmanned aerial vehicle;

[0015] The unmanned aerial vehicle obtains the data stream information of the PoI p, regards the data amount of the data packet generated in the Poisson process as a sample, and estimates the Poisson distribution parameter λ(p) and the proportional amount D(p) of the PoI p.

[0016] In some possible implementation manners of the first aspect, the data collection of the PoI by the unmanned aerial vehicle includes:

[0017] In the same time slot, the unmanned aerial vehicle transmits data to at most C PoIs, in the time slot t i , the unmanned aerial vehicle first accesses the throwing boxes of all the PoIs within a sensing radius ρ s to obtain the data stream ages of all the PoIs, and all the PoIs within the sensing radius ρ s are denoted as a set If , the unmanned aerial vehicle u selects C PoIs with the largest data stream ages from to collect data; if , the unmanned aerial vehicle u selects all the PoIs in to collect data, and all the PoIs selected by the unmanned aerial vehicle u in the time slot t i are denoted as a set

[0018] In the time slot t i , the approximate data transmission rate of the unmanned aerial vehicle u located at and the PoI p located at (x(p), y(p)) The Shannon capacity, which takes into account the path loss exponent, is defined as follows:

[0019]

[0020] in, Let B represent the approximate data transmission rate, B represent the total effective bandwidth, q0 represent the average transmission power of the UAV, and z represent the average data transmission rate. 2 Let g be the Gaussian distributed white noise power, g0 represent the channel gain at the reference distance, α represent the path loss exponent, and dist(u,p,i) represent the distance between the UAV u and PoI p in time slot t. i At discrete grid scale l grid Euclidean distance;

[0021] For each piece of data to be collected Drones will Within a given time period, each data packet in the data stream is collected from its drop box in a first-in, first-out order, and the data volume of each data packet is recorded. After that, the data stream age of the PoI will be automatically updated.

[0022] In some possible implementations of the first aspect, the state space of the partially observable Markov decision process model includes the location of the PoI, the data generation parameters, the remaining data amount of the PoI in the current time slot, the data stream age in the current time slot, and the location and energy of all UAVs in the current time slot.

[0023] The observation space of a partially observable Markov decision process model includes the location of the Point of Interest (PoI), the estimated values ​​of the data generation parameters, the estimated value of the remaining data volume of the PoI in the current time slot, the estimated age of the data stream in the current time slot, and the location and energy of the UAV in the current time slot.

[0024] In a partially observable Markov decision process model, the option space represents the agent's macroscopic active behavior, and the action space represents the agent's microscopic behavior, specifically the agent's movement vector in the time slot.

[0025] The reward function of a partially observable Markov decision process model includes the reward for the agent to collect data, the penalty for energy consumption, the reward given by the edge system, and the penalty for the agent to move out of the work area.

[0026] Among the possible implementations of the first aspect, training for multi-agent hierarchical deep reinforcement learning includes:

[0027] Before training begins, each agent is equipped with an option-value network for macro-level option decisions. And configure a pair of executor-commentator networks within each option. Each network also has a corresponding target network; in addition, each agent is also equipped with a corresponding option internal experience replay pool and the option experience replay pool

[0028] The training consists of EP rounds, in each round, the environment is set to the initial state s0, and each agent obtains the initial observation of the environment The option termination identifier f of each agent u is set to True;

[0029] Each round has M time slots, in each time slot, if the option termination identifier f u is True, the agent selects the macro option according to the following formula:

[0030]

[0031] Then, f u is set to False, for the case that f u is False, the agent maintains the original option and does not need to select the option, next, each agent generates an action according to the policy function inside the option and adds Gaussian noise to the action, next, each agent interacts with the environment according to the following steps:

[0032] The UAV moves according to the action added with noise; ;

[0033] The UAV collects data on the PoI according to the data collection module;

[0034] If the UAV is within the communication range of any edge node, the edge system exchanges, synchronizes and updates the task state information;

[0035] The UAV estimates the parameters of the PoI according to the parameter estimation model;

[0036] The PoI generates a new data packet according to the data flow model and updates the data flow age;

[0037] Then, the environment state moves to s i+1 , and each agent obtains the local observation of the new state Next, the agent decides whether to stop the current option, if the option is stopped, f u is set to True, and the stop conditions of different options are as follows:

[0038] ω = 1, that is, when the option of the agent is to collect data, the stop condition is i - i ω ≥ M c , that is, the agent continuously collects M cAfter the time slot data, the options need to be re-determined;

[0039] ω = 2, that is, when the option of the agent is to access the edge node, the stopping condition is that the agent successfully accesses the edge node, that is, , the options need to be re-determined;

[0040] After that, each agent will store the experience In the experience replay pool inside the corresponding option of each agent , update the value network parameters inside the option by gradient descent method, update the policy network parameters inside the option by gradient ascent method, and then update the target network inside the option by soft update factor τ;

[0041] If the agent u has an option termination at the current time slot t i , the environment will give the agent the reward corresponding to the option After that, the agent will store the experience In the option experience replay Next, the agent samples a batch of experiences From the And use the loss function shown in the following formula to update the network parameters according to the gradient descent method

[0042] Finally, update the target option-value target network parameters by soft update factor τ Ω .

[0043] In a second aspect, an embodiment of the present application provides a multi-unmanned aerial vehicle dynamic task intelligent allocation device based on an edge system and data flow age, which comprises:

[0044] A setting module is configured to set a working scenario, wherein the working scenario comprises a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected, and a plurality of edge nodes, the plurality of edge nodes form an edge system for indirect information sharing between the unmanned aerial vehicles;

[0045] An allocation module is configured to, in the working scenario, cause the unmanned aerial vehicles to go to the data points PoI to be collected to perform tasks, and cause the unmanned aerial vehicles to go to the edge nodes to share task state information, wherein the two macro-initiative behaviors are regarded as options, the micro-behaviors of the unmanned aerial vehicles in a time slot are regarded as actions, and an optimization target of a mobile crowd sensing task of the unmanned aerial vehicles is determined as minimizing the average data flow age of all PoI in a data collection process under an energy constraint, each unmanned aerial vehicle is regarded as an agent, a multi-agent partially observable Markov decision process model based on the options is established, and thus macro-option decision and micro-action decision are made through multi-agent hierarchical deep reinforcement learning.

[0046] In a third aspect, an electronic device is provided, and the electronic device includes at least one processor, and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0047] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to enable a computer to perform the method described above.

[0048] According to the embodiments of the present application, the following technical effects are achieved at least:

[0049] First, the data stream age in a dynamic scene is established and defined as a data freshness indicator, which is used to evaluate the dynamic response performance of the unmanned aerial vehicle cooperative data collection. The data stream age considers both the data amount in the data packet and the residence time of the data packet, and is more capable of reflecting the data freshness of the data stream composed of multiple data packets than the traditional information age.

[0050] Second, based on the existing edge system assisted indirect task state information sharing architecture, a parameter estimation based on Poisson distribution random process and a sample exchange mechanism based on the edge system are further proposed, which improves the accuracy of the dynamic parameter estimation of the PoI by the unmanned aerial vehicle mobile crowd sensing system.

[0051] Third, to solve the problem of unmanned aerial vehicle data collection in a dynamic scene, a multi-agent hierarchical deep reinforcement learning algorithm (hereinafter referred to as DRL-DCEA) is proposed. The macro and micro behaviors of the unmanned aerial vehicle are regarded as options and actions respectively. The data collection and access to the edge node by the unmanned aerial vehicle under the time requirement are regarded as two macro options. By introducing the option decision, the dynamic responsiveness of the unmanned aerial vehicle in the distributed autonomous decision is improved.

[0052] Fourth, simulation experiments verify that, in a typical dynamic mobile crowd sensing scene, compared with the baseline algorithms such as DRL-ASPT, Edics, MADDPG, greedy strategy and random strategy, the proposed DRL-DCEA algorithm has superior performance in evaluation indicators such as data stream age mean and energy consumption. For example, compared with the DRL-ASPT algorithm, the DRL-DCEA algorithm reduces the data stream age mean by 29.74% and the energy consumption by 24.33%.

[0053] It should be understood that the content described in the summary section is not intended to limit the key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0054] The above and other features, advantages, and aspects of embodiments of the present application will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The following drawings are provided to assist in understanding the present application and constitute a part of this specification. It should be noted, that the figures are provided on an as-is basis with nothing drawn to scale and the same or similar reference numerals are used throughout the various drawings to refer to same or similar elements. In the drawings:

[0055] Figure 1 A flow chart of a multi-UAV dynamic task intelligent allocation method based on edge system and data flow age provided by an embodiment of the present application;

[0056] Figure 2 A schematic diagram of data flow age change over time;

[0057] Figure 3 A schematic diagram of Poisson process parameter estimation based on task state information;

[0058] Figure 4 A box plot of PoI parameter estimation value;

[0059] Figure 5 A schematic diagram of DRL-DCEA algorithm architecture;

[0060] Figure 6 A schematic diagram of DRL-DCEA algorithm flow;

[0061] Figure 7 A schematic diagram of the impact of different numbers of UAVs on performance;

[0062] Figure 8 A schematic diagram of the impact of different perception radii on performance;

[0063] Figure 9 A structural diagram of a multi-UAV dynamic task intelligent allocation device based on edge system and data flow age provided by an embodiment of the present application;

[0064] Figure 10 A structural diagram of an exemplary electronic device capable of implementing an embodiment of the present application. DETAILED DESCRIPTION

[0065] In order to make the objects, technical solutions and advantages of embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in a clear and complete manner with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0066] In addition, the term "and / or" in the present application only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in the present application generally represents that the front and rear associated objects are in an "or" relationship.

[0067] To solve the technical problems in the background art, the embodiments of the present application provide a multi-unmanned aerial vehicle dynamic task intelligent allocation method, device and equipment based on an edge system and data stream age, and a storage medium. In the following, the multi-unmanned aerial vehicle dynamic task intelligent allocation method, device and equipment based on an edge system and data stream age provided by the embodiments of the present application will be described in detail through specific embodiments combined with the drawings.

[0068] Figure 1 The flow chart of the multi-unmanned aerial vehicle dynamic task intelligent allocation method based on an edge system and data stream age provided by the embodiments of the present application is shown in Figure 1 The multi-unmanned aerial vehicle dynamic task intelligent allocation method 100 can include the following steps:

[0069] S110, setting a working scenario, the working scenario including a plurality of unmanned aerial vehicles, a plurality of PoIs and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles.

[0070] S120, under the working scenario, the unmanned aerial vehicles going to the data points to be collected to perform tasks, the two macro-initiative behaviors of the unmanned aerial vehicles going to the edge nodes to share task state information being regarded as options, the micro-behaviors of the unmanned aerial vehicles within a time slot being regarded as actions, the optimization target of the mobile crowd sensing task of the unmanned aerial vehicles being determined as minimizing the average data stream age of all PoIs in the data collection process under the energy constraint, each unmanned aerial vehicle being regarded as an intelligent agent, a multi-agent option-based partially observable Markov decision process model being established, and macro-option decision and micro-action decision being made through multi-agent hierarchical deep reinforcement learning.

[0071] In order to facilitate further understanding, the above steps will be described in detail in combination with specific embodiments as follows:

[0072] (1) Setting a working scenario

[0073] A rectangular working area is considered, P PoIs are uniformly distributed in the rectangular working area according to geographical positions, U unmanned aerial vehicles are used to collect data from the PoIs, and the unmanned aerial vehicles do not have direct communication capability between them. N edge nodes are set to form an edge system for indirect information sharing between the unmanned aerial vehicles.

[0074] (2) Data stream model of PoI

[0075] On the basis of AoI, the Age of Data Stream (AoDS) is proposed to measure the information freshness of multiple data packets generated at different times in a PoI. The index considers the data volume of each data packet and the time it is stored in the PoI, and is defined as:

[0076]

[0077] where A i (p) represents the data stream age of PoI p at t i , m i (p) represents the set of data packets contained in PoI p at t m , δ m is the data volume in data packet m, and ι a is the time when the data packet is generated. The change of data stream age over time is described, where m a , m c , and m a are generated at t b , t c , and t d , respectively, m a and m b are collected at t e , and m c is generated at t i . Figure 2

[0078] In the case of discrete time slots, the data stream age can also be calculated incrementally, that is:

[0079]

[0080] where M is the set of data packets collected by the UAV. In the case of discrete time slots, for ease of expression and calculation, the time when each PoI generates a new data packet is set as the end of the time slot.

[0081] (3) UAV parameter estimation model for PoI

[0082] The process of adding new data to a PoI is regarded as a Poisson process. Specifically, the new data (such as taking photos) in PoI p is regarded as a random event, and these events are independent of each other and have equal occurrence probability. Therefore, the number of new data in time slot t i follows a Poisson distribution, where K i (p) represents the number of new data in PoI p in time slot t i , and λ(p) represents the probability of each event.(p)~Poisson(λ(p)).

[0083] Assuming that for the same PoI p, the amount of data added each time is a fixed value D(p) (e.g., each photo taken occupies the same amount of storage space), then the time slot t i The amount of newly added data is δ i (p)=K i (p)D(p), these data are packaged into data packets m i The data is then added to the data stream of the PoI p, and the data stream is placed in the drop box of that PoI, waiting to be collected by a drone that visits in the future.

[0084] drones from Figure 3 The data stream information of PoI p obtained in the fused TSI is shown. The data volume of these packets generated in the Poisson process is regarded as a sample, and the Poisson distribution parameter λ(p) and the proportionality D(p) of p are estimated. Let the data volume sample of PoI p by the UAV be . According to the first and second moments of the Poisson process, we have:

[0085]

[0086] Therefore, the estimated rate of data growth of the PoI can be calculated. Poisson distribution parameter estimator and estimates of the data growth rate Figure 4 This demonstrates the parameter estimation results for a Point of Interest (PoI) in an example scenario with specific parameters. It can be assumed that the more samples the UAV observes in a given PoI data packet, the more accurate the parameter estimation for that PoI will be.

[0087] (4) Data acquisition model of UAV

[0088] Assuming each drone can use multiple antennas and frequency division multiplexing mechanisms to transmit data simultaneously with multiple Points of Interest (PoIs), to ensure the transmission rate between the drone and PoIs, it is assumed that within the same time slot, the drone can transmit data with at most C PoIs. In time slot t... i In the process, the drone first visits its sensing radius ρ s To obtain all PoIs (denoted as set), use a drop box containing all PoIs. The age of the data stream. If... >C, then the drone u will Select the C PoIs with the oldest data stream age for data collection; if The drone will choose Data is collected from all PoIs within the time slot t, and the drone u is placed in the time slot t. iThe set of all PoIs selected for data collection is denoted as According to the foregoing definitions, there are Since the transmission data volume required for the UAV to access the PoI drop box is small, for the sake of simplicity, the transmission delay problem is not considered here.

[0089] At time slot t i , the UAV u in position and the PoI p in position (x(p), y(p)) have an approximate data transmission rate defined by the Shannon capacity considering the path loss exponent, as shown in the following formula:

[0090]

[0091] where B represents the total effective bandwidth (when the UAV transmits data with multiple PoIs at the same time, it is assumed that the total bandwidth is evenly distributed), q0 represents the average transmission power of the UAV, z 2 is the white noise power of Gaussian distribution, g0 represents the channel gain at the reference distance (for example, one unit distance, which is 1 m in this paper), a represents the path loss exponent, and dist(u, p, i) represents the Euclidean distance (corresponding to the grid number) between the UAV u and the PoI p at time slot t i in the discrete grid scale l grid , that is:

[0092]

[0093] For each PoI to be collected data, the UAV will collect each data packet in the data stream from its drop box in the First Input First Output (FIFO) order within the time , and record the data volume of each data packet, after which the data stream age of the PoI will be automatically updated.

[0094] (5) Multi-agent hierarchical deep reinforcement learning

[0095] In a dynamic scenario, the timeliness of UAV data collection becomes more important. To make the UAV task planning more flexible under the timeliness requirement, two kinds of macro-initiative behaviors of the UAV are considered, including:

[0096] · The UAV goes to the data collection point to perform the task;

[0097] · The UAV goes to the edge node to share the task state information.

[0098] ​Therefore, it is necessary to design the macro and micro level action selection methods of the UAV accordingly. The UAV regards the above two high-level macro behaviors as action options, reflecting the overall goal of the UAV in a period (several time slots in the future); the micro behavior represents the specific action of the UAV in each time slot, including the selection of moving direction, distance, speed and other detailed behaviors.

[0099] The mobile crowd sensing task execution process of the UAV is divided into M time slots, and the time span of each time slot is Δt. The optimization goal of the mobile crowd sensing task is to minimize the average data stream age of all PoIs in the collection process under the energy constraint, that is:

[0100]

[0101] Wherein, A i (p) represents the data stream age of the PoI. It can be seen that A i (p) is related to the data packet generation time and data volume contained in the PoI data stream, and each data packet is generated by the foregoing process and is transferred to the UAV with the data collection of the UAV. Since the problem introduces time dynamics, its complexity is no less than that of the multi-traveler problem, so the above optimization problem is very difficult to solve. Therefore, it is necessary to find a better UAV path planning and task allocation strategy to deal with the above time dynamics. Since the multi-agent reinforcement learning method has shown superior performance in sequence decision problems, and considering the hierarchical reinforcement learning ability to decompose complex behavior decision problems and high interpretability, the present application proposes to use multi-agent hierarchical reinforcement learning to solve the above complex optimization problem.

[0102] For the above timing decision problem, each UAV is regarded as an agent, and a multi-agent option-based partially observable Markov decision process (game) model is established The meanings and compositions of the elements will be explained below.

[0103] 1. The state space S of the environment is composed of the following parts:

[0104] The position of the PoI, the data generation parameter, the remaining data volume of the PoI in the current time slot, the data stream age in the current time slot, and the position and energy of all UAVs in the current time slot.

[0105] 2. The observation space O of the UAV is partially observable to the environment, and is composed of the following parts:

[0106] The position of the PoI, the data generation parameter estimate, the remaining data volume estimate of the PoI in the current time slot, the data stream age estimate in the current time slot, and the position and energy of the UAV in the current time slot.

[0107] 3, The option space Ω and the action space A are defined as follows:

[0108] The option space of the agent is Ω = {1, 2}, which represents the macro-level behavior of the agent. The option ω = 1 represents the agent going to the PoI to collect data, and the option ω = 2 represents the agent going to the edge system to exchange task state information.

[0109] The action space of the agent is the movement vector of the agent in a time slot, i.e. and is subject to the condition constraint. The length of the action vector of a single agent is 2.

[0110] 4, The reward function is as follows:

[0111] At the end of each time slot, the environment will give each agent a reward:

[0112]

[0113] The four terms in the above formula are the reward for the agent to collect data, the energy consumption penalty, the reward given by the edge system, and the penalty for the agent to move outside the working area (i.e., the out-of-bound behavior).

[0114] If the option ω starts at time slot t i and ends, the environment will give each agent an option reward for learning the parameters of the option-value network. The option reward is set as follows:

[0115]

[0116] where i ω represents the time slot when the option ω starts, and i―i ω represents the number of time slots that the option lasts. is defined as the average of the data collection reward of the agent within the option duration, and is defined as:

[0117]

[0118] Similarly, is defined as the average of the edge reward within the option duration, and is defined as:

[0119]

[0120] In addition, and represent the proportion coefficients of the data collection reward and the edge reward in the option reward, respectively.

[0121] The application proposes a "DRL-DCEA" algorithm, which is a kind of edge-assisted mobile crowd-sensing dynamic data collection algorithm based on multi-agent hierarchical deep reinforcement learning, wherein DCEA is the abbreviation of Data Collection or Edge Accessing, representing the two macro actions of "collecting data or accessing edge nodes". Figure 5 is the architecture of the DRL-DCEA algorithm. In this architecture, the actor-critic network architecture is adopted inside each option of each agent u, wherein is used to estimate the micro action Q value, is used to generate the micro action, wherein ω ∈ {1, 2} is used to represent two options; in each agent, there is additionally an used to estimate the value of the option and used for option decision.

[0122] The flow of the DRL-DCEA algorithm is shown in Figure 6 , and specifically includes the following steps:

[0123] Before training, each UAV is regarded as an agent, and each agent is equipped with an option-value network for macro option decision and a pair of actor-critic networks inside each option Each network also has a corresponding target network. In addition, each agent is also equipped with a corresponding intra-option experience replay pool and an option experience replay pool

[0124] The training of the algorithm consists of EP rounds, and in each round, the environment is set to the initial state s0, and each agent obtains the initial observation of the environment The option termination identifier f u of each agent (Boolean variable) is set to True.

[0125] There are M time slots in each round, and in each time slot, if the option termination identifier f u is True, the agent needs to select the macro option according to the following formula:

[0126]

[0127] After that, f u is set to False. For the case where f u is False, the agent maintains the original option and does not need to select the option. Next, each agent generates an action according to the policy function inside the option and adds Gaussian noise to the action. Next, each agent interacts with the environment according to the following steps:

[0128] (1) the UAV collects data from the PoI according to the data collection model; moves;

[0129] (2) the UAV collects data from the PoI according to the data collection model;

[0130] (3) if the UAV is within the communication range of any edge node, the UAV exchanges, synchronizes, and updates the task state information with the edge system;

[0131] (4) the UAV estimates the parameters of the PoI according to the parameter estimation model;

[0132] (5) the PoI generates a new data packet according to the data flow model and updates the data flow age.

[0133] Then, the environment state moves to s i+1 , and each agent obtains a local observation of the new state Next, the agent determines whether to stop the current option according to β ω , and if the option is stopped, sets f u ← True. The stop conditions of β ω for different options are as follows:

[0134] ω = 1, i.e., when the option of the agent is to collect data, the stop condition is i - i ω ≥ M c . That is, after the agent continuously collects M c time slot data, the option needs to be re-determined;

[0135] ω = 2, i.e., when the option of the agent is to access the edge node, the stop condition is that the agent successfully accesses the edge node, i.e. , the option needs to be re-determined.

[0136] After that, each agent stores the experience in the experience replay pool corresponding to the option inside the agent, updates the value (critic) network parameters inside the option according to the gradient descent method, updates the policy (actor) network parameters inside the option according to the gradient ascent method, and then updates the target network inside the option according to the soft update factor τ.

[0137] If the agent u has an option termination at the current time slot t i , the environment will give the agent a reward corresponding to the option. After that, the agent stores the experience in the option experience replay Next, the agent samples a batch of experiences from the experience replay and uses the loss function shown in the following formula to update the network parameters according to the gradient descent method That is:

[0138]

[0139] Finally, the soft update factor τ Ω Update target option-value target network parameters.

[0140] Further, the proposed DRL-DCEA algorithm and DRL-ASPT algorithm, MADDPG algorithm, Edics algorithm and greedy strategy, random strategy are compared in the test scene, Figure 7 、 Figure 8 The changes of the two performance indicators are shown when the number of UAVs and the sensing radius of UAVs are different.

[0141] Firstly, the influence of the number of UAVs on the performance of the algorithm is evaluated when the number of UAVs is 1, 2, 3 and 4, as shown in Figure 7 The results show that, except for the MADDPG algorithm, the average data flow age of each PoI decreases with the increase of the number of UAVs, because with the increase of the number of UAVs, the opportunity of collecting data packets in PoI is also more. In general, the performance of the greedy strategy and the random strategy is the worst, among the remaining algorithms, the performance of MADDPG, Edics, DRL-ASPT and DRL-DCEA presents a trend from weak to strong, in addition, the performance of DRL-DCEA on the average data flow age is better than that of other algorithms, for example, when the number of UAVs is 4, the average data flow age of DRL-DCEA is reduced by 34.43% than that of DRL-ASPT which is the second. In addition, the evaluation from the perspective of energy consumption, as shown in Figure 7 , also presents a similar law as the average data flow age. It should be pointed out that the DRL-DCEA proposed in the present application reduces the energy consumption by 11.40% than the second-ranked DRL-ASPT algorithm when the number of UAVs is 4. This shows that the DRL-DCEA algorithm introduces a multi-agent hierarchical reinforcement learning architecture, and obtains a better response performance in a dynamic scene, improves the data freshness index, and makes the decision of multi-agent more flexible and efficient.

[0142] Next, the influence of the sensing radius on the performance of the algorithm is evaluated when the sensing radius is 8, 10, 12 and 14, as shown in Figure 8The results show that with the gradual increase of the perception radius of the UAV, the four algorithms based on reinforcement learning (DRL-ECDA, DRL-ASPT, Edics and MADDPG) all show a trend of gradually decreasing the mean of data stream age. When the perception radius changes from 12 to 14, the downward trend of the mean of data stream age weakens, and even slightly increases. This is because by increasing the perception radius, the increased PoI distance from the UAV is also relatively far, which makes the data collection speed decrease, so that the UAV is difficult to collect more data. In summary, when the perception radius is 10, DRL-ECDA reduces the mean of data stream age by 29.74% compared with DRL-ASPT. On the other hand, with the change of the perception radius, the DRL-ECDA proposed in the application also shows the optimal performance in the energy consumption index among the compared algorithms, for example, when the perception radius is 8, DRL-ECDA reduces the energy consumption by 24.33% compared with the second DRL-ASPT.

[0143] In summary, according to the embodiments of the application, at least the following technical effects are achieved:

[0144] Firstly, the data stream age in the dynamic scene is established and defined as the data freshness index, which is used to evaluate the dynamic response performance of the UAV cooperative data collection. The data stream age considers not only the data amount in the data packet, but also the residence time of the data packet, which is different from the traditional information age, and can better reflect the data freshness of the data stream composed of multiple data packets.

[0145] Secondly, based on the existing technology of edge system assisted indirect task state information sharing architecture, the parameter estimation based on Poisson distribution random process and the sample exchange mechanism based on edge system are further proposed, which improves the accuracy of the dynamic parameter estimation of the PoI in the UAV mobile crowd sensing system.

[0146] Thirdly, in order to solve the problem of UAV data collection in dynamic scene, an algorithm based on multi-agent hierarchical deep reinforcement learning (hereinafter referred to as DRL-DCEA) is proposed. The macro and micro behaviors of the UAV are regarded as options and actions respectively. The data collection and access to the edge node under the time effectiveness requirement are regarded as two macro options. By introducing the option decision, the dynamic responsiveness of the UAV in distributed autonomous decision-making is improved.

[0147] Fourthly, the simulation experiment verifies that, in a typical dynamic mobile crowd sensing scene, compared with the baseline algorithms such as DRL-ASPT, Edics, MADDPG, greedy strategy and random strategy, the DRL-DCEA algorithm has superiority in evaluation indexes such as data stream age mean and energy consumption. For example, compared with the DRL-ASPT algorithm, the DRL-DCEA algorithm reduces the data stream age mean by 29.74% and the energy consumption by 24.33%.

[0148] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0149] The above is the introduction of the method embodiment, and the following will further illustrate the scheme of the present application through the device embodiment.

[0150] Figure 9 The structural diagram of a multi-unmanned aerial vehicle dynamic task intelligent allocation device based on an edge system and a data stream age provided for an embodiment of the present application is shown in FIG. 9, which can include: Figure 9

[0151] The setting module 910 is configured to set a working scene, the working scene including a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles.

[0152] The allocation module 920 is configured to, in the working scene, cause the unmanned aerial vehicles to go to the data points PoI to be collected to perform tasks, and cause the unmanned aerial vehicles to go to the edge nodes to share task state information, the two macro-initiative behaviors being regarded as options, micro-behaviors of the unmanned aerial vehicles in a time slot being regarded as actions, and an optimization target of a mobile crowd sensing task of the unmanned aerial vehicles being determined as minimizing a data stream age mean of all PoI in a data collection process under an energy constraint, each unmanned aerial vehicle being regarded as an agent, a multi-agent option-based partially observable Markov decision process model being established, and macro-option decision and micro-action decision being made through multi-agent hierarchical deep reinforcement learning.

[0153] It can be understood that, Figure 9 It can be understood that, Figure 1 ​The functions of each step in the multi-UAV dynamic task intelligent allocation method 100 shown and the corresponding technical effects can be achieved. For brevity, details are not repeated here.

[0154] Figure 10 A structural diagram of an exemplary electronic device capable of implementing an embodiment of the present application is shown in FIG. 10. Figure 10 As shown, the electronic device 1000 can include a computing unit 1001 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for operation of the electronic device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0155] A plurality of components in the electronic device 1000 are connected to the I / O interface 1005, including an input unit 1006 such as a keyboard, a mouse, etc., an output unit 1007 such as various types of displays, a speaker, etc., a storage unit 1008 such as a magnetic disk, an optical disk, etc., and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the method 100. For example, in some embodiments, the method 100 can be implemented as a computer program product including a computer program tangibly embodied in a computer-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method 100 described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the method 100 by any other appropriate means, such as by means of firmware.

[0157] The various embodiments described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0158] Program code to implement methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0159] In the context of the present application, a computer-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] It should be noted that the present application also provides a non-transitory computer readable storage medium having computer instructions stored therein, wherein the computer instructions are used to make a computer execute the method 100 and achieve the corresponding technical effects achieved by the embodiments of the present application executing the method. For brevity, the description will not be repeated here.

[0161] In addition, the present application also provides a computer program product, the computer program product comprises a computer program, the computer program realizes the method 100 when being executed by a processor.

[0162] It should be understood that the steps can be reordered, added, or deleted using the various forms of flowcharts shown above. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and the present application is not limited herein.

[0163] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for multi-UAV dynamic task intelligent allocation based on edge system and data flow age, characterized in that, The method comprises: setting a working scene, the working scene comprising a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles; under the working scene, the unmanned aerial vehicles go to the data points PoI to be collected to perform tasks, two macro-initiative behaviors of the unmanned aerial vehicles going to the edge nodes to share task state information are regarded as options, micro behaviors of the unmanned aerial vehicles in a time slot are regarded as actions, and an optimization target of a mobile crowd sensing task of the unmanned aerial vehicles is determined as minimizing a mean value of data flow ages of all the PoI in a data collection process under energy constraints, each unmanned aerial vehicle is regarded as an agent, a multi-agent partially observable Markov decision process model based on the options is established, macro option decision and micro action decision are made through multi-agent hierarchical deep reinforcement learning; the unmanned aerial vehicles collect data of the PoI, and the data collection comprises: The unmanned aerial vehicle transmits data to at most C PoIs in the same time slot, and in time slot t i , the unmanned aerial vehicle first accesses the drop box of all PoIs within its sensing radius ρ s to obtain the data stream ages of all PoIs, and all PoIs within the sensing radius ρ s are denoted as set If , the unmanned aerial vehicle u selects C PoIs with the largest data stream ages from for data collection; if , the unmanned aerial vehicle u selects all PoIs within for data collection, and all PoIs selected by the unmanned aerial vehicle u in time slot t i are denoted as set In time slot t i In the middle, in Approximate data transmission rate between the UAV u at position (x(p), y(p)) and PoIp at position (x(p), y(p)). The Shannon capacity, which takes into account the path loss exponent, is defined as follows: wherein, represents an approximate data transmission rate, B represents a total effective bandwidth, q0represents an average transmission power of the UAV, z 2 is a white noise power of a Gaussian distribution, g0represents a channel gain at a reference distance, a represents a path loss exponent, dist(u, p, i) represents an Euclidean distance between the UAV u and the PoI p in the time slot t i at the discrete grid scale l grid . For each piece of data to be collected Drones will Within a given time period, each data packet in the data stream is collected from its drop box in a first-in, first-out order, and the data volume of each data packet is recorded. After that, the data stream age of the PoI will be automatically updated.

2. The method of claim 1, wherein, a definition of the data flow age is as follows: Among them, A i (p) indicates that PoI p at t i The age of the data stream at any given moment It refers to t i The set of data packets contained within PoI p at time point δ m Let ι be the amount of data in data packet m. m The time when the data packet was generated.

3. The method of claim 1, wherein, a process of new data of the PoI is regarded as a Poisson process, and specifically: The newly added data in the PoI p is regarded as a random event, and the events are independent of each other, and the occurrence probability is equal, and the time slot t i The number of times of occurrence of the newly added data in the time slot t i In the PoI p, the number of times of occurrence of the newly added data is K i (p), and the probability of each event is λ(p), then K i (p) ~ Poisson(λ(p)); Assume that the data volume of each new data added for the same PoI p is a fixed value D(p), then the time slot t i The inner new data volume is δ i (p) = K i (p)D(p), these data are packaged as data packet m i , and added to the data stream of PoI p, the data stream will be placed in the throwing box of the PoI, waiting for the future access of the unmanned aerial vehicle to collect; the unmanned aerial vehicles obtain data flow information of the PoI p, the data amount of data packets generated in the Poisson process is regarded as a sample, and a Poisson distribution parameter λ(p) and a proportional amount D(p) of the PoI p are estimated.

4. The method of claim 1, wherein, a state space of the partially observable Markov decision process model comprises a position of the PoI, a data generation parameter, a residual data amount of the PoI in a current time slot, a data flow age in the current time slot, positions and energies of all the unmanned aerial vehicles in the current time slot; an observation space of the partially observable Markov decision process model comprises a position of the PoI, an estimated value of a data generation parameter, an estimated value of a residual data amount of the PoI in a current time slot, an estimated value of a data flow age in the current time slot, positions and energies of the unmanned aerial vehicles in the current time slot; an option space of the partially observable Markov decision process model represents macro-initiative behaviors of the agents, and an action space represents micro behaviors of the agents, specifically, a movement vector of the agents in the time slot; a reward function of the partially observable Markov decision process model comprises a reward of data collection of the agents, a penalty of energy consumption, a reward given by the edge system and a penalty of the agents moving out of a working area.

5. The method of claim 4, wherein, training of the multi-agent hierarchical deep reinforcement learning comprises: Before training starts, each agent is equipped with an option-value network for macro option decision making And a pair of performer-critic networks are configured inside each option Each network also has a corresponding target network; In addition, each agent is also equipped with a corresponding intra-option experience replay pool And an option experience replay pool The training consists of EP number of episodes, in each episode, the environment is set to the initial state s0, each agent obtains the initial observation of the environment Option termination identifier f for each agent u is set to True; There are M time slots in each round, and in each time slot, if the option termination identifier f u is True, the agent selects a macro option according to the following formula: After that, f u Set to False for f u If the value is False, the agent retains the original option and does not need to select an option. Next, each agent generates an action according to the policy function inside the option. Gaussian noise is then added to the action. Next, each agent interacts with the environment according to the following steps: The UAV follows the added noise moves; the unmanned aerial vehicles collect data of the PoI according to a data collection model; if the unmanned aerial vehicles are in a communication range of any edge node, the unmanned aerial vehicles exchange, synchronize and update task state information with the edge system; the unmanned aerial vehicles estimate parameters of the PoI according to a parameter estimation model; the PoI generates new data packets according to a data flow model, and updates a data flow age; Then, the environment state moves to s i+1 Each agent gets a local observation of the new state Then, the agent decides whether to stop the current option or not. If the option is stopped, set f u ← True. The stopping conditions for different options are as follows: ω = 1, i.e., when the option of the agent is to collect data, the stopping condition is i - i ω ≥ M c , i.e., the agent needs to re-decide the option after collecting M c time slot data continuously. ω = 2, i.e. when the option of the agent is to access the edge node, the stopping condition is that the agent successfully accesses the edge node, i.e. the option needs to be decided again; Afterwards, each agent stores experiences into an experience replay pool within each respective option Within each option, the value network parameters are updated using gradient descent and the policy network parameters are updated using gradient ascent, followed by updating the target network within each option using a soft update factor τ. If the agent u is in the current time slot t i An option termination occurs, the environment will give the agent the reward corresponding to the option The agent will store the experience Into the option experience replay Next, the agent samples a batch of experiences From the option experience replay And updates the network parameters according to the gradient descent method using the loss function shown in the following formula Finally, the soft update factor τ Ω Update target options - value target network parameters.

6. An apparatus for multi-UAV dynamic task intelligent allocation based on edge system and data flow age, characterized in that, the apparatus comprises: a setting module configured to set a working scene, the working scene comprising a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles; the setting module is configured to set a working scene, the working scene comprising a plurality of unmanned aerial vehicles, a plurality of data points PoI to be collected and a plurality of edge nodes, the plurality of edge nodes forming an edge system for indirect information sharing between the unmanned aerial vehicles; The allocation module is used to send the unmanned aerial vehicle to the data point to be collected to perform a task under the working scene, and two macro-initiative behaviors of the unmanned aerial vehicle going to the edge node to share task state information and the unmanned aerial vehicle in a time slot are regarded as options, micro behaviors of the unmanned aerial vehicle are regarded as actions, an optimization target of the mobile crowd sensing task of the unmanned aerial vehicle is determined as minimizing the average data flow age of all PoIs in the data collection process under the energy constraint, each unmanned aerial vehicle is regarded as an intelligent agent, a multi-agent option-based partially observable Markov decision process model is established, and thus macro option decision and micro action decision are made through multi-agent hierarchical deep reinforcement learning. The data collection of the unmanned aerial vehicle on the PoI includes: The UAV transmits data to at most C PoIs in the same time slot, and in time slot t i The UAV first accesses the drop box of all PoIs within its sensing radius p s to obtain the data stream age of all PoIs, and all PoIs within the sensing radius p s are denoted as set If the UAV u will select the C PoIs with the largest data stream age from to collect data; if the UAV u will select all PoIs within to collect data, and all PoIs selected by the UAV u in time slot t i are denoted as set In time slot t i The UAV u in position (x(p), y(p)) transmits data to the PoI in position (x(p), y(p)) at an approximate data transmission rate The UAV u in position (x(p), y(p)) transmits data to the PoI in position (x(p), y(p)) at an approximate data transmission rate defined by the Shannon capacity taking into account the path loss exponent, as shown in the following equation: wherein, denotes an approximate data transmission rate, B denotes a total effective bandwidth, q0denotes an average transmission power of the UAV, z 2 is a white noise power of a Gaussian distribution, g0denotes a channel gain at a reference distance, a denotes a path loss exponent, dist(u, p, i) denotes an Euclidean distance between the UAV u and the PoI p in the time slot t i at a discrete grid scale l grid . For each piece of data to be collected Drones will Within a given time period, each data packet in the data stream is collected from its drop box in a first-in, first-out order, and the data volume of each data packet is recorded. After that, the data stream age of the PoI will be automatically updated.

7. An electronic device, comprising: The electronic device includes: At least one processor; and The memory is in communication with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in any one of claims 1-5.

8. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method in any one of claims 1-5.

Citation Information

Patent Citations

  • Unmanned aerial vehicle task unloading method and system based on reinforcement learning in edge calculation

    CN111787509A

  • Unmanned aerial vehicle task allocation method based on edge system and reinforcement learning

    CN117707198A