A flight control and computing offloading method and system for multi-uav mobile edge computing
By optimizing the flight trajectory and computational offloading decisions of UAVs, and by using deep reinforcement learning algorithms to optimize the multi-UAV mobile edge computing system, the problems of system energy consumption and computational task latency have been solved, and efficient computing services have been achieved in field environments and emergency rescue scenarios.
Patent Information
- Application Number
- CN202211119514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-09-14
AI Technical Summary
In scenarios such as wilderness environments, emergency rescue and disaster relief, and military applications, multi-UAV mobile edge computing systems face challenges such as limited power supply and strict requirements for the completion time of computing tasks. Existing solutions are unable to effectively reduce system energy consumption and computing task completion latency.
By jointly optimizing the UAV flight trajectory, computational task offloading decisions, and offloading ratios, deep reinforcement learning algorithms are used to construct a state space, action space, and reward function suitable for multi-UAV assisted MEC systems, thereby optimizing UAV trajectories and computational offloading strategies to minimize the total system cost.
It achieves a significant reduction in system energy consumption and computation task completion time, provides flexible and low-cost computing services, and meets the stringent requirements of computing tasks.
Smart Images

Figure CN115454527B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of mobile communication, in particular, to a flight control and computing offloading method and system for multi-unmanned aerial vehicle mobile edge computing. BACKGROUND
[0002] The rapid development of mobile communication networks and mobile Internet of Things promotes the unprecedented growth of intelligent user terminals, and also provides a strong platform for many new intelligent applications. Various new applications have emerged, such as face recognition, virtual reality games, remote medical treatment, etc. However, these applications are all computationally intensive and delay sensitive, and usually require high computing power. The limited battery power and low computing capacity of users make it difficult for them to handle these applications. In order to solve this conflict, mobile edge computing (MEC) has gradually been valued. Compared with cloud computing, mobile edge computing servers are deployed at the edge of infrastructure-based mobile communication networks, which can provide task offloading services for users and improve user experience. However, the mobile edge computing system deployed in the ground mobile communication network is limited by the fixed location deployment, and lacks sufficient flexibility, especially in infrastructure-limited scenarios such as field environment, emergency rescue, military field, etc. It cannot be quickly and flexibly deployed. Therefore, a flexible and low-cost mobile edge computing system deployed in an unmanned aerial vehicle communication network, i.e. an unmanned aerial vehicle mobile edge computing system, is proposed, which is suitable for various application scenarios.
[0003] Regarding the unmanned aerial vehicle mobile edge computing system, most existing solutions consider single unmanned aerial vehicle deployment. However, in scenarios such as virtual reality tourism parks or large intelligent factories, multiple unmanned aerial vehicles need to be deployed to achieve greater network coverage and computing service guarantee. In a multi-unmanned aerial vehicle mobile edge computing system, many solutions provide methods to minimize the energy consumption of unmanned aerial vehicles. However, in typical unmanned aerial vehicle mobile edge computing system application scenarios such as field environment, emergency rescue, military field, etc., mobile terminals also face the problem of limited power supply, and the energy consumption of mobile terminals is also crucial. At the same time, there are strict requirements for the completion time of computing tasks in the above typical application scenarios.
[0004] Therefore, how to provide a method for significantly reducing system energy consumption and computing task completion time by reasonably planning the trajectory of unmanned aerial vehicles and formulating computing offloading decisions is a problem that needs to be solved by those skilled in the art. SUMMARY
[0005] The application provides a flight control and computing offloading method for multi-unmanned aerial vehicle mobile edge, which reduces system energy consumption and total time of completing computing tasks.
[0006] To achieve the above purpose, the application provides a flight control and computing offloading method for multi-unmanned aerial vehicle mobile edge, which specifically comprises the following steps: obtaining initial information; constructing a solution model according to the obtained initial information; simulating and solving the energy consumption and delay problem by the solution model to obtain the best unmanned aerial vehicle trajectory and the offloading decision and computing task offloading ratio of the user terminal; and performing actions corresponding to the best unmanned aerial vehicle trajectory and the offloading decision and computing task offloading ratio of the user terminal.
[0007] As above, wherein the obtaining of the initial information comprises obtaining the number and position of user terminals in the service area, obtaining the initial number of unmanned aerial vehicles, and using a feedforward neural network to predict the computing task amount of all user terminals in the next flight period according to the obtained number and position of user terminals.
[0008] As above, wherein the constructing of the solution model according to the obtained initial information specifically comprises the following sub-steps: constructing a channel model; and constructing the solution model according to the obtained initial information in response to the completion of the construction of the channel model.
[0009] As above, wherein the user terminal set is defined as The unmanned aerial vehicle set is The horizontal position coordinate of the unmanned aerial vehicle m in the nth time slot is L m,n =[x m,n ,y m,n ], The construction of the channel model comprises: determining the path loss and the probability that the link can be regarded as a line-of-sight link According to the path loss and the probability that the link can be regarded as a line-of-sight link The wireless channel gain is determined.
[0010] As above, wherein the path loss is specifically represented as:
[0011]
[0012] The probability that the link can be regarded as a line-of-sight link is specifically represented as:
[0013]
[0014] where the position coordinates of the user terminal are denoted as w s = [x s ,y s ]. The horizontal position coordinates of the UAV m at the nth time slot are denoted as L m,n = [x m,n ,y m,n ], denote the horizontal distance between the ground user terminal s and the UAV m at the nth time slot, f c is the carrier frequency, d o = max{294.05 log 10 H - 432.94, 18}, pi = 233.98 log 10 H - 0.95.
[0015] As above, where the wireless channel gain g s,m,n between the ground user terminal s and the UAV m at the nth time slot is denoted as:
[0016]
[0017] where, and denote the path loss of LoS and NLoS links, respectively, denotes the probability that the wireless channel is a LoS link, denotes the probability that the wireless channel is a NLoS link, and
[0018] As above, where the construction of the decision model includes: defining the communication frequency band bandwidth between the UAV and the user terminal as W, different UAVs use different frequency bands, so there is no interference between the UAVs. Therefore, at the nth time slot, the rate r s,m,n at which the ground user terminal s offloads tasks to the UAV m is denoted as:
[0019]
[0020] where B s,m,n is the bandwidth allocated by the UAV m to the user terminal s at time slot n, P s is the transmit power of the user terminal s, and σ 2 is the noise power.
[0021] As above, where the construction of the decision model further includes: defining the offloading decision variable z s,m,n , using z s,m,n = 1 to represent the device m selected by the s-th user terminal to perform the computing task at time slot n, and m = 0 represents selecting the terminal itself as the computing device, i.e. no offloading, and m ≠ 0 represents offloading the computing task to the UAV m. The set of devices for computing task offloading is defined as Additionally define a variable p s,m,n ∈[0,1] to represent the proportion of user terminals s to unmanned aerial vehicles m to unload tasks; when m = 0, p s,m,n = 0.
[0022] As above, wherein the solving model performs simulation solving of the energy consumption and delay problem, and obtains the optimal unmanned aerial vehicle trajectory and the user terminal unloading decision and the computing task unloading proportion, and specifically includes the following sub-steps: initializing the new actor network, the old actor network and the critic network parameters; in response to the initialization of the actor network and the critic network parameters, it is judged whether the maximum number of cycles is reached; if the maximum number of cycles is reached, it is directly ended, and the network parameters corresponding to the new actor network, the old actor network and the critic network are output, and the optimal unmanned aerial vehicle trajectory and the user terminal unloading decision and the computing task unloading proportion are obtained.
[0023] A multi-unmanned aerial vehicle mobile edge computing flight control and computing offloading system for any of the above-mentioned methods, specifically comprising: an information acquisition module, an optimization problem modeling module, a simulation solving and actual execution module; the information acquisition module is used to acquire initial information; the optimization problem modeling module is used to construct a solving model according to the acquired initial information; the simulation solving module is used to solve the simulation solving of the energy consumption and delay problem by the solving model, and obtain the optimal unmanned aerial vehicle trajectory and the user terminal unloading decision and the computing task unloading proportion; the actual execution module is used to execute the action corresponding to the optimal unmanned aerial vehicle trajectory and the user terminal unloading decision and the computing task unloading proportion.
[0024] The present application has the following beneficial effects:
[0025] The present application is based on a deep reinforcement learning algorithm, and the present application proposes a state space, an action space and a reward function suitable for a multi-unmanned aerial vehicle assisted MEC system. The flight action (flight speed and angle) that the unmanned aerial vehicle should take in each time slot and the unloading decision and unloading proportion of the computing task are obtained, and the total system cost (system energy consumption and delay) is minimized. BRIEF DESCRIPTION OF DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art according to these drawings.
[0027] Figure 1 It is a model diagram of the multi-unmanned aerial vehicle mobile edge computing flight control and computing offloading system provided by the embodiments of the present application.
[0028] Figure 2 is an internal structure diagram of a flight control and computing offloading system for multi-unmanned aerial vehicle mobile edge computing provided by an embodiment of the present application;
[0029] Figure 3 is a flowchart of a method of flight control and computing offloading for multi-unmanned aerial vehicle mobile edge computing provided by an embodiment of the present application;
[0030] Figure 4 is still another flowchart of a method of flight control and computing offloading for multi-unmanned aerial vehicle mobile edge computing provided by an embodiment of the present application. DETAILED DESCRIPTION
[0031] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0032] In a multi-unmanned aerial vehicle assisted edge computing cellular network, a base station first collects the number and position information of user terminals according to the actual scene, and predicts the computing task amount in the near future according to the actual computing task amount of user terminals returned by unmanned aerial vehicles. According to the prediction data, an algorithm model is trained. After the training is completed, the optimal trajectory and computing offloading strategy of the unmanned aerial vehicle in each time slot are calculated through the model, and then the communication bandwidth resource allocation ratio and the unmanned aerial vehicle computing power resource allocation ratio are obtained. Finally, the base station sends data to the unmanned aerial vehicle for actual execution, thereby providing services for ground user terminals.
[0033] Scene assumption:
[0034] First, consider an MEC system containing S ground user terminals and M unmanned aerial vehicle assistance, and the system model diagram is as shown in Figure 1 The user terminal set and the unmanned aerial vehicle set are defined as and The flight cycle of each unmanned aerial vehicle is set to T, and the entire flight cycle is discretized into N time slots. The position coordinates of the user terminal are represented as w s =[x s ,y s ]. The horizontal position coordinates of the unmanned aerial vehicle m in the nth time slot are L m,n =[x m,n ,y m,n ], The horizontal distance between the ground user terminal s and the unmanned aerial vehicle m in the nth time slot is The range of the task performed by the multi-unmanned aerial vehicle is represented as an area of x max ×ymax The service area, of which x max ,y max These represent the length and width of the flight mission area, respectively. The distance between the UAV m and m′ is defined as R. m,m′,n And introduce a minimum safety distance R min R m,m′,n ≥R min Assume the wireless channel is quasi-static, meaning the channel conditions remain unchanged within a single time slot, and the uplink and downlink use different frequency bands. Considering the limitations of different terrains, the path loss between the UAV and the user terminal is determined by the probabilities of line-of-sight (LoS) links and non-line-of-sight (NLoS) links.
[0035] Example 1
[0036] like Figure 2 The image shows a flight control and computational offloading system for the mobile edge of multiple unmanned aerial vehicles (UAVs) provided in this application embodiment. It includes a ground base station and UAVs, as well as various modules set in the ground base station and UAVs. Specifically, each module is an information acquisition module 210, an optimization problem modeling module 220, a simulation solution module 230, and an actual execution module 240.
[0037] The optimization problem modeling module and simulation solution module are deployed in the ground base station, while the information acquisition module and the actual execution module coexist in both the ground base station and the UAV. The ground base station uses the acquired information and stored algorithms for offline training, while the UAV receives the results calculated by the base station and executes them online.
[0038] After the service area is planned, the information acquisition module 210 in the ground base station first obtains the number and location of user terminals in the service area, then uses a feedforward neural network to predict the computational workload of all user terminals in the next flight cycle, initializes the number of UAVs, and outputs all the information obtained above to the optimization problem modeling module 220.
[0039] After receiving information from the information acquisition module 210, the optimization problem modeling module 220 determines the wireless channel gain based on the user terminal location coordinates and the UAV location coordinates, and constructs a channel model based on the wireless channel gain. Based on the user terminal location coordinates, the UAV location coordinates, the computational workload for predicting the next flight cycle of all user terminals, and the channel model, it constructs a decision and computation model. The decision and computation model is used as a solution model to solve the energy consumption and time delay problem, and the solution model is output to the simulation solution module 230.
[0040] After the channel model, decision and calculation model construction are completed, the simulation solving module 230 performs simulation, solves the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio, and outputs the obtained solution to the actual execution module 240.
[0041] The actual execution module 240 of the ground base station sends data such as the best flight trajectory, the user terminal unloading decision and the calculation task unloading ratio to the actual execution module 240 of the unmanned aerial vehicle. The unmanned aerial vehicle takes corresponding actions according to the data to provide services for the ground user terminal, and uses the information acquisition module 210 in the unmanned aerial vehicle to collect the actual calculation task amount of the user terminal in the current flight period, thereby providing training data for the information acquisition module 210 in the ground base station to predict the calculation task amount of the user terminal in the next flight period.
[0042] Embodiment Two
[0043] As Figure 3 shown, a flight control and calculation unloading method of a multi-unmanned aerial vehicle mobile edge computing system is provided, which specifically includes the following sub-steps:
[0044] Step S310: Obtain initial information.
[0045] Specifically, obtaining initial information includes obtaining the number and position of user terminals in the service area, obtaining the initial number of unmanned aerial vehicles, and using a feedforward neural network to predict the calculation task amount of all user terminals in the next flight period according to the obtained number and position of user terminals.
[0046] Step S320: Construct a solving model according to the obtained initial information.
[0047] Specifically, step S320 includes the following sub-steps:
[0048] Step S3201: Construct a channel model.
[0049] The wireless channel gain is determined according to the user terminal position coordinates and the unmanned aerial vehicle position coordinates, and the channel model is constructed after the wireless channel gain is determined.
[0050] Path loss and the probability of being considered as a line-of-sight link are respectively represented as:
[0051]
[0052]
[0053] wherein H is the flight height of the unmanned aerial vehicle, f c is the carrier frequency, and d o= max{294.05 log 10 H - 432.94, 18}, pi = 233.98 log 10 H - 0.95.
[0054] Therefore, the wireless channel gain between the ground user terminal s and the UAV m in the nth time slot is represented as:
[0055]
[0056] wherein, and denote the path loss of LoS and NLoS links, respectively, denotes the probability that the wireless channel is a LoS link, denotes the probability that it is a NLoS link, and
[0057] Step S3202: In response to completing the construction of the channel model, the construction of the solving model is performed according to the obtained initial information.
[0058] Specifically, the decision model and the calculation model are constructed according to the user terminal position coordinates, the UAV position coordinates, the predicted calculation task amount of all user terminals in the future short time, and the channel model.
[0059] wherein the construction of the decision model comprises: defining the unloading decision variable z s,m,n , using z s,m,n = 1 to represent the device m selected by the s-th user terminal to perform the calculation task in the nth time slot, when m = 0, it represents selecting the terminal itself as the calculation device, i.e. no unloading, and when m ≠ 0, it represents that the calculation task is unloaded to the UAV m. The set of devices for calculation task unloading is defined as In addition, the variable p s,m,n ∈ [0, 1] is defined to represent the proportion of the task unloaded by the user terminal s to the UAV m. Obviously, when m = 0, p s,m,n = 0. It is set that each user terminal will unload the task to at most one UAV in a single time slot, and therefore there is the following constraint, which completes the construction of the decision model:
[0060]
[0061] In response to completing the construction of the decision model and the channel model, the communication frequency band bandwidth between the UAV and the user terminal is defined as W, different UAVs use different frequency bands, and therefore there is no interference between the UAVs. Therefore, in the nth time slot, the rate r s,m,n of the task unloaded by the ground user terminal s to the UAV m is represented as:
[0062]
[0063] where B s,m,n is the bandwidth allocated to user terminal s by UAV m in time slot n, P s is the transmit power of user terminal s, σ 2 is the noise power.
[0064] Define the number of CPU cycles consumed per bit of task as C f , the amount of computing task generated by the s-th user terminal in time slot n as D s,n , then the time for the ground user terminal s to transmit computing data to the UAV m in the n-th time slot, that is, the computing task offloading completion time is expressed as:
[0065]
[0066] In order to unify the variable expression, when m = 0,
[0067] The present application balances the computing task offloading completion time of user terminals by allocating different bandwidths to different user terminals Try to keep the completion time of all user terminals accessing the UAV m close to each other, and ensure the fairness between users. Through The expression of bandwidth allocation B s,m,n is obtained:
[0068]
[0069] Further, the local completion time of the computing task of the ground user terminal s in time slot n is expressed as:
[0070]
[0071] where f s represents the computing power of user terminal s, and the unit is cycle number / second, which is generally the frequency of a micro computing processor.
[0072] The expected computing completion time of the computing task offloaded by the UAV m for the user terminal s in time slot n is expressed as:
[0073]
[0074] where f m is the computing power allocated by the UAV m to each user terminal, and the unit is cycle number / second. In order to unify the variable expression, when m = 0,
[0075] The multiple computing tasks offloaded to multiple ground terminals of a UAV are computed in parallel by multiple virtual machines created by the UAV computing processor. Since there is mutual interference due to sharing of common computing resources in the same physical computing processor, the expected computing completion time of the computing task offloaded by user terminal s to UAV m in time slot n is expressed as:
[0076]
[0077] where τ is the degradation coefficient caused by I / O interference between virtual machines, which can generally be set to 0.2, and c is the number of virtual machines. In order to unify the variable expression, when m = 0,
[0078] Each user terminal has different expected computing completion time due to different computing task amounts. After the UAV completes the edge computing service for the user terminal, it will immediately remove the virtual machine and release the computing resources. Therefore, the number of virtual machines c actually existing in the UAV m is constantly changing, i.e., increasing with time, within a single time slot. First, define a set Sort each element in ascending order, and define the sorted set as and additionally define Therefore, the actual computing completion time of the computing task offloaded by user terminal s to UAV m in time slot n is expressed as:
[0079]
[0080] where, in order to unify the variable expression, when m = 0,
[0081] In response to the determination of formula 11, since the amount of UAV computing result data is small, the time and energy consumption required by the downlink can be ignored. It is set that the local execution of the computing task and the offloading of the computing task can be performed simultaneously, and the total time of the computing task of user terminal s in time slot n is expressed as:
[0082]
[0083] It is set that the computing task of user terminal s in time slot n must be completed within the maximum limit time t max Therefore, there is the following constraint:
[0084]
[0085] The computing energy consumption of user terminal s locally in time slot n is expressed as:
[0086]
[0087] wherein, is the effective switched capacitance of the user terminal.
[0088] At time slot n, the user terminal s unloads the communication energy consumption of the computing task to the UAV m, i.e., the communication energy consumption of the computing task is unloaded is expressed as:
[0089]
[0090] At time slot n, the UAV m provides the computing energy consumption required for the computing service for the user terminal s is expressed as:
[0091]
[0092] wherein, φ m is the effective switched capacitance of the UAV.
[0093] Considering the high mobility of the UAV, the propulsion energy consumption of the UAV m within the time slot n is expressed as:
[0094]
[0095] wherein, κ1, κ2, κ3 are respectively parameters related to the hardware of the UAV, v m,n is the flight speed of the UAV m at time slot n, in order to unify the variable expression, when m = 0,
[0096] So far, the computer model is completed.
[0097] Further, the local computing energy consumption and the unloaded communication energy consumption of the ground user terminal, the computing energy consumption of the UAV, the flight energy consumption of the UAV, and the total time of completing the computing task are jointly considered, and the total cost of the system is defined as the weighted sum thereof. In particular, for the total time of completing the computing task, it is not necessary to minimize it, but only to ensure that it does not exceed the maximum limit time. By reasonably planning the flight speed and flight angle of the UAV, the unloading strategy (unloading decision and unloading proportion) of the user terminal, the total cost of the system is minimized.
[0098] wherein the joint optimization problem is in the following form:
[0099]
[0100] wherein, In order to ensure the optimizability, the parameters ω s , ω m , ω f and ωt respectively represent the weight of user terminal energy consumption, the weight of UAV computing energy consumption, the weight of UAV flight energy consumption and the weight of total time of computing task completion, to ensure that the four are in the same order of magnitude.
[0101] Step S330: solving the model to simulate the energy consumption delay problem, and obtaining the best UAV trajectory and the user terminal unloading decision and computing task unloading ratio.
[0102] Based on the existing proximal policy optimization algorithm, a joint optimization method of computing task unloading and UAV trajectory control is proposed to obtain the best UAV trajectory and the user terminal unloading decision and computing task unloading ratio.
[0103] The proximal policy optimization algorithm mainly consists of multiple loops. First, define the actor network (new actor network and old actor network) and critic network, and complete initialization (the two actor network initialization parameters are the same). Then set the maximum number of training rounds, and complete the initialization of the agent state and experience buffer. After all the initialization is completed, a single round of training will be entered, with a time step as the unit, and the time step size is equal to the length of a single time slot. The agent obtains the action distribution to be taken in each state through the new actor network, and samples to obtain the actual action. Then input the action into the environment to get the reward and enter the next state, and store the state, action and reward pair in the buffer. Repeat the cycle. If a violation occurs, proceed directly to the next round of training. If a violation occurs, proceed directly to the next round of training. Otherwise, when the time step reaches the set batch size or the last step, first synchronize the new actor network parameters to the old actor network, and then use the state value predicted by the critic network to backtrack the state value stored in the buffer:
[0104]
[0105] wherein G i is the value of the i-th state in the experience buffer, γ is the decay factor, generally set to 0.9-1, r j+1 is the reward obtained at the j+1 time step, e n+1 is the n+1 state, Q(·|θ Q ) is the critic network. When n=N, Q(e n+1 |θ Q )=0.
[0106] wherein the collision between UAV m and m' is defined as R m,m′,n <R min or the UAV m exceeds the flight task area, i.e. x m,n >x max or y m,n >y maxFor violating the operation.
[0107] wherein in each time step, the state, action and reward are defined as follows:
[0108] The state space of the agent is defined as respectively represent the horizontal position coordinate [x m,n ,y m,n ] of the UAV m at time slot n, the horizontal distance between the UAV m and the user terminal s at time slot n, and the computing task amount of the user terminal s at time slot n.
[0109] The action space is defined as respectively represent the flight angle and flight speed (i.e. the UAV trajectory) taken by the UAV m at time slot n, the computing task offloading decision and the offloading ratio selected by the user terminal s at time slot n. In processing the offloading decision part, we use the indicator function to continuous the discrete decision for the convenience of neural network processing, as follows:
[0110]
[0111] wherein, is a continuous set defined according to the number of UAVs, and when , z s,m=0,n = 1.
[0112] The reward r n obtained by the agent in each time step is defined as:
[0113]
[0114] wherein ξ is a penalty term introduced to prevent the UAV from violating the operation. When there is no violation of the operation, the approximate average reward obtained in each time slot is set as that is, if the order of magnitude of -r n is k, then when any UAV collides or exceeds the flight task area, and thereby correcting its wrong action and ending the round.
[0115] As shown in FIG. 4, the overall algorithm specifically includes the following sub-steps: Figure 4 Step S3401: initialize the actor network and critic network parameters.
[0116] wherein the actor network includes a new actor network and an old actor network, and the two actor networks have the same initialized parameters.
[0117]
[0118] The actor network and the critic network are networks provided in the proximal policy optimization algorithm, and are also parameters commonly used in the proximal policy optimization algorithm, and details are not described herein.
[0119] Step S3402: In response to initializing the actor network and the critic network parameters, it is determined whether the maximum number of cycles is reached.
[0120] The maximum number of cycles is a maximum number of rounds of training set in advance.
[0121] If the maximum number of cycles is reached, it is directly ended, and the network parameters corresponding to the new actor network, the old actor network and the critic network are output, and the best unmanned aerial vehicle trajectory and the user terminal unloading decision and the computing task unloading ratio are obtained.
[0122] If the maximum number of cycles is not reached, step S1403 is performed.
[0123] Step S3403: The agent state and the experience buffer are initialized.
[0124] The agent and the experience buffer are included in the algorithm, and specific meanings are not described herein.
[0125] Step S3404: In response to completing the initialization of the agent state and the experience buffer, it is determined whether the maximum step length is exceeded or the action is ended.
[0126] If the maximum step length is exceeded or the action is ended, step S3402 is returned, otherwise step S3405 is performed.
[0127] Step S3405: The action is obtained according to the actor network sampling.
[0128] Step S3406: In response to sampling the action, the action is performed, the reward is obtained, and the next state is entered.
[0129] Step S3407: The obtained state, action and reward are stored in the experience buffer.
[0130] Step S3408: In response to storing the obtained state, action and reward in the experience buffer, it is determined whether a batch size is full or whether the maximum step length is reached.
[0131] If the batch size is full or the maximum step length is reached, step S3409 is performed. Otherwise, step S3404 is performed.
[0132] Step S3409: The old actor network parameters are updated.
[0133] Step S3410: in response to updating the old actor network parameters, updating the new actor network parameters and the critic network parameters for a specified number of times.
[0134] The specified number of times is a pre-set number of times, which can be adjusted according to actual conditions, and the specific value is not limited herein.
[0135] The critic network parameters are updated using the time difference error
[0136]
[0137] B t is the batch size.
[0138] The probability distribution p i and p′ i of all actions a i in the buffer are obtained using the new actor network and the old actor network respectively, and the advantage value A i of each state is obtained using the state value obtained by the critic network, A i =G i -Q(e i |θ Q ), and the parameters of the new actor network are updated.
[0139]
[0140] Wherein, ε is the clipping ratio, clip represents outputting 1-ε when p i / p′ i is less than 1-ε, outputting 1+ε when it is greater than 1+ε, and not processing when it is between the two.
[0141] The calculation method of G i is referred to formula 19.
[0142] Step S3411: in response to updating the new actor network parameters and the critic network parameters, emptying the experience buffer, and returning to execute step S3404.
[0143] Wherein, until the maximum number of loops is reached, the action sampled by the actor network is output, that is, the best UAV trajectory and the user terminal unloading decision and the calculation task unloading ratio are obtained.
[0144] Step S340: performing the action corresponding to the best UAV trajectory and the user terminal unloading decision and the calculation task unloading ratio.
[0145] The best flight trajectory, user terminal unloading decision and calculation task unloading ratio and other data are sent to the actual execution module of the UAV, and the UAV takes corresponding actions according to the data to provide services for the ground user terminal.
[0146] Further, the information acquisition unit in the unmanned aerial vehicle collects the actual calculation task amount of the current flight period of the user terminal, and provides training data for the information acquisition unit in the base station to predict the calculation task amount of the next flight period of the user terminal.
[0147] The present application has the following beneficial effects:
[0148] The present application is based on a deep reinforcement learning algorithm, and the present application proposes a state space, an action space and a reward function suitable for a multi-unmanned aerial vehicle assisted MEC system. The flight action (flight speed and angle) of the unmanned aerial vehicle in each time slot, the offloading decision and the offloading ratio of the calculation task are obtained, and the total system cost (system energy consumption and time delay) is minimized.
[0149] Although the examples described with reference to the present application are described, they are only for the purpose of explanation and not limitation of the present application, and changes, additions and / or deletions to the embodiments can be made without departing from the scope of the present application.
[0150] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for flight control and computing offloading of multi-UAV mobile edge computing, characterized in that, Specifically comprising the following steps: Obtaining initial information; According to the obtained initial information, the construction of the solving model is carried out; The simulation solving of the energy consumption and time delay problem is carried out by the solving model, and the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio are obtained; Performing actions corresponding to the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio; According to the obtained initial information, the construction of the solving model includes the following sub-steps: Constructing a channel model; In response to completing the construction of the channel model, the construction of the solving model is carried out according to the obtained initial information; The constructing the channel model comprises: determining a wireless channel gain g according to the user terminal position coordinate and the UAV position coordinate s,m,n is represented as: wherein, and PLLoSand PLNLoSdenote the path loss for LoS and NLoS links, respectively, PLLoSdenotes the path loss for LoS links, PLNLoSdenotes the path loss for NLoS links, and According to the obtained initial information, the construction of the solving model comprises: constructing a decision model; the construction of the decision model comprises: defining an unloading decision variable z s,m,n , wherein z s,m,n =1 represents that the s th user terminal selects the device m to execute the computing task in the time slot n, when m=0, it represents selecting the terminal itself as the computing device, that is, no unloading, and when m≠0, it represents that the computing task is unloaded to the unmanned aerial vehicle m; a set of computing task unloading devices is defined as In addition, a variable p s,m,n ∈[0,1] is defined to represent the proportion of the user terminal s unloading the task to the unmanned aerial vehicle m; when m=0, p s,m,n =0; it is set that each user terminal unloads the task to at most one unmanned aerial vehicle in a single time slot, and therefore there is a constraint as follows, and the construction of the decision model is completed: The joint optimization problem is as follows: Wherein, the user terminal set is The UAV set is The horizontal position coordinate of the UAV m in the nth time slot is L m,n = [x m,n ,y m,n ], α m,n represents the flight angle of the UAV m in the time slot n, v m,n represents the flight speed of the UAV m in the time slot n, represents the computing energy consumption required by the UAV m to provide computing services for the user terminal s in the time slot n, represents the propulsion energy consumption of the UAV m in the time slot n, represents the energy consumption of the user terminal s unloading computing tasks to the UAV m in the time slot n, represents the computing energy consumption of the user terminal s locally in the time slot n, represents the total computing task completion time of the user terminal s in the time slot n, and it is set that the computing task of the user terminal s must be completed within the maximum limit time t max , the parameter ω s represents the user terminal energy consumption weight, ω m represents the UAV computing energy consumption weight, ω f represents the weight of the UAV flight energy consumption, ω t represents the weight of the total computing task completion time; The simulation solving of the energy consumption and time delay problem is carried out by the solving model, and the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio are obtained, including the following sub-steps: Initialize actor network and critic network parameters; In response to initializing the actor network and the critic network parameters, it is judged whether the maximum number of cycles is reached; If the maximum number of cycles is reached, directly end, output the network parameters corresponding to the new actor network, the old actor network and the critic network, and obtain the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio. 2.The method of claim 1, wherein, Obtaining initial information includes obtaining the number and location of user terminals in the service area, obtaining the initial number of unmanned aerial vehicles, and using a feedforward neural network to predict the calculation task amount of all user terminals in the next flight cycle according to the obtained number and location of user terminals. 3.The method of claim 2, wherein, According to the obtained initial information, the construction of the solving model includes the following sub-steps: Constructing a channel model; In response to completing the construction of the channel model, the construction of the solving model is carried out according to the obtained initial information. 4.The method of claim 3, wherein, Then, the channel model is constructed, including: Determining path loss And probability of being considered as line-of-sight link According to path loss And probability of being considered as line-of-sight link Determining a wireless channel gain. 5.The method of claim 4, wherein, Path loss Specifically represented as: Probability of being visual as line-of-sight Specifically represented as: wherein the position coordinates of the user terminal are denoted as w s = [x s ,y s ], the horizontal position coordinates of the unmanned aerial vehicle m in the nth time slot are denoted as L m,n = [x m,n ,y m,n ], denote the horizontal distance between the ground user terminal s and the unmanned aerial vehicle m in the nth time slot, H is the flight height of the unmanned aerial vehicle, f c is the carrier frequency, d o = max{294.05 log 10 H - 432.94, 18}, and p1 = 233.98 log 10 H - 0.
95. 6.The method of claim 5, wherein, The wireless channel gain g between the ground user terminal s and the unmanned aerial vehicle m in the nth time slot s,m,n is represented as: wherein, and denote the path loss for LoS and NLoS links, respectively, denotes the probability that the wireless channel is a LoS link, denotes the probability that it is a NLoS link, and 7. The method of claim 6, wherein the method further comprises: The construction of the decision model comprises: defining the communication frequency band bandwidth between the unmanned aerial vehicle and the user terminal as W, different unmanned aerial vehicles use different frequency bands, and therefore there is no interference between the unmanned aerial vehicles; therefore, in the nth time slot, the rate r at which the ground user terminal s unloads tasks to the unmanned aerial vehicle m s,m,n is represented as: where B s,m,n is the bandwidth allocated to the user terminal s by the drone m in time slot n, P s is the transmit power of the user terminal s, σ 2 is the noise power, g s,m,n denotes the wireless channel gain between the ground user terminal s and the drone m in the nth time slot. 8.The method of claim 7, wherein, The construction of the decision model also comprises: defining an offloading decision variable z s,m,n , wherein z s,m,n =1 represents the device m selected by the s-th user terminal to perform the computing task at time slot n, wherein m=0 represents selecting the user terminal itself as the computing device, i.e., no offloading, and wherein m≠0 represents offloading the computing task to the drone m; a set of devices for computing task offloading is defined as In addition, a variable p s,m,n ∈[0,1] is defined to represent the proportion of the user terminal s offloading the task to the drone m; when m=0, p s,m,n =0. 9.The method of claim 7, wherein, The simulation solving of the energy consumption and time delay problem is carried out by the solving model, and the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio are obtained, including the following sub-steps: Initialize new actor network, old actor network and critic network parameters; In response to initializing the actor network and the critic network parameters, it is judged whether the maximum number of cycles is reached; If the maximum number of cycles is reached, directly end, output the network parameters corresponding to the new actor network, the old actor network and the critic network, and obtain the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio. 10.A multi-UAV mobile edge computing flight control and computing offloading system for performing the method of any one of claims 1-9. Specifically comprising: Information acquisition module, optimization problem modeling module, simulation solving and actual execution module; The information acquisition module is used to obtain initial information; The optimization problem modeling module is used to construct the solving model according to the obtained initial information; The simulation solving module is used to carry out the simulation solving of the energy consumption and time delay problem by the solving model, and the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio are obtained; The actual execution module is used to perform actions corresponding to the best unmanned aerial vehicle trajectory and the user terminal unloading decision and calculation task unloading ratio.