A noma-based unmanned aerial vehicle assisted MEC resource optimization method

Through the NOMA-based drone-assisted MEC resource optimization method, the Poisson point process and deep reinforcement learning are used to optimize drone deployment and resource allocation, which solves the computing task processing problem of the drone-assisted MEC system when facing a large number of service demands, and improves user service quality and system capacity.

CN116528250BActive Publication Date: 2025-10-10NORTHEASTERN UNIV CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310399486.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-10-10
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

Existing drone-assisted MEC systems are unable to effectively handle computing tasks when faced with the service demands of a large number of smart devices, resulting in poor user service quality. Traditional multiple access technology leads to low spectrum utilization, making it difficult to meet the access needs of massive IoT devices.

Method used

A NOMA-based drone-assisted MEC resource optimization method is adopted. The user task distribution is simulated through the Poisson point process, drone pre-deployment is performed, a system model is built, and a deep reinforcement learning algorithm is used to optimize resource allocation to ensure service quality when the user location is unknown.

Benefits of technology

It improves user service quality, reduces drone deployment time, relieves the computing pressure of ground base stations, meets users' service needs and QoE requirements, and increases system capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116528250B_ABST
    Figure CN116528250B_ABST
Patent Text Reader

Abstract

The application provides a UAV-assisted MEC resource optimization method based on NOMA, relates to the fields of communication and reinforcement learning, and comprises the following steps: S1: obtaining the distribution of user computing tasks based on a Poisson point process; S2: pre-deploying a UAV according to the distribution of the user computing tasks; S3: constructing a UAV-assisted MEC system model based on NOMA; S4: obtaining an optimization problem based on the UAV-assisted MEC system model, wherein the optimization problem is to minimize the weighted sum of system energy consumption and task completion delay; and S5: solving the optimization problem by using a deep reinforcement learning algorithm to obtain an optimal resource allocation scheme. The application can be widely popularized in the fields of communication and reinforcement learning, can be applied to problem solving in a large-scale user massive data scene, and has a certain reference value for the research on future UAV-assisted MEC networks based on NOMA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of communications and reinforcement learning, and more specifically, to a NOMA-based drone-assisted MEC resource optimization method. Background Art

[0002] With the gradual maturity of 5G technology and the explosive growth of smart devices, a range of services and applications have emerged and are being widely used on these devices, leading to a surge in data traffic. However, the limited computing power of these devices cannot meet the needs of compute-intensive applications. Uploading tasks to a central cloud results in excessive latency, and existing cellular networks cannot meet the computing demands of high-traffic applications. Mobile Edge Computing (MEC) servers offer high computing power, allowing devices to offload tasks to nearby MEC servers, effectively alleviating congestion in ground infrastructure and improving user QoE. Drones can be deployed anywhere at any time and move in a controlled manner, especially over high-probability line-of-sight links. Compared to traditional fixed-location MEC, integrating MEC servers into drones offers low cost, high flexibility, and easy deployment, making them suitable for temporary events, emergencies, and on-demand services.

[0003] In existing drone-assisted MEC systems, the focus is on situations where the user's location is known, and the drone trajectory and resource allocation are jointly optimized based on the user's specific location. Most of them use traditional multiple access technology, and each channel can only provide services to one user.

[0004] Nowadays, smart devices are constantly developing towards miniaturization and portability, and user locations are uncertain. If only the situation where the user location is known is considered, it will be difficult to ensure that users receive good service quality when faced with a large number of service demands. Using traditional multiple access technology, each channel can only provide services to one user, and the spectrum utilization rate is low, which is difficult to meet the access needs of massive IoT devices in the future. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to propose a NOMA-based drone-assisted MEC resource optimization method to solve the problem that due to the huge number of smart devices and the explosive growth of data traffic, when faced with a large number of service demands from users, the ground base station may not be able to handle a large number of computing tasks.

[0006] The technical means adopted in the present invention are as follows:

[0007] A NOMA-based drone-assisted MEC resource optimization method includes the following steps:

[0008] S1: Obtain the distribution of user computing tasks based on the Poisson point process;

[0009] S2: Pre-deploy UAVs based on the distribution of the user computing tasks; the UAV pre-deployment includes determining the UAV coverage radius and the number and location of the UAVs;

[0010] S3: Construct a NOMA-based drone-assisted MEC system model that meets the user's task latency requirements and drone capacity requirements; the drone-assisted MEC system model includes a channel model and an offloading model.

[0011] S4: Obtaining an optimization problem based on the UAV-assisted MEC system model, wherein the optimization problem is to minimize the weighted sum of system energy consumption and task completion delay;

[0012] S5: Solve the optimization problem using a deep reinforcement learning algorithm to obtain an optimal resource allocation solution.

[0013] Furthermore, S1 specifically includes the following steps:

[0014] S1-1: Set a rectangular area of ​​size a×b and a user distribution density of λ / ㎡. The number and location of users are obtained by a two-dimensional Poisson point process.

[0015] Furthermore, S2 specifically includes the following steps:

[0016] S2-1: Determination of drone coverage radius:

[0017] Assume that the maximum number of users that the drone can accommodate is ρ max , the maximum coverage is r max , the coverage radius of the drone is expressed as the distance from the drone to the ρth max The horizontal distance from the user is expressed as the coverage radius from the drone to the ρth user. max The expectation of the horizontal distance between users is

[0018]

[0019] Among them, (ρ) 1 / ω It is Pochammer notation;

[0020] S2-2: Determination of the number and positions of drones: Considering densely packed drones in a rectangular venue, use regular hexagons inscribed in a circle to cover the rectangular area. The number of regular hexagons used is the number of drones, and the position of the center of the regular hexagon is the horizontal position of the drone.

[0021] Furthermore, S3 specifically includes the following steps:

[0022] S3-1: Establish a UAV-assisted mobile edge computing system, the mobile edge computing system comprising M ground users generating computing tasks, N UAVs and K ground base stations;

[0023] The set of ground users of the computing tasks is M={1, 2,..., M}, the set of UAVs is N={1, 2,..., N}, and the set of ground base stations is K={1, 2,..., K}. In a three-dimensional Euclidean coordinate system, the horizontal coordinates of the mth ground user are represented as u m =(x m , y m ), m∈M, and the three-dimensional coordinates of the nth UAV are represented as w n =(q n , h n ), wherein q n =(x n , y n ), n∈N;

[0024] S3-2: Establish a channel model: using a free space path loss model, the corresponding channel power gain mainly depends on the air-to-ground distance;

[0025] S3-3: Establish an offloading model, respectively calculate the computing delay and the energy consumed by user m in local computing, offloading to UAV n for computing, and offloading to a ground base station for computing, and then obtain the total energy consumption and the total time delay of the system.

[0026] Further, S3-2 specifically comprises the following steps:

[0027] Let the bandwidth between the UAV and the user be B uav , and the bandwidth between the ground base station and the user be B bs . Each UAV divides the bandwidth into L={1, 2,..., L} mutually orthogonal sub-channels. The ground user offloads the computing task to the UAV in a non-orthogonal multiple access manner or to the ground base station in a frequency division multiple access manner. The users of the same sub-channel of the UAV share the bandwidth, and the users of the same ground base station are equally allocated the bandwidth. If the mth ground user decides to offload its task to the l∈L sub-channel of the nth UAV, the data rate is represented as:

[0028]

[0029] If the mth ground user decides to offload its task to the kth ground base station, the data rate will be:

[0030]

[0031] wherein B krepresents the bandwidth allocated to users who offload the computational tasks to the ground base station k, r n,l represents the number of users in the lth sub-channel of the nth drone, Indicates the index of the user with the minimum received power in the i-th subchannel, i∈{1, 2, ..., r n,l}, R m,n,l represents the transmission rate of the lth subchannel when the mth ground user offloads the task to the nth UAV, represents the transmission power of the mth ground user transmitting data to the nth UAV, and the index of the user in the subchannel is i, pm represents user m transmitting data to the ground base station with a fixed transmission power, g m,k represents the channel power gain between ground user m and ground base station k, N0 represents the noise power between the user and the UAV, and N1 represents the noise power between the user and the ground base station.

[0032] Furthermore, S3-3 specifically includes the following steps:

[0033] Local computing: When user m's computing task is computed locally, the computing delay is:

[0034]

[0035] The energy consumed is:

[0036]

[0037] Among them, f m represents the local computing power of user m, κ m It represents the effective capacitance coefficient at user m affected by the chip architecture;

[0038] Offload the calculation to drone n; when the calculation task of user m is offloaded to drone n for calculation, the calculation delay is:

[0039]

[0040] The energy consumed is:

[0041]

[0042] Among them, f m,n represents the computing resources allocated by drone n to user m, κ n represents the effective capacitance coefficient of UAV n;

[0043] Offload the calculation to the ground base station; when the calculation task of user m is offloaded to the ground base station for calculation, the calculation delay is:

[0044]

[0045] The energy consumed is:

[0046]

[0047] Among them, f m,k represents the computing resources allocated by ground base station k to user m;

[0048] The total energy consumption of the system is:

[0049]

[0050] The total system delay is:

[0051]

[0052] Furthermore, S4 specifically includes the following steps:

[0053] set up represents the unloading decision of the UAV, a bs ={a m,1 ,a m,2 ,...a m,k} represents the offloading decision of the base station, represents the user's transmission power, f uav ={f 1,n ,f 2,n ,...,f m,n} represents the resource allocation of the UAV, f bs ={f 1,k ,f 2,k ,...,f m,k} represents the resource allocation of the base station;

[0054] The joint problem of minimizing system energy consumption and delay is formulated as:

[0055]

[0056] The constraints are:

[0057] C1:

[0058] C2:

[0059] C3:

[0060] C4:

[0061] C5:

[0062] C6:

[0063] C7:γ∈[0,1]

[0064] Among them, γ is the weight of energy consumption, δ e and δ t is the normalization factor.

[0065] Furthermore, S5 specifically includes the following steps:

[0066] S5-1: Define the four key elements of reinforcement learning: environment, state, action, and reward;

[0067] S5-2: Convert the problem of minimizing the weighted sum of system energy consumption and delay into the problem of maximizing the reward value in reinforcement learning.

[0068] Furthermore, in S5-1,

[0069] Environmental factors are MEC nodes in the NOMA-based drone-assisted MEC system. The environmental parameters that MEC nodes need to use when calculating reward values ​​and performing state transfers mainly include: the initial number of users M and their positions u m =(x m ,y m ), m∈Μ, the number of drones N and their locations w n =(q n ,h n ); Each user's local computing power f m , effective capacitance coefficient κ m ; Bandwidth B of each drone uav , the number of sub-channels L, the maximum number of users that can be accommodated ρ max , maximum computing resources; bandwidth B of ground base station bs , maximum computing resources The channel power gain g between the user m,k ; Channel power gain per unit distance β0, noise power density N0 between the user and the drone, noise power density N1 between the user and the ground base station;

[0070] The state element is a NOMA-based drone-assisted MEC system model. When the user computing task is offloaded in the current time slot, the state information S can be observed. (t) Consists of four parts Among them D m ={D1,D2,...,D m} represents the data size of each user's computing task, represents the maximum allowable delay of each user's computing task, represents the remaining resources of each drone, f bs Indicates the remaining resources of the ground base station;

[0071] The action element consists of the MEC node's decision to offload the user's computing tasks and allocate resources, which can be expressed as an action vector where a m ={0,1,2,...,n+1} represents the offloading decision of user m, i.e. local computing, offloading to ground base station, offloading to UAV n, l m ={1,2} represents the channel selection of the UAV, p m ={p1,p1,...,p m} represents the transmission power of user m, represents the computing resources allocated by the drone to user m, represents the computing resources allocated to user m by the ground base station;

[0072] Reward element; in the reinforcement learning model, the agent is in state S at each step of exploring the best unloading decision action towards the goal state (t) Next, perform a possible action A (t) After acting on the environment, you will receive an instant reward in the form of environmental feedback.

[0073] Compared with the prior art, the present invention has the following advantages:

[0074] The NOMA-based drone-assisted MEC resource optimization allocation method provided by the present invention takes into account the situation where the specific location of the user is unknown, deploys drones according to the uncertainty of user distribution, and ensures that all users can obtain better service quality; considers combining Poisson distribution with drone deployment, deploys drones in advance, reduces the time of drone deployment, and can directly perform computing services after users arrive, thereby improving service quality; the NOMA-based drone-assisted MEC resource optimization allocation research proposed by the present invention, drones provide services to users at the same time as ground base stations, which is used to improve system capacity and greatly alleviate the computing pressure of ground base stations to meet users' service needs and QoE requirements.

[0075] Based on the above reasons, the present invention can be widely promoted in the fields of communications and reinforcement learning, and can be applied to problem solving in scenarios with large-scale user and massive data. It has certain reference value for future research on NOMA-based drone-assisted MEC networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0077] Figure 1 This is a system model diagram of the present invention.

[0078] Figure 2 This is the user distribution diagram of the present invention.

[0079] Figure 3 This is a pre-deployment diagram of the UAV of the present invention.

[0080] Figure 4 This is a diagram of the convergence process of the PPO algorithm of the present invention.

[0081] Figure 5 This is a graph showing the changes in delay and energy consumption under different weight coefficients of the present invention.

[0082] Figure 6 This is a comparison chart of the weighted sum of system energy consumption and delay under different numbers of drones in the present invention.

[0083] Figure 7 This is a comparison chart of the weighted sum of system energy consumption and delay under different optimization objectives of the present invention.

[0084] Figure 8 This is a delay comparison diagram under different unloading conditions of the present invention.

[0085] Figure 9 This is a comparison chart of energy consumption under different unloading conditions of the present invention.

[0086] Figure 10 This is a weighted and comparative diagram under different unloading conditions of the present invention. DETAILED DESCRIPTION

[0087] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0088] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0089] To address the problem of ground base stations being unable to handle a large number of computing tasks due to the explosive growth of data traffic caused by the huge number of smart devices and the large number of user service demands, this paper studies a NOMA-based drone-assisted MEC system. This system optimizes resource allocation under a static distribution of user tasks and uses a Poisson point process to simulate the initial distribution of user tasks within the venue. First, drones are pre-deployed and the drone coverage radius is calculated based on the Poisson distribution. To meet the user distribution requirements under different scenarios, drones are densely distributed within the venue, and the number and location of drones are determined based on proven mathematical conclusions. When a user arrives, the drone and ground base station begin to provide services to the user and jointly optimize offloading decisions, power control, and resource allocation based on the initial distribution of user tasks. The optimization goal is to minimize the weighted sum of the total energy consumption and task completion delay of the entire system while meeting the service requirements and QoE of all users.

[0090] The present invention provides a NOMA-based UAV-assisted MEC resource optimization allocation method, comprising the following steps:

[0091] S1: Obtain the distribution of user computing tasks based on the Poisson point process;

[0092] S2: Based on the distribution of user tasks in S1, drone pre-deployment is performed, including determining the drone coverage radius and the number and location of drones;

[0093] S3: Based on S2, a NOMA-based drone-assisted MEC system model is constructed, including a channel model and an offloading model, and meeting the user task latency requirements and drone capacity requirements;

[0094] S4: According to the system model in S3, the optimization goal of the present invention is obtained, that is, minimizing the weighted sum of system energy consumption and task completion delay.

[0095] S5: Use deep reinforcement learning algorithm to solve the optimization problem in S4 and obtain the optimal resource allocation solution.

[0096] Furthermore, S1 specifically includes the following steps:

[0097] S1-1: Set a rectangular area of ​​size a×b, the user distribution density is λ / ㎡, and the number and location of users are obtained by a two-dimensional Poisson point process.

[0098] Furthermore, step S2 specifically includes the following steps:

[0099] S2-1: Determination of the UAV coverage radius: In a uniform ω-dimensional Poisson point process with intensity λ, in the area A∈R ω The probability that there are ν nodes in is given by:

[0100]

[0101] where μ(A) is the standard Lebesgue measure of A, which can be used to compute the distance to the ρth neighbor in a straightforward way.

[0102] Theorem 1 (Euclidean distance to the ρth neighbor):

[0103] At R with strength λ ω In the Poisson point process, the distance R between the point and the ρth neighboring point ρ Distributed according to the generalized Gamma distribution:

[0104]

[0105] Among them, c ω r ω is the volume of the ω-dimensional sphere with radius r, Γ(ρ)=(ρ-1)!.

[0106] Assume that the maximum number of users that the drone can accommodate is ρ max , the maximum coverage is r max , the coverage radius of the drone is expressed as the distance from the drone to the ρth max The horizontal distance of each user is not a constant, and each coverage radius cannot be determined during the initial deployment. Therefore, the coverage radius is expressed as the distance from the drone to the ρth user. max The expectation of the horizontal distance between users is

[0107]

[0108] Among them, (ρ) 1 / ω is Pochammer notation.

[0109] S2-2: Determining the number and positions of drones: Consider densely covering a rectangular venue with drones. Since the coverage range of drones is a circle, the problem becomes a problem of circle covering the rectangle. However, the rectangle cannot be completely covered by the circle. To completely cover the rectangular area, the total capacity of drones must be redundant. Therefore, how to use the least number of circles to fully cover the rectangular area is the next problem to be solved.

[0110] To cover a rectangular area with a circle, you first need to cover the rectangular area with a polygon. According to the minimum covering circle model, the circle obtained by limiting the radius of the circle with a regular polygon can cover the entire area with the least number of circles.

[0111] According to the theorem, when an area is covered by a regular hexagon inscribed in a circle, the common area between the intersection of two adjacent circles is the smallest. Therefore, fewer circles are required in the same rectangle, which is more in line with the condition of the minimum number of circles required in the problem.

[0112] Therefore, we use regular hexagons inscribed in the circle to cover the rectangular area. The number of regular hexagons used is the number of drones, and the position of the center of the regular hexagon is the horizontal position of the drone.

[0113] Furthermore, step S3 specifically includes the following steps:

[0114] S3-1: Consider a drone-assisted mobile edge computing system, which consists of three parts: M ground users that generate computing tasks, N drones, and K ground base stations. The set of ground users, drones, and ground base stations are represented as M = {1, 2, ..., M}, N = {1, 2, ..., N}, and K = {1, 2, ..., K}, respectively. Using a three-dimensional Euclidean coordinate system, the horizontal coordinate of the mth ground user is represented as u m =(x m ,y m ),m∈Μ, the three-dimensional coordinates of the n-th UAV are represented by w n =(q n ,h n ), where q n =(x n ,y n ),n∈Ν。

[0115] S3-2: Channel Model: In this system model, the UAV LoS link communication is dominant, assuming no obstacles in the air. A free-space path loss model is used, and the corresponding channel power gain depends primarily on the air-to-ground distance.

[0116] The distance from the mth ground user to the nth UAV is:

[0117]

[0118] The channel power gain between ground user m and UAV n is:

[0119]

[0120] Among them, β0 represents the channel power gain when the reference distance is 1m, and the UAV only provides services to ground users within its coverage area.

[0121] Assume that the bandwidth between the drone and the user is B uav , the bandwidth between the ground base station and the user is B bs Each UAV divides its bandwidth into L = {1, 2, ..., L} mutually orthogonal sub-channels. Ground users offload computing tasks to UAVs using non-orthogonal multiple access (NOMA) or to ground base stations using frequency division multiple access. UAV users on the same sub-channel share the bandwidth, while users on the same ground base station share their bandwidth equally. If the mth ground user decides to offload its task to the l∈Lth sub-channel of the nth UAV, the data rate is expressed as:

[0122]

[0123] If the mth ground user decides to offload its task to the kth ground base station, the data rate will be:

[0124]

[0125] Among them, B k represents the bandwidth allocated to users who offload the computational tasks to the ground base station k, r n,l represents the number of users in the lth sub-channel of the nth drone, Indicates the index of the user with the minimum received power in the i-th subchannel, i∈{1,2,...,r n,l}. R m,n,l represents the transmission rate of the lth subchannel when the mth ground user offloads the task to the nth UAV, represents the transmission power of the mth ground user transmitting data to the nth UAV, and the index of the user in the subchannel is i, p m Indicates that user m transmits data to the ground base station with a fixed transmission power, g m,k represents the channel power gain between ground user m and ground base station k, N0 represents the noise power between the user and the UAV, and N1 represents the noise power between the user and the ground base station.

[0126] S3-3: Offloading model: The computing task of ground user m is represented as Among them D mdenotes the size of the computing task, denotes the maximum allowable delay.

[0127] Introducing the offloading decision variable a m,0 ,a m,l ,a m,k ∈{0,1},a m,0 = 1 indicates that user m computes locally, otherwise 0; denotes that user m offloads to the l-th sub-channel of its nearest n-th drone, otherwise 0;a m,k = 1 indicates that user m offloads to the ground base station k, otherwise 0. Thus, for the entire computing task of user m, we have

[0128]

[0129] The k-th ground base station sub-channel bandwidth is

[0130]

[0131] (1) Local computing

[0132] When the computing task of user m is computed locally, its computing delay is:

[0133]

[0134] The consumed energy is:

[0135]

[0136] where f m denotes the local computing capability of user m, κ m denotes the effective capacitance coefficient of user m affected by the chip architecture.

[0137] (2) Computing on the drone n

[0138] When the computing task of user m is computed on the drone n, its computing delay is:

[0139]

[0140] The consumed energy is:

[0141]

[0142] where f m,n denotes the computing resource allocated by drone n to user m, κ n denotes the effective capacitance coefficient of drone n. Thus, for all users connected to drone n, under the condition that the maximum computing resource of the drone is , it should be satisfied that:

[0143]

[0144] (3) Offloading calculation to the ground base station

[0145] When the computing task of user m is offloaded to the ground base station for calculation, the calculation delay is:

[0146]

[0147] The energy consumed is:

[0148]

[0149] Among them, f m,k represents the computing resources allocated by ground base station k to user m. For all users connected to ground base station k, the maximum computing resources at the ground base station are Under the conditions, the following conditions should be met:

[0150]

[0151] Therefore, to complete the calculation task I m The delay is expressed as:

[0152]

[0153] To complete the calculation task I m The energy consumed is:

[0154]

[0155] The total energy consumption of the system is:

[0156]

[0157] The total system delay is:

[0158]

[0159] In this model, the energy consumption of the drone hovering is not considered because in this scheme, the position change of the drone is not considered. The drone is always hovering in the air, and the hovering energy consumption is constant, which has no effect on the change of system energy consumption.

[0160] Furthermore, step S4 specifically includes the following steps:

[0161] S4-1: Set represents the unloading decision of the UAV, a bs ={a m,1 ,a m,2 ,...a m,k} represents the offloading decision of the base station, denotes the transmit power of the user, f uav 1,n 2,n m,n denotes the resource allocation of the UAV, f bs 1,k 2,k m,k denotes the resource allocation of the base station;

[0162] The joint problem of minimizing system energy consumption and latency is formulated as:

[0163] (γ is the weight of energy consumption, δ e and δ t are normalization factors to make energy consumption and latency reach similar scales.)

[0164] P:

[0165] C1:

[0166] C2:

[0167] C3:

[0168] C4:

[0169] C5:

[0170] C6:

[0171] C7:γ∈[0,1]

[0172] Further, step S5 specifically comprises the following steps:

[0173] S5-1: Based on the optimization problem of solving the optimal offloading and resource allocation scheme to minimize the system energy consumption and latency weighted sum established in step S4, the present application adopts a proximal policy optimization algorithm (PPO) to solve the optimization problem, which is based on the Actor-critic framework and can solve reinforcement learning problems in both discrete action space and continuous action space. Since the model of the present application has the characteristics of high-dimensional discrete action space, it is suitable for solving with this algorithm.

[0174] The present application first defines in detail the four key elements of environment, state, action and reward in reinforcement learning, and then converts the problem of solving the minimum of the system energy consumption and latency weighted sum into the problem of solving the maximum of the reward value in reinforcement learning. ​​​​​​

[0175] Environment: The environment here specifically refers to the MEC nodes in the NOMA-based drone-assisted MEC system of the present invention. The environmental parameters that the MEC nodes need to use when calculating the reward value and performing state transfer mainly include: the initial number of users M and their positions u m =(x m ,y m ), m∈Μ, the number of drones N and their locations w n =(q n ,h n ); Each user's local computing power f m , effective capacitance coefficient κ m ; Bandwidth B of each drone uav , the number of sub-channels L, the maximum number of users that can be accommodated ρ max , maximum computing resources; bandwidth B of ground base station bs , maximum computing resources The channel power gain g between the user m,k ; Channel power gain per unit distance β0, noise power density N0 between the user and the drone, and noise power density N1 between the user and the ground base station.

[0176] State: Based on the NOMA-based drone-assisted MEC system model constructed by the present invention, when the user computing task is unloaded in the current time slot, the state information S can be observed. (t) Consists of four parts Among them D m ={D1,D2,...,D m} represents the data size of each user's computing task, represents the maximum allowable delay of each user's computing task, represents the remaining resources of each drone, f bs Indicates the remaining resources of the ground base station.

[0177] Action: In the system of the present invention, at time slot t, the action of the MEC node consists of the offloading of the user's computing tasks and the resource allocation decision of the MEC node, which can be expressed as an action vector where a m ={0,1,2,...,n+1} represents the offloading decision of user m, i.e. local computing, offloading to ground base station, offloading to UAV n, l m ={1,2} represents the channel selection of the UAV, p m ={p1,p1,...,p m} represents the transmission power of user m, represents the computing resources allocated by the drone to user m, represents the computing resources allocated by the ground base station to user m.

[0178] Reward: In the reinforcement learning model, the agent is rewarded in state S for each step in exploring the best unloading decision action towards the goal state. (t) Next, perform a possible action A (t) After acting on the environment, you will get an instantaneous reward R of environmental feedback (t) The goal of reinforcement learning is to obtain the maximum cumulative reward. Therefore, the reward function designed by the present invention should be negatively correlated with the objective function of the optimization problem. Therefore, the present invention defines the reward R (t) for:

[0179]

[0180] In reinforcement learning, the agent explores and selects the unloading action A with the goal of maximizing the cumulative reward. (t) Therefore, the reward value corresponding to the optimal decision action (which minimizes the weighted sum of energy consumption and delay) in any state is the highest, realizing the transformation of the problem.

[0181] Figure 1 This is a system model diagram of the NOMA-based drone-assisted MEC resource optimization allocation method described in the present invention. The drone-assisted mobile edge computing system consists of three parts: M ground users that generate computing tasks, N drones, and K ground base stations. The set of ground users, drones, and ground base stations are represented as M = {1, 2, ..., M}, N = {1, 2, ..., N}, and K = {1, 2, ..., K} respectively. Using a three-dimensional Euclidean coordinate system, the horizontal coordinate of the mth ground user is represented as u m =(x m ,y m ),m∈Μ, the three-dimensional coordinates of the n-th UAV are represented by w n =(q n ,h n ), where q n =(x n ,y n ), n∈N. Users offload computing tasks to drones using non-orthogonal multiple access or to ground base stations using frequency division multiple access. UAV users on the same subchannel share bandwidth, while users on the same ground base station share bandwidth evenly.

[0182] Figure 2 This is the user distribution diagram in the NOMA-based drone-assisted MEC resource optimization allocation method described in the present invention.

[0183] Figure 3This is the drone pre-deployment diagram in the NOMA-based drone-assisted MEC resource optimization allocation method of the present invention. The drone coverage radius is obtained based on the user distribution density and the number of users that the drone can accommodate, and the number and position of drones are obtained by covering the rectangular area with a regular hexagon inscribed in a circle. Figure 3 As shown in the figure, the red rectangle represents the rectangular area, and the center of the circle represents the horizontal position of the drone.

[0184] Figure 4 This figure shows the convergence process of the PPO algorithm in the NOMA-based UAV-assisted MEC resource optimization allocation method described in this invention. The horizontal axis represents the number of training episodes, and the vertical axis represents the average reward value in a training episode. The figure clearly shows that when the task data volume is 20kb, the average reward value obtained by the PPO algorithm gradually increases with the increase in the number of training episodes. After 300,000 rounds of iterative training, the PPO algorithm has gradually converged to the optimal solution.

[0185] Figure 5 This is a graph showing the changes in latency and energy consumption under different weight coefficients in the NOMA-based drone-assisted MEC resource optimization allocation method described in the present invention. Figure 5 The weight coefficient in represents the weight of energy consumption. As can be seen, as the weight coefficient increases, energy consumption decreases while latency increases. In practice, the weight coefficient can be adjusted based on the level of concern for energy consumption and latency, as well as the actual energy consumption or latency requirements. For example, if a computing task requires a latency of less than 8.5 seconds, the figure shows that an energy consumption weight of less than 0.5 should be selected.

[0186] Figure 6 This figure compares the weighted sum of system energy consumption and latency for different numbers of drones in the NOMA-based drone-assisted MEC resource optimization allocation method described in this invention. It can be seen that when the number of drones is four (i.e., the drone deployment method designed in this invention), the weighted sum of system latency and energy consumption is minimized. When the number of drones is less than four, the system performance is relatively low due to the small number of drones, and the weighted sum is higher. When the number of drones is greater than four, as the number of drones increases, the hovering energy consumption increases, resulting in a rapid increase in the weighted sum and a significant decrease in system performance. This verifies the rationality of the drone deployment scheme.

[0187] Figure 7This figure compares the weighted sum of system energy consumption and latency under different optimization objectives in the NOMA-based drone-assisted MEC resource optimization and allocation method described in the present invention. It can be seen that compared with optimizing only latency or optimizing only energy consumption, the solution of the present invention (i.e., optimizing the weighted sum of energy consumption and latency) has better system performance. Furthermore, as the amount of mission data continues to increase, the difference in the weighted sum of latency and energy consumption of the solution of the present invention compared to the other two methods becomes increasingly larger, indicating that the system performance is improving.

[0188] Figure 8-10 This is a comparison chart of latency, energy consumption, and weighted sum under different unloading conditions in the NOMA-based drone-assisted MEC resource optimization allocation method described in the present invention.

[0189] Figure 8 、 Figure 9 The figure compares latency and energy consumption under three different offloading conditions: local computing, local computing combined with a ground base station, and local computing combined with a ground base station and a drone. The figure shows that as the amount of task data increases, latency and energy consumption also increase, and the performance gap between the proposed solution and the other two solutions also widens.

[0190] Figure 10 The figure shows a comparison of the delay and weighted sum of energy consumption under the three methods when the energy consumption weight is 0.7. When the task data volume is 45kb, the performance of this solution is improved by 30% compared with the local computing system and by 20% compared with the local + ground base station.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A NOMA-based UAV-assisted MEC resource optimization method, characterized in that: The steps include: S1: Obtain the distribution of user computing tasks based on the Poisson point process; S2: Pre-deploy UAVs based on the distribution of the user computing tasks; the UAV pre-deployment includes determining the UAV coverage radius and the number and location of the UAVs; S2-1: Determination of drone coverage radius: Assume that the maximum number of users that the drone can accommodate is , the maximum coverage is , the coverage radius of the drone is expressed as the distance from the drone to the first The horizontal distance from the user is expressed as the coverage radius from the drone to the The expectation of the horizontal distance between users is in, is the Pochammer notation, λ is the intensity, is the dimension of the sphere; S2-2: Determining the number and positions of drones: Considering densely packed drones in a rectangular venue, use regular hexagons inscribed in a circle to cover the rectangular area. The number of regular hexagons used is the number of drones, and the position of the center of the regular hexagon is the horizontal position of the drones; S3: Build a NOMA-based drone-assisted MEC system model that meets the user task latency requirements and drone capacity requirements; the drone-assisted MEC system model includes a channel model and an offloading model; S3-1: Establish a UAV-assisted mobile edge computing system, which includes M ground users that generate computing tasks, N UAVs, and K ground base stations; The set of ground users of the computing task is , the set of drones is , the set of ground base stations is , using the three-dimensional Euclidean coordinate system, the horizontal coordinate of the mth ground user is expressed as , the three-dimensional coordinates of the nth UAV are expressed as ,in ; S3-2: Establish channel model: Using the free space path loss model, the corresponding channel power gain mainly depends on the air-to-ground distance; each UAV divides the bandwidth into Mutually orthogonal sub-channels are used. Ground users offload computing tasks to drones using non-orthogonal multiple access or to ground base stations using frequency division multiple access. UAV users on the same sub-channel share bandwidth, while users on the same ground base station share bandwidth evenly. S3-3: Establish an uninstallation model and calculate the user Compute locally and offload to drones The computation delay and energy consumed are calculated and offloaded to the ground base station, and then the total energy consumption and total system delay are obtained; S4: Obtaining an optimization problem based on the UAV-assisted MEC system model, wherein the optimization problem is to minimize the weighted sum of system energy consumption and task completion delay; S5: Solve the optimization problem using a deep reinforcement learning algorithm to obtain an optimal resource allocation solution.

2. The NOMA-based drone-assisted MEC resource optimization method according to claim 1 is characterized in that: S1 specifically includes the following steps: S1-1: Set a rectangular area of ​​size a×b and a user distribution density of λ / ㎡. The number and location of users are obtained by a two-dimensional Poisson point process.

3. The NOMA-based drone-assisted MEC resource optimization method according to claim 1 is characterized in that: S3-2 specifically includes the following steps: Assume that the bandwidth between the drone and the user is , the bandwidth between the ground base station and the user is , each drone divides the bandwidth into The ground users offload the computing tasks to the UAVs in a non-orthogonal multiple access manner or to the ground base station in a frequency division multiple access manner. The users of the same sub-channel of the UAVs share the bandwidth, and the users of the same ground base station share their bandwidth evenly. A ground user decides to offload its tasks to The first drone sub-channels, the data rate is expressed as: in, g m,n Indicates ground users With the n Channel power gain between UAVs; If the A ground user decides to offload its tasks to If there are terrestrial base stations, the data rate will be: in, Indicates offloading computing tasks to ground base stations The bandwidth allocated to users on Indicates the The first drone The number of users in a sub-channel, Indicates the number of The index of the user with the minimum received power, , Indicates the The first ground user offloads the task to the The first drone The transmission rate of the sub-channels, Indicates the A ground user to The transmission power of the drone to transmit data, and the index of the user in the sub-channel is , Represents a user Transmit data to the ground base station with fixed transmission power, Indicates ground users With ground base station k The channel power gain between represents the noise power between the user and the drone, Represents the noise power between the user and the ground base station.

4. The NOMA-based drone-assisted MEC resource optimization method according to claim 3 is characterized in that: S3-3 specifically includes the following steps: Local computing; when the user When the computation task is computed locally, the computation delay is: The energy consumed is: in, Represents a user local computing power, Represents a user The effective capacitance coefficient is affected by the chip architecture; Unloading to drone Calculate on; when the user Offloading computing tasks to drones When the calculation is performed on , the calculation delay is: in, Indicates the size of the computing task; The energy consumed is: in, Indicates drone Assign to user computing resources, Indicates drone The effective capacitance coefficient; Offload the calculation to the ground base station; when the user When the computation task is offloaded to the ground base station for computation, the computation delay is: The energy consumed is: in, Indicates ground base station Assign to user computing resources; The total energy consumption of the system is: The total system delay is: 。 5. The NOMA-based drone-assisted MEC resource optimization method according to claim 4 is characterized in that: S4 specifically includes the following steps: set up represents the unloading decision of the drone, represents the offloading decision of the base station, Indicates the user's transmit power, represents the resource allocation of the drone, Indicates the resource allocation of the base station; The joint problem of minimizing system energy consumption and delay is formulated as: The constraints are: in, is the weight of energy consumption, and is the normalization factor.

6. The NOMA-based drone-assisted MEC resource optimization method according to claim 1 is characterized in that: S5 specifically includes the following steps: S5-1: Define the four key elements of reinforcement learning: environment, state, action, and reward; S5-2: Convert the problem of minimizing the weighted sum of system energy consumption and delay into the problem of maximizing the reward value in reinforcement learning.

7. The NOMA-based drone-assisted MEC resource optimization method according to claim 6 is characterized in that: In S5-1, Environmental factors are MEC nodes in the NOMA-based drone-assisted MEC system. The environmental parameters that MEC nodes need to use when calculating reward values ​​and performing state transfers mainly include: the initial number of users M and their locations , the number of drones N and their locations ; Local computing power of each user , effective capacitance coefficient ; Bandwidth per drone , number of sub-channels L, maximum number of users that can be accommodated , maximum computing resources; bandwidth of ground base stations , maximum computing resources , and the channel power gain between users ; Channel power gain per unit distance , the noise power density between the user and the drone , the noise power density between the user and the ground base station ; The state element is a NOMA-based drone-assisted MEC system model. When the user computing task is offloaded in the current time slot, the state information can be observed. Consists of four parts ,in Indicates the data size of each user's computing task, represents the maximum allowable delay of each user's computing task, represents the remaining resources of each drone, Indicates the remaining resources of the ground base station; The action element consists of the MEC node's offloading of the user's computing tasks and resource allocation decisions, which can be expressed as an action vector ,in represents the offloading decision of user m, i.e. local computing, offloading to ground base station, offloading to drone , Indicates the channel selection of the drone, Represents a user The transmission power, Indicates that the drone is assigned to the user computing resources, Indicates that the ground base station is assigned to the user computing resources; Reward factor: In the reinforcement learning model, the agent is rewarded at each step in the state when exploring the best unloading decision action towards the goal state. Next, perform a possible action After acting on the environment, you will receive an instant reward in the form of environmental feedback.

Citation Information

Patent Citations

  • Unmanned aerial vehicle auxiliary edge unloading decision-making method based on deep reinforcement learning

    CN115309467A