A method and system for resource allocation and task offloading in NOMA-assisted multi-UAV edge computing
By using NOMA-assisted multi-UAV edge computing system and combining it with MSAC algorithm to optimize UAV trajectory and resource allocation, the problem of low OMA spectrum efficiency is solved, resulting in lower total maximum latency and better service support, making it suitable for computationally intensive and latency-sensitive tasks.
Patent Information
- Application Number
- CN202411224585.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-09-03
AI Technical Summary
In existing drone edge computing scenarios, traditional orthogonal multiple access (OMA) technology faces problems such as low spectrum efficiency and limited user access, and fails to effectively utilize the flexibility of drones. Especially when communication resources are scarce and user equipment computing power is limited, it cannot meet the needs of computationally intensive and latency-sensitive services.
A NOMA-assisted multi-UAV edge computing system is adopted. By deploying an offload architecture, defining user device clusters, establishing a network communication model, and combining the MSAC algorithm of multi-agent reinforcement learning, the system optimizes UAV trajectories, computing resource allocation, and task offload ratio to achieve joint optimization.
It effectively reduces the total maximum latency of the system, improves spectrum efficiency, supports computationally intensive and latency-sensitive services, and makes up for the shortcomings of traditional OMA scenarios, especially when communication resources are limited or ground base stations are damaged.
Smart Images

Figure CN119300047B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of UAV edge computing technology, specifically to a method and system for resource allocation and task offloading in NOMA-assisted multi-UAV edge computing. Background Technology
[0002] In recent years, the rapid development of IoT applications has created a contradiction between the limited computing power of user devices and the surge in demand. To address the challenges posed by computationally intensive tasks, combining drones with mobile edge computing can enable timely task processing, better adapt to complex task requirements, and provide user devices with more flexible and higher-quality services.
[0003] Edge computing has the advantages of low latency, decentralization, and high security and reliability. Combining drones with edge computing technology can be used in fields such as intelligent transportation and environmental monitoring. These services generate a lot of data. By using edge computing technology, the huge computing tasks of user devices can be offloaded to drones equipped with edge servers, relieving the computing pressure on user devices and reducing latency.
[0004] Existing research on UAV edge computing scenarios largely employs traditional Orthogonal Multiple Access (OMA) technology. However, given the scarcity of communication resources, spectral efficiency remains a challenge. Due to the precious and scarce nature of communication resources, traditional OMA suffers from low spectral efficiency and limited user access. Non-orthogonal Multiple Access (NOMA), by having users share the same time-frequency and other resource blocks, achieves efficient resource utilization and improves spectral efficiency, making it effectively applicable to the massive data access of the future Internet of Things (IoT). Furthermore, existing research on multi-UAV edge computing scenarios does not address the joint optimization of offloading strategies, resource allocation, and UAV 3D flight trajectories, failing to fully utilize the flexibility of UAVs. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a resource allocation and task offloading method and system for NOMA-assisted multi-UAV edge computing, which minimizes the total maximum latency of all user devices in the NOMA-assisted multi-UAV edge computing system.
[0006] The technical solution adopted in this embodiment of the invention is: a resource allocation and task offloading method for NOMA-assisted multi-UAV edge computing, comprising the following steps:
[0007] S10. Deploy an offloading architecture in the NOMA-assisted multi-drone edge computing scenario, define relevant parameters of drones and user equipment, cluster user equipment, and establish a network communication model.
[0008] S20. Based on the network communication model, establish a computing model and determine a joint optimization strategy for UAV trajectory, computing resource allocation, and task offloading ratio;
[0009] S30. Based on the computational model and joint optimization strategy, determine the objective function and constraints of the optimization problem;
[0010] S40. Transform the optimization problem into a Markov decision process, and determine the environment, state, action space and reward of the Markov decision process; normalize the state, and use the MSAC algorithm to solve the Markov decision process, and realize the joint optimization of task unloading ratio, resource allocation and multi-UAV trajectory based on the solution results;
[0011] The MSAC algorithm is a multi-agent SAC algorithm based on the single-agent SAC algorithm. In the MSAC algorithm, before the state is input into the neural network, the state information of each UAV is extracted and formed into a standard array, so that multiple UAVs can share the neural network. The location of the UAV that is serving and connected to the neural network, the channel gain of the user equipment it serves, the task size, and the remaining power of the UAV are input into the designated neuron.
[0012] In another embodiment of the present invention, the technical solution adopted is: a resource allocation and task offloading system for NOMA-assisted multi-UAV edge computing, comprising:
[0013] The architecture deployment unit deploys the offloading architecture in the NOMA-assisted multi-drone edge computing scenario, defines the relevant parameters of drones and user equipment, clusters user equipment, and establishes a network communication model.
[0014] The strategy determination unit is used to establish a computational model based on the network communication model and determine the joint optimization strategy for UAV trajectory, computational resource allocation, and task offloading ratio.
[0015] The optimization problem determination unit is used to determine the objective function and constraints of the optimization problem based on the computational model and joint optimization strategy.
[0016] The transformation and solution unit is used to transform the optimization problem into a Markov decision process, determine the environment, state, action space and reward of the Markov decision process; normalize the state, use the MSAC algorithm to solve the Markov decision process, and realize the joint optimization of task unloading ratio, resource allocation and multi-UAV trajectory based on the solution results.
[0017] The MSAC algorithm is a multi-agent SAC algorithm based on the single-agent SAC algorithm. In the MSAC algorithm, before the state is input into the neural network, the state information of each UAV is extracted and formed into a standard array, so that multiple UAVs can share the neural network. The location of the UAV that is serving and connected to the neural network, the channel gain of the user equipment it serves, the task size, and the remaining power of the UAV are input into the designated neuron.
[0018] Compared with the prior art, the main advantages of the present invention include:
[0019] 1. In this invention, each UAV is equipped with an edge server to provide communication and computing services to user devices. By combining NOMA technology with a multi-UAV edge computing system, reasonable clustering of users is achieved by considering the payload capacity of each UAV, and the offloading strategy, computing resource allocation, and UAV 3D flight trajectory are jointly optimized. This provides better support for computationally intensive and latency-sensitive services, making up for the shortcomings of traditional OMA scenarios where communication resources are limited, user device computing capabilities are limited, or ground base stations are damaged and unable to complete communication and computing services under natural disasters, thereby reducing the total maximum latency of all user tasks in the system.
[0020] 2. This invention constructs a NOMA-assisted multi-UAV communication and computing network system model, defines relevant parameters for UAVs and user equipment, clusters user equipment, and establishes a network communication model. Based on the network communication model, a computing model is established, and joint optimization strategies for UAV trajectory, computing resource allocation, and task offloading are determined. According to the system model, the objective function and constraints of the optimization problem are determined. The optimization problem is transformed into a Markov decision process, and the MSAC algorithm in multi-agent reinforcement learning is used to jointly optimize the task offloading ratio, resource allocation, and UAV 3D trajectory for solution. By comparing the performance of the MSAC algorithm and benchmark algorithms, it is concluded that this invention can effectively reduce the total maximum latency and obtain a suitable offloading strategy and reasonable resource allocation.
[0021] 3. This invention is effectively applicable to computationally intensive and latency-sensitive services, making up for the shortcomings of traditional OMA scenarios where communication resources are limited, user equipment computing power is limited, or ground base stations are damaged and unable to complete communication computing services under natural disasters. It has practical significance for reducing the total maximum latency of the system in multi-UAV communication scenarios. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the NOMA-assisted multi-UAV edge computing resource allocation and offloading method provided in an embodiment of the present invention;
[0023] Figure 2This is a schematic diagram of the NOMA-assisted multi-UAV edge computing system architecture environment provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram illustrating the principle of the MSAC algorithm provided in this embodiment of the invention and its differences from the SAC algorithm;
[0025] Figure 4 This is a schematic diagram of the MSAC algorithm structure for implementing multivariable optimization in a NOMA-assisted multi-UAV edge computing system provided by an embodiment of the present invention;
[0026] Figure 5 This is the convergence graph of the MSAC algorithm provided in the embodiments of the present invention;
[0027] Figure 6 This is an example diagram of the optimized flight trajectory of multiple UAVs provided in an embodiment of the present invention;
[0028] Figure 7 This is a graph showing the relationship between total latency and computing power of each user device, provided in an embodiment of the present invention.
[0029] Figure 8 This is a graph showing the relationship between the total latency and the total computing task scale of each user device, as provided in this embodiment of the invention.
[0030] Figure 9 This is a schematic diagram of the NOMA-assisted multi-UAV edge computing resource allocation and offloading strategy system provided in an embodiment of the present invention. Detailed Implementation
[0031] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1
[0033] This embodiment of the NOMA-assisted multi-UAV edge computing resource allocation and task offloading method minimizes the total maximum latency of the multi-UAV edge computing network by combining NOMA technology, providing better support for computationally intensive user services. It overcomes the shortcomings of traditional OMA scenarios, such as limited communication resources, limited computing power of user equipment, or the inability to complete communication and computing services due to damage to ground base stations under natural disasters. By defining relevant parameters of UAVs and user equipment, user equipment is clustered and a network communication model is established. Based on the network communication model, a computing model is established, and a joint optimization strategy for UAV trajectory, computing resource allocation, and task offloading is determined. According to the system model and optimization strategy, the objective function and constraints of the optimization problem are determined. The optimization problem is transformed into a Markov decision process. Based on state normalization, the Multi-agent Soft Actor-Critic (MSAC) algorithm is used to solve the Markov decision process. Based on the solution results, the joint optimization of task offloading ratio, resource allocation, and multi-UAV trajectory is achieved.
[0034] Please see Figure 1 This embodiment specifically includes the following steps:
[0035] S10. Deploy an offloading architecture in a NOMA-assisted multi-drone edge computing scenario, define relevant parameters for drones and user equipment, cluster user equipment, and establish a network communication model.
[0036] In this embodiment, the parameters related to the drone and user equipment include location, mission parameters, computing power, and mission offloading ratio.
[0037] The deployed offloading architecture is as follows Figure 2 As shown, this includes M drones serving K user devices, each drone being available This indicates that the user equipment is available. The position of the drone m is represented by q. m =(x m ,y m ,h m The position of user equipment k is represented by q. k =(x k ,y k ,0), x m Let y be the x-coordinate of the drone m. m Let h be the ordinate of the drone m. m Let m be the flight altitude of the drone, and x be the altitude of the drone. k Let y be the x-coordinate of user equipment k. k Let k be the ordinate of the user equipment.
[0038] In this embodiment, continuous time is divided into N equal-length and very small time slots. In the nth time slot, the task size of user equipment k is D. k (n), the proportion of tasks unloaded to the drone (i.e., the unloading ratio) is defined as R. k (n), R k (n)∈[0,1].
[0039] Furthermore, during the process of UAVs serving ground user equipment, the user equipment is clustered according to the number of UAVs and the initial location of the user equipment. Each cluster includes one UAV and several user equipment, and a service indicator β is defined. m,k (n)∈{0,1}, specifically:
[0040]
[0041] In this step, NOMA technology improves channel resource utilization by allowing user equipment to share the same time-frequency and other resource blocks. According to the NOMA protocol guidelines, Successive Interference Cancellation (SIC) technology is used at the receiver to detect the signal. Using α... k,l (n)∈{0,1} represents the SIC decoding order of user equipment k and user equipment l, specifically:
[0042]
[0043] Among them, g m,k (n), g m,l (n) represent the channel gain of user equipment k and user equipment l served by UAV m, respectively.
[0044] Therefore, in time slot n, the transmission rate from user equipment k to the drone m providing computing services can be r. m,k (n) represents, specifically:
[0045]
[0046] Where B is the bandwidth allocated to each UAV, N0 represents the power spectral density of the additive white Gaussian noise received, and P k (n) represents the transmit power of user equipment k in time slot n, P l (n) represents the transmit power of user equipment l in time slot n.
[0047] S20. Based on the network communication model, establish a computational model and determine a joint optimization strategy for UAV trajectory, computational resource allocation, and task offloading ratio.
[0048] In the joint optimization strategy determined in this step, the latency calculated locally by user equipment k in time slot n is available. Specifically, it means:
[0049]
[0050] Among them, f UE For each user device, the CPU frequency is C, where C is the number of CPU cycles required to process each bit, and D is... k (n) is the task size of user equipment k in time slot n, R k (n) represents the offloading ratio of user equipment k in time slot n.
[0051] In time slot n, the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m is used. This indicates that the computational latency for drone m to process this part of the upload task is available. Specifically, it means:
[0052]
[0053] Among them, f m,k (n) represents the computing resources allocated from UAV m to user device k in time slot n. Since the computation results from the edge server are typically very small, the transmission delay returned to the user device is negligible.
[0054] The total energy consumption of drone m in processing user equipment tasks across N time slots during the entire service period. It indicates that the drone's flight energy consumption is... Specifically, it means:
[0055]
[0056] Where k is the CPU capacitance coefficient of each UAV. M is the cube of the computing resources allocated from the drone m to the user equipment k. UAV Let v(n) be the mass of the drone, v(n) be the average velocity of the drone, and t be the average velocity of the drone. fly This refers to the flight time of the drone.
[0057] S30. Based on the computational model and joint optimization strategy, determine the objective function and constraints of the optimization problem.
[0058] The objective function is:
[0059]
[0060] The constraints are as follows:
[0061]
[0062] Where, q m(n) represents the position of UAV m at time slot n, X size Y is the length of the service area. size H represents the width of the service area. range For the flight altitude range of the drone; α l,k θ(n) represents the SIC decoding order of user equipment l and user equipment k; θ(n) is the elevation angle of the UAV flight. f is the azimuth angle of the drone's flight. UAV E represents the total computing resources available for the drone. UAV This represents the total energy of the drone's battery.
[0063] S40. Transform the optimization problem into a Markov decision process and determine the environment, state, action space, and reward of the Markov decision process.
[0064] The environment includes drones and user equipment; the status includes drone location information in the environment, channel gain between drones and user equipment, the size of computing tasks generated by each user equipment, and the remaining battery power of each drone.
[0065] The action space is determined by the task offloading ratio strategy of each user in each cluster, the computing resource allocation of the edge server carried by the drone, and the drone's flight trajectory.
[0066] The rewards include both incentives and penalties for performing actions in the target direction.
[0067] The SAC algorithm is an advanced single-agent off-policy algorithm designed for continuous action spaces. It extends soft-value functions and entropy maximization to the standard actor-critic architecture, where entropy represents uncertain states to promote efficient exploration and stable learning through stochastic action choices.
[0068] This embodiment is based on the single-agent SAC algorithm to form a multi-agent SAC (Multi-agent SoftActor-Critic) algorithm, namely the MSAC algorithm. To allow multiple drones to share the neural network, the state information of each drone must be extracted and formed into a standard array before inputting the state into the neural network. The specific differences between the MSAC algorithm and the SAC algorithm are as follows: Figure 3As shown, the drone currently connected to the neural network, i.e., the drone that is currently serving, needs to latch its input neurons. For example, in the MSAC algorithm, when drone 1 is serving and connected to the neural network, drone 1's position information needs to be input into the first neuron; when drone 2 is serving, drone 2's position information must also be input into the same neuron. In other words, the position of the drone currently serving and connected to the neural network, as well as the channel gain, task size, and remaining battery power of the drone it serves, must be input into the designated neurons. However, in the SAC algorithm, this input is not required into the designated neurons.
[0069] This step normalizes the state and uses the MSAC algorithm to solve the Markov decision process. Based on the solution results, the joint optimization of task unloading ratio, resource allocation, and multi-UAV trajectories is achieved.
[0070] Please see Figure 4 This embodiment provides a process for joint optimization of task offloading ratio, resource allocation, and multi-UAV trajectories based on the MSAC algorithm, including: transforming the optimization problem into a Markov Decision Process (MDP), and proposing a method based on the MSAC algorithm to solve the MDP problem by jointly optimizing the task offloading ratio, resource allocation, and multi-UAV trajectories; simulating the system environment and verifying convergence to prove the feasibility and practicality of the algorithm and model.
[0071] Specifically, the MSAC algorithm is used to solve the Markov decision process as follows:
[0072] 1) A Markov Decision Process (MDP) is defined as a quintuple (S, A, P, R, γ), where S is a finite set of system states, A is a finite set of actions, P is the state transition probability, R is the reward function, and γ is the discount factor used to calculate the cumulative reward, γ∈(0,1).
[0073] 2) Deep reinforcement learning is the interaction process between an intelligent agent (i.e., a drone) and its environment. When the agent performs a task, it interacts with the environment, generating new states and receiving rewards from the environment. Deep reinforcement learning continuously corrects the agent's action strategy based on the data generated by the interaction. After multiple iterations, the agent continuously executes actions in the direction that maximizes the reward until the task is completed.
[0074] The states include the drone's location information in the environment, the channel gain between the drone and the user equipment, the workload of the user equipment, and the drone's remaining battery power; the states are defined as follows:
[0075] s n ={q m (n),qi (n),g m,k (n),g i,l (n),D k (n),D l (n),E m (n),E i (n)},
[0076]
[0077] Where m represents the drone currently in service, i represents other drones, k represents the user equipment associated with the drone currently in service, and l represents the remaining user equipment.
[0078] 3) The action space is mainly determined by the user task offloading ratio, the allocation of UAV computing resources, and the UAV flight trajectory. The action space a n Defined as:
[0079]
[0080] Among them, R k (n) represents the offloading ratio of user equipment k in time slot n, f m,k θ(n) represents the computing resources allocated to user equipment k by UAV m in time slot n, and θ(n) represents the elevation angle of the UAV's flight. This represents the azimuth angle of the UAV's flight, where 0 ≤ θ(n) ≤ π.
[0081] 4) During interaction with the environment, the agent will continuously execute action strategies along the path that maximizes the cumulative reward. In this embodiment, the negative of the total maximum delay is set as the reward. When the task within the current time slot is not completed, a penalty is required. Therefore, the reward function is set as follows:
[0082]
[0083] Here, λ is the penalty coefficient. When the agent's action cannot complete the computation task within the current time slot, λ increases from 0, reducing the reward and serving as a penalty. The delay calculated locally by user equipment k in time slot n; This refers to the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m in time slot n. The computational delay for the uploading task of user equipment k by UAV m in time slot n.
[0084] 5) Based on the MSAC algorithm, perform joint optimization of task offloading ratio, resource allocation and multi-UAV trajectory to solve the MDP problem.
[0085] When the intelligent agent (i.e., the drone) begins to serve the user equipment, it seeks the optimal action based on the current state, and adopts optimized user equipment task offloading ratio, computing resource allocation and drone flight trajectory to provide services to the user equipment. After the action is executed, the intelligent agent will receive the reward for performing the action and update the state of the drone after performing the action.
[0086] Set the relevant parameters: allocate a bandwidth of B = 4MHz to each drone, a maximum transmit power of P = 20dBm for the user equipment, and set the CPU frequencies of the user equipment and drones to f0 and f0 respectively. UE =0.8GHz and f UAV =10GHz, number of time slots N=60, C=1000 cycles / bit, UAV CPU capacitance coefficient κ=10 -28 Drone mass M UAV = 9.65kg, Total battery energy of the drone E UAV =100kJ. Data simulation of the algorithm was performed using Python 3.9 and PyTorch 2.0.
[0087] Figure 5 The convergence of the proposed algorithm is demonstrated with a learning rate lr = 0.000009, a discount factor γ = 0.95, a batch size = 128, and a soft update coefficient τ = 0.005. It can be seen that in the first 300 training iterations, JOORT consistently maintained a latency exceeding 300 seconds, which gradually decreased thereafter. After approximately 1100 training iterations, the latency stabilized at around 140 seconds, indicating that the algorithm had reached convergence.
[0088] Figure 6 An example of optimized multi-drone flight trajectories is given. Observation shows that when the user equipment moves linearly in a directional direction, each drone exhibits an upward trend while simultaneously approaching the user equipment it serves in the horizontal direction. After reaching the upper limit of flight altitude, each drone adjusts its flight trajectory in the horizontal direction to approach its respective user equipment.
[0089] Figure 7 The relationship between total latency and the computing power of each user device is shown. As the computing power of user devices increases, the total latency shows a significant decreasing trend. The proposed MSAC algorithm also significantly outperforms other algorithms in terms of latency performance. As can be seen from the figure, even if other variables are optimized but the UAV flight trajectory is not optimized, i.e., a circular trajectory is performed, the overall system performance is poor, which is not conducive to latency-sensitive tasks. Therefore, trajectory optimization is an important optimization variable that cannot be ignored.
[0090] Figure 8The relationship between total latency and the total computational task size per user device is illustrated. As the number of tasks increases, the total latency also increases, and the gap between MSAC and other schemes widens. Clearly, MSAC outperforms other schemes, and its performance improves with the increase of optimization variables, especially for multi-UAV 3D flight trajectories, a key optimization factor. When the task size is small, the drawbacks of random user clustering are not very noticeable. However, as the task size increases, the latency of random user clustering increases significantly because user devices are assigned to unsuitable UAVs, thus amplifying transmission and overall latency.
[0091] In summary, this invention integrates UAVs with edge servers and utilizes NOMA technology to assist in a multi-UAV edge computing network. Within this system architecture, user clustering is achieved, a computational model is established based on the network model, and a joint optimization strategy is determined for UAV trajectories, computational resource allocation, and task offloading ratios. The objective function and constraints of the optimization problem are defined, transforming the optimization problem into a Markov decision process. The MSAC algorithm employed in this invention is a multi-agent SAC (Multi-agent Soft Actor-Critic) algorithm based on the single-agent SAC algorithm, allowing multiple UAVs to share the neural network. Simultaneously, this invention uses NOMA technology instead of OMA technology to improve spectral efficiency. By combining NOMA technology, the total maximum latency of all user devices within the multi-UAV edge computing network is reduced, providing better support for computationally intensive and latency-sensitive services. This overcomes the shortcomings of traditional OMA scenarios, such as limited communication resources, limited user device computing capabilities, or the inability to provide communication and computing services due to damage to ground base stations caused by natural disasters. This has practical significance for reducing the total maximum latency of the system in multi-UAV communication scenarios.
[0092] Example 2
[0093] Please see Figure 9 Based on the same inventive concept as Embodiment 1, this embodiment provides a resource allocation and task offloading system for NOMA-assisted multi-UAV edge computing, including an architecture deployment unit 01, a strategy determination unit 02, an optimization problem determination unit 03, and a transformation and solution unit 04.
[0094] The architecture deployment unit is used to define the relevant parameters of drones and user equipment, cluster user equipment, and establish a network communication model in NOMA-assisted multi-drone edge computing scenarios.
[0095] The NOMA-assisted multi-drone edge computing network architecture includes M drones serving K user devices, with each drone capable of providing... This indicates that the user equipment is available. The position of the drone m is represented by q.m =(x m ,y m ,h m The position of user equipment k is represented by q. k =(x k ,y k ,0), x m Let y be the x-coordinate of the drone m. m Let h be the ordinate of the drone m. m Let m be the flight altitude of the drone, and x be the altitude of the drone. k Let y be the x-coordinate of user equipment k. k Let be the ordinate of user equipment k. In this embodiment, continuous time is divided into N equal-length and very small time slots. In the nth time slot, the task size of user equipment k is D. k (n), the proportion of tasks unloaded to the drone (i.e., the unloading ratio) is defined as R. k (n), R k (n)∈[0,1].
[0096] User equipment is clustered based on the number of drones and the initial location of the user equipment. Each cluster includes one drone and several user devices. A service indicator β is defined. m,k (n)∈{0,1}, specifically:
[0097]
[0098] NOMA technology improves channel resource utilization by allowing users to share the same time-domain and other resource blocks. According to the NOMA protocol guidelines, Successive Interference Cancellation (SIC) is used at the receiver to detect signals. Using α... k,l (n)∈{0,1} represents the SIC decoding order of user equipment k and user equipment l, specifically:
[0099]
[0100] Among them, g m,k (n), g m,l (n) represent the channel gain of user equipment k and user equipment l served by UAV m, respectively.
[0101] The strategy determination unit is used to establish a computational model based on the network communication model, and to determine the joint optimization strategy for UAV trajectory, computational resource allocation and task offloading ratio.
[0102] At time slot n, the transmission rate from user equipment k to the drone m providing computing services is available as r. m,k (n) represents, specifically:
[0103]
[0104] Where B is the bandwidth allocated to each UAV, N0 represents the power spectral density of the additive white Gaussian noise received, and P k (n) represents the transmit power of user equipment k in time slot n, P l (n) represents the transmit power of user equipment l in time slot n.
[0105] At time slot n, the available latency for tasks computed locally by user equipment k is [not specified]. Specifically, it means:
[0106]
[0107] Among them, f UE For each CPU frequency, C is the number of CPU cycles required to process each unit byte, and D... k (n) is the total task size of user equipment Uk in time slot n, R k (n) represents the offloading ratio of user equipment UK in time slot n.
[0108] At time slot n, user equipment k uploads the offloaded portion of the task to drone m for processing. The transmission time required to upload to drone m is... The computational latency required for drone m to process this part of the task is Over the entire service time (N time slots), the total energy consumption of drone m in processing user equipment tasks and its flight energy consumption can be respectively used as... and express.
[0109] The optimization problem determination unit is used to determine the objective function and constraints of the optimization problem based on the computational model and joint optimization strategy.
[0110] The transformation and solution unit is used to transform the optimization problem into a Markov decision process, normalize the state, solve the Markov decision process using the MSAC algorithm, and generate a policy method based on the solution results.
[0111] It is understood that the system provided in this embodiment is used to execute the method provided in the above embodiments and achieve the same technical effect, and will not be described in detail here.
[0112] This invention reduces the total maximum latency of all user devices in a multi-UAV edge computing network by combining NOMA technology, providing better support for computationally intensive and latency-sensitive services. It overcomes the shortcomings of traditional OMA scenarios, such as limited communication resources, limited computing power of user devices, or the inability to complete communication computing services due to damage to ground base stations caused by natural disasters. It has practical significance for reducing the total maximum latency of the system in multi-UAV communication scenarios.
[0113] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A method for resource allocation and task offloading in NOMA-assisted multi-UAV edge computing, characterized in that, Includes the following steps: S10. Deploy an offloading architecture in the NOMA-assisted multi-drone edge computing scenario, define relevant parameters of drones and user equipment, cluster user equipment, and establish a network communication model. S20. Based on the network communication model, establish a computing model and determine a joint optimization strategy for UAV trajectory, computing resource allocation, and task offloading ratio; S30. Based on the computational model and joint optimization strategy, determine the objective function and constraints of the optimization problem; S40. Transform the optimization problem into a Markov decision process, and determine the environment, state, action space and reward of the Markov decision process; normalize the state, and use the MSAC algorithm to solve the Markov decision process, and realize the joint optimization of task unloading ratio, resource allocation and multi-UAV trajectory based on the solution results; The MSAC algorithm is a multi-agent SAC algorithm based on the single-agent SAC algorithm. In the MSAC algorithm, before the state is input into the neural network, the state of each UAV is extracted and formed into a standard array, so that multiple UAVs can share the neural network. The location of the drone that is being served and connected to the neural network, the channel gain of the user equipment it serves, the task size, and the remaining battery power of the drone are input into the designated neuron. The objective function determined in step S30 is: The constraints are: in, The delay calculated locally by user equipment k in time slot n; This refers to the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m in time slot n. The computational latency for UAV m to process the upload task of user equipment k in time slot n; q m (n) represents the position of UAV m at time slot n, X size Y is the length of the service area. size H represents the width of the service area. range For the flight altitude range of the drone; α k,l (n) represents the SIC decoding order of user equipment k and user equipment l, α l,k (n) represents the SIC decoding order of user equipment l and user equipment k, g m,k (n), g m,l θ(n) represents the channel gain of user equipment k and user equipment l served by UAV m, respectively; θ(n) is the elevation angle of the UAV flight. R is the azimuth angle of the drone's flight. k (n) represents the offloading ratio of user equipment k in time slot n; f m,k (n) represents the computing resources allocated to user equipment k by UAV m in time slot n, f UAV Let m be the total computing resources of the drone; during the N time slots of the entire service time, the total energy consumption of drone m in processing the tasks of the user equipment it serves is... The drone's flight energy consumption is E UAV This refers to the total energy of the drone's battery. Define β m,k (n)∈{0,1} is the service indicator, specifically:
2. The method according to claim 1, characterized in that, In step S10, the relevant parameters of the UAV and user equipment include location, mission parameters, computing power, and mission offloading ratio; The deployed offloading architecture consists of M drones serving K user devices, each drone using It indicates that the user equipment uses The position of drone m is represented by q. m =(x m ,y m ,h m The position of user equipment k is represented by q. k =(x k ,y k ,0), x m Let y be the x-coordinate of the drone m. m Let h be the ordinate of the drone m. m Let m be the flight altitude of the drone, and x be the altitude of the drone. k Let y be the x-coordinate of user equipment k. k Let k be the ordinate of user equipment k.
3. The method according to claim 1, characterized in that, In step S10, the continuous time is divided into N time slots of equal length. During the process of the UAV serving the ground user equipment, the user equipment is clustered according to the number of UAVs and the initial position of the user equipment. Each cluster includes one UAV and several user equipment. At time slot n, the user equipment k transmits data to the drone m providing computing services at a rate of r. m,k (n) is represented as: Where B is the bandwidth allocated to each UAV, N0 represents the power spectral density of the additive white Gaussian noise received, and P k (n) represents the transmit power of user equipment k in time slot n; α k,l (n) represents the SIC decoding order of user equipment k and user equipment l, α k,l (n)∈{0,1};g m,k (n), g m,l (n) represent the channel gains of user equipment k and user equipment l served by UAV m, respectively; and we have:
4. The method according to claim 3, characterized in that, In the joint optimization strategy determined in step S20, the latency of user equipment k is calculated locally in time slot n. for: Among them, f UE For each user device, the CPU frequency is C, where C is the number of CPU cycles required to process each bit, and D is... k (n) is the task size of user equipment k in time slot n, R k (n) represents the offloading ratio of user equipment k in time slot n; At time slot n, the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m is: The computational latency for drone m to process the upload task from user equipment k is... Among them, f m,k (n) represents the computing resources allocated to user equipment k by UAV m in time slot n; Over the entire service time, in N time slots, the total energy consumption of drone m for handling the tasks of user equipment it serves is... The drone's flight energy consumption is Where κ is the CPU capacitance coefficient of each UAV. M is the cube of the computing resources allocated from the drone m to the user equipment k. UAV Let v(n) be the mass of the drone, v(n) be the average velocity of the drone, and t be the average velocity of the drone. fly This refers to the flight time of the drone.
5. The method according to claim 1, characterized in that, The environment described in step S40 includes drones and user equipment; the state includes drone location information in the environment, channel gain between drones and user equipment, the size of computing tasks generated by each user equipment, and the remaining battery power of each drone. The action space is determined by the task offloading ratio of each user in each cluster, the computing resource allocation of the edge server carried by the drone, and the drone's flight trajectory. The rewards include both incentives and penalties for performing actions in the target direction.
6. The method according to claim 5, characterized in that, Step S40, which uses the MSAC algorithm to solve the Markov decision process, includes: A Markov Decision Process (MDP) is a quintuple (S, A, P, R, γ), where S is a finite set of states, A is a finite set of actions, P is the state transition probability, R is the reward function, and γ is a discount factor used to calculate the cumulative reward, γ∈(0,1]. When a drone performs a mission, it interacts with the environment, generates new states, and receives rewards from the environment. Deep reinforcement learning continuously corrects the drone's action strategy based on the data generated by the interaction. After multiple iterations, the drone continues to perform actions in the direction of maximizing the reward until the mission is completed. Action space a n Defined as: Among them, R k (n) represents the task offloading ratio of user equipment k in time slot n, f m,k θ(n) represents the computing resources allocated to user equipment k by UAV m in time slot n, and θ(n) represents the elevation angle of the UAV's flight. This represents the azimuth angle of the UAV's flight, where 0 ≤ θ(n) ≤ π. Set the negative of the total maximum delay as the reward, and set a penalty when the task within the current time slot is not completed; set the reward function as follows: Wherein, λ is the penalty coefficient. When the drone's actions cannot complete the computation task within the current time slot, λ increases from 0, reducing the reward and serving as a penalty. The delay calculated locally by user equipment k in time slot n; This refers to the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m in time slot n. The computational delay for UAV m to process the upload task of user equipment k in time slot n; The MSAC algorithm is used to jointly optimize the task offloading ratio, resource allocation, and multi-UAV trajectories to solve the Markov decision problem. When a UAV starts serving a user device, it seeks the optimal action based on the current state and provides services to the user device by optimizing the user device task offloading ratio, computational resource allocation, and UAV flight trajectory. After the action is executed, the UAV will receive the reward for performing the action and update the UAV's state after performing the action.
7. A resource allocation and task offloading system for NOMA-assisted multi-UAV edge computing, characterized in that, include: The architecture deployment unit deploys the offloading architecture in the NOMA-assisted multi-drone edge computing scenario, defines the relevant parameters of drones and user equipment, clusters user equipment, and establishes a network communication model. The strategy determination unit is used to establish a computational model based on the network communication model and determine the joint optimization strategy for UAV trajectory, computational resource allocation and task offloading ratio; The optimization problem determination unit is used to determine the objective function and constraints of the optimization problem based on the computational model and joint optimization strategy. The transformation and solution unit is used to transform the optimization problem into a Markov decision process, determine the environment, state, action space and reward of the Markov decision process; normalize the state, use the MSAC algorithm to solve the Markov decision process, and realize the joint optimization of task unloading ratio, resource allocation and multi-UAV trajectory based on the solution results. The MSAC algorithm is a multi-agent SAC algorithm based on the single-agent SAC algorithm. In the MSAC algorithm, before the state is input into the neural network, the state of each UAV is extracted and formed into a standard array, so that multiple UAVs can share the neural network. The location of the drone that is being served and connected to the neural network, the channel gain of the user equipment it serves, the task size, and the remaining battery power of the drone are input into the designated neuron. The objective function determined by the problem-solving unit is: The constraints are: in, The delay calculated locally by user equipment k in time slot n; This refers to the transmission time required for user equipment k to upload the offloaded proportion of tasks to drone m in time slot n. The computational latency for UAV m to process the upload task of user equipment k in time slot n; q m (n) represents the position of UAV m at time slot n, X size Y is the length of the service area. size H represents the width of the service area. range For the flight altitude range of the drone; α k,l (n) represents the SIC decoding order of user equipment k and user equipment l, α l,k (n) represents the SIC decoding order of user equipment l and user equipment k, g m,k (n), g m,l θ(n) represents the channel gain of user equipment k and user equipment l served by UAV m, respectively; θ(n) is the elevation angle of the UAV flight. R is the azimuth angle of the drone's flight. k (n) represents the offloading ratio of user equipment k in time slot n; f m,k (n) represents the computing resources allocated to user equipment k by UAV m in time slot n, f UAV Let m be the total computing resources of the drone; during the N time slots of the entire service time, the total energy consumption of drone m in processing the tasks of the user equipment it serves is... The drone's flight energy consumption is E UAV This refers to the total energy of the drone's battery. Define β m,k (n)∈{0,1} is the service indicator, specifically:
8. The system according to claim 7, characterized in that, In the transformation and solution unit, the environment includes UAVs and user equipment; the state includes the location information of UAVs in the environment, the channel gain between UAVs and user equipment, the size of the computing tasks generated by each user equipment, and the remaining power of each UAV. The action space is determined by the task offloading ratio strategy of each user in each cluster, the computing resource allocation of the edge server carried by the drone, and the drone's flight trajectory. The rewards include both incentives and penalties for performing actions in the target direction.
Citation Information
Patent Citations
Task offloading optimization method and system for multi-unmanned aerial vehicle assisted mobile edge computing
CN117729564A