An Unmanned Aerial Vehicle-Assisted Joint Optimization Method and Device for Terahertz Communication Networks
By building a drone-assisted terahertz communication network system model and using deep reinforcement learning algorithms, the drone location, calculation offload ratio and computing resource allocation scheme are optimized, and the problem of how to reduce user delay under service quality and resource constraints is solved, and network capacity is improved and delay is reduced.
Patent Information
- Application Number
- CN202210454105.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-27
AI Technical Summary
How to jointly optimize the drone location, calculate the unloading ratio and calculate the resource allocation plan in real time under the constraints of service quality and resource, so as to minimize the sum of the delays of all users.
Build a drone-assisted terahertz communication network system model, and optimize the drone location, calculate the offload ratio and computing resource allocation scheme based on deep reinforcement learning algorithms to minimize the sum of delays of all users in the communication network system.
Under the constraints of user service quality and resource, network capacity is effectively improved, delayed is reduced, and the needs of various delay-sensitive services are met.
Smart Images

Figure CN114980160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly relates to a method and apparatus for jointly optimizing a terahertz communication network assisted by an unmanned aerial vehicle. Background Art
[0002] With the rapid development of Internet of Things (IoT) technologies, many latency-sensitive applications such as remote healthcare, autonomous driving, virtual reality, and augmented reality have gradually entered people's daily lives, and these applications generate a large number of computationally intensive tasks. Although the CPU performance in the new generation of IoT devices is getting stronger and stronger, it is still unable to process computationally intensive tasks in a short period of time. To solve the problem of limited computing power of IoT devices, cloud computing technology transfers computationally intensive tasks from the user side to the cloud server for computing and processing, effectively reducing the latency. However, it is expected that by 2025, the number of IoT devices will reach 75 billion, and transmitting a large amount of data to the cloud server will consume a large amount of network resources and bring great computing pressure to the cloud server. Therefore, cloud computing technology can no longer meet the real-time computing and processing of a large amount of data. To make up for the deficiencies of cloud computing, mobile edge computing (MEC) technology transfers the functions of the core network to the network edge by deploying edge access points (E-APs) on the IoT device side, reducing the requirements for the bandwidth of the backhaul link and effectively improving the quality of service.
[0003] Traditional E-APs are deployed at fixed positions, and their coverage range and the number of users that can be served simultaneously are limited. With the development of unmanned aerial vehicle (UAV) technology, deploying a server on a UAV has become an effective way to improve system capacity. When the number of users exceeds the capacity limit of E-APs or users are outside the coverage range of E-APs, the UAV can carry a server to provide computing offloading services for users. Compared with the traditional architecture, the UAV-assisted architecture has higher scalability and flexibility.
[0004] To better support computationally intensive applications, it is necessary to reduce the transmission latency from the user to the server. The rate of terahertz communication can reach dozens of Gb / s, which is significantly better than the current ultra-wideband technology. Therefore, terahertz communication technology has attracted much attention and has become a key technology to meet the real-time service requirements of mobile heterogeneous network systems. Due to the sensitivity of the terahertz band to channel congestion, deploying a server on a UAV can effectively reduce the impact of obstacles on the communication link. Therefore, it is very promising to carry a server on a UAV to provide computing offloading services in the terahertz band.
[0005] Currently, how to jointly optimize the UAV position, computing offloading ratio, and computing resource allocation scheme in real time under service quality and resource constraints to minimize the sum of delays of all users is an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a method and device for jointly optimizing an unmanned aerial vehicle (UAV)-assisted terahertz communication network to solve the problem of jointly optimizing the UAV position, computing offloading ratio, and computing resource allocation scheme.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] On the one hand, the present invention provides a method for jointly optimizing a UAV-assisted terahertz communication network, and the method for jointly optimizing the UAV-assisted terahertz communication network includes:
[0009] Construct a system model of a UAV-assisted terahertz communication network; wherein, in the communication network system model, the UAV carries a server to provide computing offloading services for users in the terahertz band;
[0010] Based on the communication network system model, under the service quality of users and resource constraints, with the goal of minimizing the sum of delays of all users in the communication network system, construct an optimization objective function;
[0011] Based on a preset deep reinforcement learning algorithm, obtain the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function, realize the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieve the purpose of improving network capacity and reducing delays.
[0012] Further, in the communication network system model, the path loss PL(f, D) of the terahertz communication link between the server carried by the UAV and the user is expressed as:
[0013]
[0014] where L abs (f, D) represents the molecular absorption loss, L spread (f, D) represents the transmission loss, D represents the distance between the user and the UAV server, c is the speed of light in a vacuum state, k abs (f) is the medium absorption coefficient related to the frequency, and f represents the terahertz carrier frequency.
[0015] Further, the optimization objective function is expressed as:
[0016]
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023] wherein, T i represents the total delay of the i-th user, N represents the number of users, x uav and y uav represent the coordinate information of the UAV, and α i represents the offloading ratio of the i-th user, and β i represents the proportion of computing resources allocated to the i-th user. represents the computing offloading vector, represents the computing resource allocation vector, represents the local computing energy consumption, represents the uploading energy consumption, represents the standby energy consumption of the user waiting for the server to process data, and t i,max represents the maximum tolerable delay of the i-th user, and E i,max represents the maximum tolerable energy consumption of the i-th user, represents the set of users that cannot be served by the E-APs, represents the preset coordinate threshold of the UAV;
[0024] C1 means that the total delay of each user does not exceed the maximum tolerable delay, thus ensuring the quality of service of the user; C2 means that the position of the UAV is within the preset specified range; C3 and C4 mean that the sum of the computing resources allocated to each user does not exceed the total computing resources; C5 means that the user can offload any proportion of partial tasks to the server for processing; C6 means that the energy consumed by the user is within the specified range.
[0025] Furthermore, obtaining the optimal UAV position, computing offloading ratio and computing resource allocation scheme that meet the optimization objective function based on the preset deep reinforcement learning algorithm includes:
[0026] Taking the UAV, the server and all users as agents, the UAV-assisted terahertz communication network system model serves as the environment, and the UAV position, computing offloading ratio and computing resource allocation scheme serve as the action output of the agent. The preset deep reinforcement learning algorithm is used to train the agent to obtain the optimal UAV position, computing offloading ratio and computing resource allocation scheme that meet the optimization objective function.
[0027] Further, the preset deep reinforcement learning algorithm is the DDPG (deep deterministic policy gradient) algorithm.
[0028] Further, using the preset deep reinforcement learning algorithm to train the agent includes:
[0029] Step 1: Initialize the state space, action space, and deep neural network parameters of the system;
[0030] Step 2: The agent selects and executes an action according to the current state and the policy network;
[0031] Step 3: After the agent executes the action, return the reward and the new state, and put the state conversion process into the experience cache space;
[0032] Step 4: Sample a preset number of state transition data in the experience cache space as the training data for training the Q network and the training policy network;
[0033] Step 5: Calculate the gradients of the cost functions of the Q network and the policy network respectively;
[0034] Step 6: Update the target neural network parameters.
[0035] Further, initializing the state space, action space, and deep neural network parameters of the system includes:
[0036] Model the user resource requirements and channel states as a finite state Markov model;
[0037] Create two target neural networks μ′(F, ω′) and Q′(F, G, λ′) for the policy network μ(F, ω) and the Q network Q(F, G, λ) respectively for parameter update.
[0038] Further, after the agent executes the action, the returned reward includes:
[0039] After the agent executes the action, determine whether the preset conditions are met. When the preset conditions are met, obtain the immediate reward according to the environment; wherein, the preset conditions include: the delay of each user meets the quality of service constraint; the position of the UAV is within the specified range; the computing resources allocated to each user do not exceed the total resource amount; the computing offloading ratio is within the preset range; the total energy consumption of each user meets the energy saving requirements.
[0040] The expression of the immediate reward R is:
[0041]
[0042] where, Tn denotes the latency of the nth user, and N denotes the number of users.
[0043] Furthermore, the step of separately calculating the gradients of the cost functions of the Q-network and the policy network includes:
[0044] The gradients of the cost functions of the Q-network and the policy network are separately calculated, and the stochastic gradient descent method is used to update the neural network parameters.
[0045] On the other hand, the present invention also provides a joint optimization device for an unmanned aerial vehicle (UAV)-assisted terahertz communication network. The UAV-assisted terahertz communication network joint optimization device includes:
[0046] A communication network system model construction module, configured to construct a UAV-assisted terahertz communication network system model; wherein, in the communication network system model, a UAV carries a server to provide computing offloading services for users in the terahertz band;
[0047] An optimization objective function construction module, configured to construct an optimization objective function with the goal of minimizing the sum of the latencies of all users in the communication network system under user quality of service and resource constraints, based on the communication network system model constructed by the communication network system model construction module;
[0048] A joint optimization module, configured to obtain an optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function constructed by the optimization objective function construction module based on a preset deep reinforcement learning algorithm, so as to realize the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieve the purpose of improving network capacity and reducing latency.
[0049] On yet another aspect, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0050] On still another aspect, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.
[0051] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0052] The joint optimization method for the UAV-assisted terahertz communication network of the present invention realizes the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme under user quality of service and resource constraints, makes up for the shortcomings of the limited coverage range of edge access nodes and the limited number of accessed users, effectively improves the network capacity and reduces the latency under resource constraints, and meets the requirements of various latency-sensitive services. Brief Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0054] Figure 1 It is a schematic execution flowchart of a method for jointly optimizing an unmanned aerial vehicle (UAV)-assisted terahertz communication network provided by an embodiment of the present invention;
[0055] Figure 2 It is a schematic diagram of a UAV-assisted terahertz network architecture provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic flowchart of a joint optimization algorithm based on deep reinforcement learning provided by an embodiment of the present invention. Detailed Embodiments
[0057] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0058] First Embodiment
[0059] This embodiment provides a method for jointly optimizing a UAV-assisted terahertz communication network. By leveraging the characteristics of strong UAV flexibility and high terahertz communication transmission rate, it makes up for the shortcomings of limited coverage range and limited number of access users of E-APs, and effectively improves network capacity and reduces latency under resource constraints. This method can be implemented by an electronic device. The execution process of this method is as Figure 1 shown and includes the following steps:
[0060] S1. Construct a system model of a UAV-assisted terahertz communication network; wherein, in the communication network system model, the UAV carries a server to provide computing offloading services for users in the terahertz band;
[0061] S2. Based on the communication network system model, under the user service quality and resource constraints, construct an optimization objective function with the goal of minimizing the sum of the latencies of all users in the communication network system;
[0062] S3. Based on a preset deep reinforcement learning algorithm, obtain the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function, realize the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieve the goal of improving network capacity and reducing latency.
[0063] Specifically, the communication network system model constructed in this embodiment is as follows Figure 2 As shown, in this model, the path loss PL(f, D) of the terahertz communication link between the server carried by the UAV and the user is expressed as:
[0064]
[0065] where L abs (f, D) represents the molecular absorption loss, L spread (f, D) represents the transmission loss, D represents the distance between the user and the UAV server, c is the speed of light in a vacuum, k abs (f) is the medium absorption coefficient related to the frequency, and f represents the terahertz carrier frequency.
[0066] Since the coverage range of E-APs and the number of access users are limited, there are cases where some users cannot be served by E-APs. represents the set of these users, and the task of each user can be expressed as ζ i ∈ {d i , c i , o i , t i,max}}, d i represents the size of the computing task of the i-th user, c i represents the number of CPU cycles required for the computing task of the i-th user, o i represents the size of the computing result of the i-th user, t i,max represents the maximum tolerable delay of the i-th user. To minimize the delay, this problem can be modeled as:
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074] where T i represents the total delay of the i-th user, x uav and y uav represent the coordinate information of the UAV, α i represents the offloading ratio of the i-th user, β irepresents the proportion of computing resources allocated to the \(i\)-th user, represents the computing offloading vector, represents the computing resource allocation vector, represents the local computing energy consumption, represents the uploading energy consumption, represents the standby energy consumption of the user waiting for the server to process data, \(t\) i,max represents the maximum tolerable delay of the \(i\)-th user, \(E\) i,max represents the maximum tolerable energy consumption of the \(i\)-th user; \(C1\) represents that the total delay of each user does not exceed the maximum tolerable delay, ensuring the quality of service of the user; \(C2\) represents that the position of the UAV is within the specified range; \(C3\) and \(C4\) represent that the sum of the computing resources allocated to each user does not exceed the total computing resources; \(C5\) represents that the user can offload any proportion of partial tasks to the server for processing; \(C6\) represents that the energy consumed by the user is within the specified range.
[0075] Furthermore, based on the preset deep reinforcement learning algorithm, the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function are obtained. Specifically: taking the UAV, the server, and all users as agents, the UAV-assisted terahertz communication network system model serves as the environment, and the UAV position, computing offloading ratio, and computing resource allocation scheme serve as the action outputs of the agents. A preset deep reinforcement learning algorithm is used to train the agents to obtain the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function. Among them, the preset deep reinforcement learning algorithm adopted in this embodiment is the Deep Deterministic Policy Gradient (DDPG) algorithm.
[0076] In the process of jointly optimizing the UAV position, computing offloading ratio, and computing resource allocation scheme using DDPG, considering the dynamic changes of the system state in the real environment, the system state is modeled as a first-order Markov decision model. The deterministic policy network is used to select actions according to the state, and the Q network is used to measure the performance of the selected actions. Since a single neural network will cause the learning process to be very unstable, a target neural network copy is created for each of the policy network and the Q network for network learning, and they are called target networks, which are used to calculate the corresponding target values. The target network and the training network have the same network structure, but their parameter settings are different. When executing the DDPG algorithm, the UAV-assisted terahertz communication network system model serves as the environment, and the UAV position, computing offloading ratio, and computing resource allocation scheme serve as the action outputs of the agents. The specific steps of the algorithm are as Figure 3 shown, including the following steps:
[0077] Initialize the state space, action space, and deep neural network parameters of the system; specifically: initialize the resource requirements, location information, DDPG algorithm parameters, Q-network and policy network parameters of each user, and assign the Q-network and policy network parameters to the target Q-network and target policy network respectively. Among them, the user requirements and channel state are modeled as a finite state Markov model. This system is a discrete time slot system. At the same moment, the system state does not change. The system state at the next moment is generated by the agent based on the behavior policy.
[0078] The DDPG algorithm includes four deep neural networks, namely the policy network μ(F, ω), Q-network Q(F, G, λ), target policy network μ′(F, ω′), and target Q-network Q′(F, G, λ′). ω, λ, ω′, and λ′ represent the parameters of the four deep neural networks respectively. The agent selects and executes actions according to the behavior policy. At each iteration, first obtain the channel state and resource requirement information. The agent obtains the current information, selects and executes actions according to the policy network μ(F, ω). The actions include adjusting the UAV position, calculating the offloading ratio, and calculating the resource allocation scheme. After executing the actions, return the reward R t and the new state. For DDPG, the selection of actions is a deterministic behavior policy, and the behavior of each step directly obtains a definite value through μ(F, ω).
[0079] Among them, after the agent executes the actions, it returns the reward. Specifically: after the agent executes the actions, it judges whether the preset conditions are met. When the preset conditions are met, it obtains the immediate reward according to the environment. Among them, the preset conditions include: 1) The delay of each user meets the quality of service constraint; 2) The position of the UAV is within the specified range; 3) The computing resources allocated to each user do not exceed the total resource amount; 4) The offloading ratio is within the preset range; 5) The total energy consumption of each user meets the energy saving requirements.
[0080] The expression of the immediate reward R is:
[0081]
[0082] Among them, T n represents the delay of the nth user, and N represents the number of users.
[0083] After the agent executes the actions, it returns the reward and the new state, and puts the state transition process (F t , G t , R t , F t+1 ) into the experience cache space D. F t represents the state at time t, G t represents the action at time t, and R t represents the reward obtained in the state F tExecute action G t The obtained reward, F t+1 Indicates the state F t Execute action G t The next state reached. To train the neural network, N mini-batch state transition data (F t , G t , R t , F t+1 ) are required as the training data for training the Q-network and the policy network. Calculate the gradients of the cost functions of the policy network and the Q-network respectively to update the parameters of the policy network and the Q-network;
[0084] Among them, the cost function of the Q-network is:
[0085]
[0086] Among them, Indicates the target Q value, Q(F i , μ(F i , ω′), λ′) indicates the predicted Q value. The purpose of DDPG is to make the predicted Q value gradually approach the target Q value, and N represents the number of mini-batches extracted.
[0087] The definition of the target Q value is as follows:
[0088]
[0089] Among them, ψ represents the discount factor.
[0090] Therefore, the update method of the Q-network is:
[0091]
[0092] Among them, α c Indicates the learning rate for updating the Q-network.
[0093] The role of the policy network is to maximize the Q value. Therefore, the cost function of the policy network can be defined as:
[0094]
[0095] Taking the derivative of the cost function of the policy network gives:
[0096]
[0097] Therefore, the update method of the Q-network is:
[0098]
[0099] Among them, α aRepresents the learning rate for updating the policy network.
[0100] After updating the parameters of the Q-network and the policy network, it is necessary to update the parameters of the target Q-network and the target policy network every C steps. The update principle is as follows:
[0101] λ←τλ+(1 - τ)λ′
[0102] ω←τω+(1 - τ)ω′
[0103] Where τ is the update coefficient.
[0104] In each iteration cycle, when the algorithm converges or reaches the maximum number of iterations, the algorithm terminates. The position of the UAV, the computing offloading ratio, and the computing resource allocation scheme are obtained from the action with the optimal immediate reward.
[0105] In summary, the UAV-assisted terahertz communication network joint optimization method in this embodiment is aimed at the scenario of using UAVs to provide computing offloading services for users in the terahertz band. The DDPG algorithm is used to train the neural network to jointly optimize the UAV position, the computing offloading ratio, and the computing resource allocation scheme. Thus, on the premise of meeting the user service quality, the resource utilization rate and network capacity are effectively improved, and the total delay is reduced.
[0106] Second Embodiment
[0107] This embodiment provides a UAV-assisted terahertz communication network joint optimization device, including:
[0108] A communication network system model construction module for constructing a UAV-assisted terahertz communication network system model; wherein, in the communication network system model, the UAV carries a server to provide computing offloading services for users in the terahertz band;
[0109] An optimization objective function construction module for constructing an optimization objective function with the goal of minimizing the sum of the delays of all users in the communication network system under user service quality and resource constraints, based on the communication network system model constructed by the communication network system model construction module;
[0110] A joint optimization module for obtaining the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function constructed by the optimization objective function construction module based on a preset deep reinforcement learning algorithm, realizing the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieving the goal of improving network capacity and reducing delay.
[0111] The drone-assisted terahertz communication network joint optimization device of this embodiment corresponds to the drone-assisted terahertz communication network joint optimization method of the above first embodiment; among them, the functions realized by each functional module in the drone-assisted terahertz communication network joint optimization device correspond one by one to each process step in the above drone-assisted terahertz communication network joint optimization method; therefore, it will not be elaborated here.
[0112] Third Embodiment
[0113] This embodiment provides an electronic device, which includes a processor and a memory; among them, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the first embodiment.
[0114] This electronic device may have relatively large differences due to configuration or performance, and may include one or more processors (central processing units, CPUs) and one or more memories, among which, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0115] Fourth Embodiment
[0116] This embodiment provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the method of the above first embodiment. Among them, the computer-readable storage medium may be ROM, random access memory, CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.
[0117] In addition, it should be noted that the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0118] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate for implementation in the process Figure 1means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one box or more boxes.
[0119] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions in the process Figure 1 means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one box or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions in the process Figure 1 means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one box or more boxes.
[0120] It should also be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or terminal device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.
[0121] Finally, it should be noted that the above is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A joint optimization method for an unmanned aerial vehicle (UAV)-assisted terahertz communication network, characterized in that, it includes: Constructing a system model of a UAV-assisted terahertz communication network; wherein, in the system model, the UAV carries a server to provide computing offloading services for users in the terahertz band; Based on the system model, under the user service quality and resource constraints, with the goal of minimizing the sum of the delays of all users in the communication network system, an optimization objective function is constructed as: Among them, T i is the total delay of the i-th user, N is the number of users, x uav and y uav are the UAV coordinate information, α i is the offloading ratio of the i-th user, β i is the proportion of computing resources allocated to the i-th user, is the computing offloading vector, is the computing resource allocation vector, is the local computing energy consumption, is the uploading energy consumption, is the standby energy consumption of the user waiting for the server to process data, t i,max is the maximum tolerable delay of the i-th user, E i,max is the maximum tolerable energy consumption of the i-th user, is the set of users who cannot be served by the E-APs, is the preset coordinate threshold of the UAV; C1 represents that the total delay of each user does not exceed the maximum tolerable delay; C2 represents that the position of the UAV is within a preset specified range; C3 and C4 represent that the sum of the computing resources allocated to each user does not exceed the total computing resources; C5 represents that the user can offload any proportion of partial tasks to the server for processing; C6 represents that the energy consumed by the user is within the specified range; Based on a preset deep reinforcement learning algorithm, an optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function are obtained, realizing the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieving the purpose of improving network capacity and reducing delay, including: Taking the UAV, server, and all users as agents, the UAV-assisted terahertz communication network system model serves as the environment, and the UAV position, computing offloading ratio, and computing resource allocation scheme serve as the action outputs of the agents. The DDPG algorithm is used to train the agents to obtain an optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function, including: Step 1: Initialize the state space, action space, and deep neural network parameters of the system, including: modeling the user resource requirements and channel states as a finite state Markov model; creating two target neural networks μ′(F, ω′) and Q′(F, G, λ′) for each of the policy network μ(F, ω) and Q network Q(F, G, λ) for parameter update; Step 2: The agent selects and executes an action according to the current state and the policy network; Step 3: After the agent executes an action, return the reward and the new state, and put the state transformation process into the experience cache space, including: after the agent executes an action, determine whether the preset conditions are met. When the preset conditions are met, obtain the immediate reward according to the environment; the preset conditions include that the delay of each user meets the quality of service constraint, the position of the drone is within the specified range, the computing resources allocated to each user do not exceed the total resource amount, the computing offloading ratio is within the preset range, and the total energy consumption of each user meets the energy saving requirement; the immediate reward T n is the delay of the nth user; Step 4: Sample a preset number of state transition data in the experience cache space as the training data for training the Q network and the training policy network; Step 5: Calculate the gradients of the cost functions of the Q network and the policy network respectively; Step 6: Update the target neural network parameters.
2. The joint optimization method for a UAV-assisted terahertz communication network according to claim 1, characterized in that, In the communication network system model, the path loss PL(f, D) of the terahertz communication link between the server carried by the UAV and the user is expressed as: Among them, L abs (f, D) represents the molecular absorption loss, L spread (f, D) represents the transmission loss, D represents the distance between the user and the drone server, c is the speed of light in vacuum, k abs (f) is the medium absorption coefficient related to the frequency; f represents the terahertz carrier frequency.
3. The joint optimization method for a UAV-assisted terahertz communication network according to claim 1, characterized in that, The calculating the gradients of the cost functions of the Q network and the policy network respectively includes: Calculating the gradients of the cost functions of the Q network and the policy network respectively, and using the stochastic gradient descent method to update the neural network parameters.
4. A joint optimization device for a UAV-assisted terahertz communication network, characterized in that, it includes: A communication network system model construction module is used to construct a UAV-assisted terahertz communication network system model. In the system model, the UAV carries a server to provide computing offloading services for users in the terahertz band. An optimization objective function construction module is used to construct an optimization objective function based on the system model, with the goal of minimizing the sum of the delays of all users in the communication network system under user service quality and resource constraints, as follows: Among them, T i is the total delay of the i-th user, N is the number of users, x uav and y uav are the UAV coordinate information, α i is the offloading ratio of the i-th user, β i is the proportion of computing resources allocated to the i-th user, is the computing offloading vector, is the computing resource allocation vector, is the local computing energy consumption, is the uploading energy consumption, is the standby energy consumption of the user waiting for the server to process data, t i,max is the maximum tolerable delay of the i-th user, E i,max is the maximum tolerable energy consumption of the i-th user, is the set of users who cannot be served by the E-APs, is the preset coordinate threshold of the UAV; C1 represents that the total delay of each user does not exceed the maximum tolerable delay; C2 represents that the position of the UAV is within a preset range; C3 and C4 represent that the sum of the computing resources allocated to each user does not exceed the total computing resources; C5 represents that users can offload any proportion of their tasks to the server for processing; C6 represents that the energy consumed by users is within a specified range. A joint optimization module is used to obtain the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function based on a preset deep reinforcement learning algorithm, realizing the joint optimization of the UAV position, computing offloading ratio, and computing resource allocation scheme, and achieving the goal of improving network capacity and reducing delay, including: Regarding the UAV, server, and all users as agents, the UAV-assisted terahertz communication network system model serves as the environment, and the UAV position, computing offloading ratio, and computing resource allocation scheme serve as the action outputs of the agents. The DDPG algorithm is used to train the agents to obtain the optimal UAV position, computing offloading ratio, and computing resource allocation scheme that satisfy the optimization objective function, including: Step 1: Initialize the state space, action space, and deep neural network parameters of the system, including: modeling the user resource requirements and channel states as a finite-state Markov model; creating two target neural networks μ′(F, ω′) and Q′(F, G, λ′) for each of the policy network μ(F, ω) and Q network Q(F, G, λ) for parameter update. Step 2: The agent selects and executes an action based on the current state and the policy network. Step 3: After the agent executes an action, it returns the reward and the new state, and puts the state transformation process into the experience cache space, including: after the agent executes an action, it judges whether the preset conditions are met. When the preset conditions are met, an immediate reward is obtained according to the environment; the preset conditions include that the delay of each user meets the quality of service constraint, the position of the drone is within the specified range, the computing resources allocated to each user do not exceed the total resource amount, the computing offloading ratio is within the preset range, and the total energy consumption of each user meets the energy-saving requirements; immediate reward T n is the delay of the nth user; Step 4: Sample a preset number of state transition data from the experience cache space as the training data for training the Q network and the training policy network. Step 5: Calculate the gradients of the cost functions of the Q network and the policy network respectively. Step 6: Update the target neural network parameters.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle task unloading and resource allocation method for edge computing system
CN113395654A