Task Offloading Method Based on Ultra-Reliable Low-Latency Reinforcement Learning for Heterogeneous Communication Technologies

By employing heterogeneous communication technologies and DRL algorithms, the method addresses computing resource challenges in vehicle edge computing, achieving stable and reliable task offloading with low latency.

CN115118783BActive Publication Date: 2025-07-15JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210756389.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-15
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The difficulty in offloading computing-intensive tasks caused by insufficient vehicle computing resources in the Internet of Vehicles, and existing communication technologies are difficult to meet the requirements of ultra-reliability and low latency, especially in the case of high vehicle density, network performance optimization faces challenges.

Method used

Heterogeneous communication technology combined with deep reinforcement learning, by constructing vehicle edge computing scenarios, using random network calculations and Liyapunov optimization methods, establishing a delay upper bound model, optimizing task offloading and server CPU allocation strategies, ensuring system stability and low latency.

Benefits of technology

Effectively reduce system effectiveness, control base station queue stability, ensure task transmission delay requirements, and improve network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115118783B_ABST
    Figure CN115118783B_ABST
Patent Text Reader

Abstract

The present invention discloses a task offloading method based on ultra-reliable low-latency reinforcement learning for heterogeneous communication technologies, constructs a vehicle edge computing scenario and a vehicle heterogeneous communication network, and the vehicle can offload tasks to a server for processing through three communication technologies; constructs a dynamic change model of the base station queue to ensure the stability of the base station queue; uses stochastic network calculus theory to calculate the upper bound of the system delay for offloading based on different communication technologies, and this delay includes the communication transmission time and the server processing time; establishes the utility of the vehicle edge computing system; establishes an optimization problem, the optimization goal is to minimize the system utility while ensuring the task offloading delay and the stability of the base station queue; uses SoftActor Critic reinforcement learning to learn the offloading strategy for each task and the server CPU allocation strategy. The task offloading strategy and resource allocation scheme adopted by the present invention are superior to other offloading and resource allocation schemes in reducing the system utility, controlling the system stability, and ensuring the task transmission delay requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle networking edge computing reinforcement learning, and particularly relates to a task offloading method for ultra-reliable low-latency reinforcement learning based on heterogeneous communication technology. Background Art

[0002] The Internet of Vehicles (IoV) will develop rapidly in the future 5G era. The demand of vehicle users (VUs) for immersive quality of experience (QoE) and computationally intensive services (such as online 3D games, augmented and virtual reality (AR / VR), videos or other interactive applications) has increased sharply. In addition, for an autonomous vehicle, high-definition resolution cameras, LiDARS, high-speed high-definition maps and other in-vehicle sensors will generate 1GB of data per second. For a vehicle with insufficient in-vehicle computing power, it is a huge pressure to execute these computationally intensive tasks. To solve the problem of limited in-vehicle computing resources (such as CPU), vehicle edge computing (VEC) is considered a very promising technology to alleviate the tension of in-vehicle computing resources. VEC provides an open wireless network edge platform, enabling vehicles to offload computationally intensive task loads to nearby roadside MEC servers with lower latency. Although the bottleneck of insufficient in-vehicle computing resources can be alleviated to a certain extent by VEC technology, the emerging 5G applications and the ultra-reliable low-latency (URLLC) requirements for tasks still pose a certain pressure on the development of the Internet of Vehicles. The performance requirements related to ultra-reliable low latency, including supporting up to 1000 times the server data volume, ultra-low transmission latency below 5 ms and ultra-high reliability of 99.99%, on the one hand pose a huge challenge to a single communication technology, and on the other hand also test the reliability of VEC servers.

[0003] The emerging heterogeneous V2X communication technology has brought hope for the improvement of vehicle communication capabilities. Currently, there are three widely used technologies in vehicle networking, namely dedicated short-range communication (DSRC) communication, cellular vehicle-to-everything (C-V2X) communication, and millimeter-wave (mmWave) communication. DSRC enables vehicles to communicate over short distances without involving roadside units (RSUs). DSRC mainly operates in the 5.9 GHz frequency band and is based on the 802.11p standard protocol. C-V2X allows users to benefit from the existing extensive mobile communication infrastructure. In addition to operating at 5.9 GHz, C-V2X can also operate in licensed frequency bands of cellular carriers. However, research results show that in the case of high vehicle density, these two technologies cannot support reliable latency guarantees. The next-generation wireless technology, millimeter-wave, operates in a large unused spectrum (i.e., 3 - 300 GHz), can achieve multi-gigabit transmission capabilities for autonomous driving, and can also adapt to applications with high performance requirements. Heterogeneous V2X communication integrates the advantages of the three communication technologies, providing wide-area coverage and more efficient and reliable communication transmission for vehicles. However, due to the randomness of task generation and the time-varying channel conditions in the vehicle scenario, it will greatly affect the performance of vehicle edge computing task offloading, posing a challenge to network performance optimization. In recent years, deep reinforcement learning (DRL) has been widely applied to the strategy decision-making of task offloading in vehicle networking. DRL can adjust strategies to make optimal decisions to achieve the best long-term goals without any prior information about the vehicle environment.

[0004] Therefore, the present invention proposes a task offloading scheme for ultra-reliable low-latency reinforcement learning of heterogeneous communication technology. This new scheme considers the competition factors of task offloading communication bandwidth resources and server computing resources, and uses the stochastic network calculus (SNC) method based on the moment generating function (MGF) to obtain the delay upper bounds of mmWave, DSRC, and CV2I, thus ensuring low latency of task offloading. The offloaded tasks will cause an increase in the queue length at the server side, thus making the server unstable. Lyapunov optimization has been widely used to stabilize the queues of the system. This scheme uses Lyapunov technology to ensure the reliability of the system. In addition, based on deep reinforcement learning Soft actor-critic, under the conditions of ensuring latency and reliability, it learns the offloading strategy for each task and the allocation strategy of the server CPU, and can make optimal offloading and allocation decisions, reducing the overall system consumption utility and improving network and system performance. Summary of the Invention

[0005] Object of the Invention: The present invention proposes a task offloading method based on ultra-reliable low-latency reinforcement learning of heterogeneous communication technology, which is superior to other offloading and resource allocation methods in terms of reducing system utility, controlling the stability of the base station queue, and ensuring task transmission latency requirements.

[0006] Technical solution: A task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to the present invention includes the following steps:

[0007] (1) Construct a vehicle edge computing scenario, which consists of a base station connected to a server, multiple roadside units, and vehicles; construct a vehicle heterogeneous communication network composed of three communication technologies: millimeter wave, DSRC, and CV2I. Vehicles can offload tasks to the server for processing through the three communication technologies.

[0008] (2) Construct a bounded bursty traffic model based on the theory of stochastic network calculus.

[0009] (3) Construct a dynamic change model of the base station queue to ensure the stability of the base station queue.

[0010] (4) Based on the theory of stochastic network calculus, establish communication transmission models for the three communication technologies of millimeter wave, DSRC, and CV2I, and at the same time establish a computing processing model for the CPU; according to the series theorem, perform min-plus convolution on the communication transmission model and the computing processing model to obtain a system processing model.

[0011] (5) Deduce the upper bound of the delay probability for offloading and processing based on each communication technology; the delay includes the communication transmission time and the server computing processing time.

[0012] (6) Establish the utility of the vehicle edge computing system, which consists of communication utility and computing utility.

[0013] (7) Establish an optimization problem, with the optimization goal of minimizing the system utility while ensuring the task offloading delay and the stability of the base station queue.

[0014] (8) Use SoftActor Critic reinforcement learning to learn the offloading strategy for each task and the server CPU allocation strategy.

[0015] Further, the implementation process of step (2) is as follows:

[0016] Assume that the vehicle has K types of tasks to be processed. At the beginning of each t time slot, A i (t) is the amount of task data accumulated in queue i within the time interval [t, t + 1); at the same time, given a time interval 0 ≤ s ≤ t, define the binary non-cumulative quantity A i (s, t) = A i (t - s) = A i (t) - A i (s) is the cumulative task volume of the i-th arrival queue i, and A i (s, t) is a bounded bursty traffic model, satisfying a stationary non-negative stochastic process:

[0017] A i (s, t) = λ i [ρ i (t - s) + σ i (1)

[0018] Where ρ i is the task arrival rate, σ i is the task burst size, and both are constants. λ i satisfies the Poisson distribution, and λ i represents the number of vehicles that generate the i-th task within the time interval [s, t).

[0019] Furthermore, the implementation process of step (3) is as follows:

[0020] The queue length of the base station is expressed as:

[0021]

[0022] Where q i (t) is the queue length of the i-th task at the start time of time slot t, f E is the maximum CPU processing clock rate of the server, ω i represents the number of CPU clock cycles required for the server to process each bit of task i, and α i (t) represents the proportion of CPU clock cycles allocated by the server to the i-th task, and [x] + = max(x, 0); The stability of all queues is controlled by the following definition:

[0023]

[0024] The left end of equation (3) describes the long-term time-average backlog of the queue; Equation (3) means that the strong stability of the queue corresponds to a finite average backlog with a finite average queuing delay.

[0025] Furthermore, the implementation process of step (4) is as follows:

[0026] Let β mmw (s, t) represent the total communication transmission capacity that mmWave can provide within the time interval [s, t), let C (q) represent the channel capacity of time slot q, ζ (q) represent the channel gain of time slot q, represent the signal-to-noise ratio, B represents the bandwidth, l and δ represent the transmission distance and path loss exponent respectively, and the total communication transmission capacity that mmWave can provide is:

[0027]

[0028] where η = Blog2e, denotes the proportion of task i transmitted through millimeter - wave communication; the communication transmission volume that millimeter - wave can provide for the i - th task is:

[0029]

[0030] where

[0031] Use the delay - rate model in network calculus theory to establish the total communication transmission volume that DSRC communication can provide within the time interval [s, t):

[0032]

[0033] R dsrc is the DSRC communication bandwidth, denotes the average access delay when data is transmitted through DSRC; the communication transmission volume that DSRC can provide for the i - th task is:

[0034]

[0035] where denotes the proportion of task i transmitted through DSRC communication;

[0036] Use to denote the communication transmission volume that DSRC can provide for the i - th task in the time interval [s, t):

[0037]

[0038] where R cv2i is the communication bandwidth reserved for the i - th task, where

[0039] Use to denote the computing processing volume that the server CPU can provide for task i unloaded to the server within the time interval [s, t), and this computing processing volume is equal to the task volume processed by the CPU in formula (2):

[0040]

[0041] Let the set denote the communication technologies that can be unloaded; use to denote within the time interval [s, t) via the communication technology the task volume of task i unloaded to the server, denotes the proportion of task i transmitted through communication technology g, q i (s) represents the backlog tasks in base - station queue i that have not been processed before s moment, use Indicates the computing throughput provided by the server CPU during the time interval [s, t) to , The computable throughput is calculated as:

[0042]

[0043] where

[0044] The latency of the task is the sum of the communication transmission time and the server CPU processing time. For task i, the total processing that can be obtained by the system is the sum of the communication transmission volume and the CPU computing throughput. The overall service that the system can provide to task i offloaded based on communication technology g is the minimum-plus convolution of the communication transmission of communication technology g and CPU computing throughput :

[0045]

[0046] In the formula is the minimum-plus convolution operator, which is one of the most important operators in the theory of stochastic network calculus and has the following operation rules:

[0047]

[0048] Furthermore, the implementation process of step (5) is as follows:

[0049] Use W i g (t) to represent the latency of task i offloaded based on communication technology g. Use ω i g to represent the upper bound of the probability of task i transmitted through communication technology g. Represents the probability that the task transmission and processing time W i g (t) exceeds is less than ε i , defined as follows:

[0050]

[0051] Obtain the latency upper bound The solution of is:

[0052]

[0053] In the formula is:

[0054]

[0055] where To obtain the probabilistic latency upper bound of task i First, it is necessary to calculate

[0056]

[0057] Assume that and That is, the communication and computing resources that the offloaded task can obtain are much greater than the offloading rate of the task offloaded based on communication technology g. Obtain The closed-form solution of:

[0058]

[0059] Among them, and Upper bound of delay Is determined by many factors, and the probability ε beyond the delay bound i Determines the magnitude of the delay. Each term in the numerator of the second term is related to the task burst size, and the third term is jointly determined by the remaining resources of communication technology g and server computing and the task volume of task i offloaded based on communication technology g to determine the upper bound of the delay.

[0060] Furthermore, the implementation process of the step (7) is as follows:

[0061]

[0062] Among them, T i max Is the maximum transmission and processing delay requirement for the i-th task; the control variable α(t) = [α1(t), α2(t),....α N (t)] allocates the CPU clock cycle resources, Is the communication offloading strategy, among which Condition C1 is to make the queue in a stable state; condition C2 ensures that the transmission and processing time of each type of task is within the maximum delay requirement. Since the task is offloaded through three different communication technologies, and The maximum value among the three is used as the upper bound of the transmission delay of the i-th task; constraint C3 ensures that the CPU clock cycles used to process all tasks cannot exceed the total amount of available CPU computing resources on the server; constraint C4 ensures that each task selects mmWave, DSRC, or CV2I to execute the computing task;

[0063] The Lypunov technique is used to solve this long-term random constraint C1:

[0064] Define the second-order Lypunov function L(t) and the 1-slot Lypunov drift ΔLt :

[0065]

[0066]

[0067] where \(Q(t)=[q_1(t),q_2(t),\cdots,q N (t)]\); then, the desired system utility is added to the drift to obtain the drift-plus-penalty term, that is where \(V\) is a non-negative parameter set by the system, which is used to trade off between system utility and queue backlog; for any given control parameter \(V\geq0\) with respect to the offloading workload \(\alpha i under, the drift-plus-penalty term is obtained as:

[0068]

[0069] where the original time-averaged long-term queue length condition \(C1\) is implicitly absorbed into the optimization objective, and the optimization objective of problem \(P1\) is transformed into \(F2(t)\):

[0070]

[0071] Adopt a DRL framework including state, action, and reward to formulate the problem of computing resource allocation strategy and heterogeneous communication offloading strategy in VEC:

[0072] The state space \(s\) at time \(t\) t is:

[0073]

[0074] Since and both have dimensions of \(N\)-dimensional, while has a dimension of \(4N\)-dimensional, the dimension of the state space is \(5N\)-dimensional;

[0075] The action space \(a\) at time \(t\) t is:

[0076]

[0077] where \(\alpha i (t)\) and both need to satisfy the constraint conditions in formula (30) and Add a dummy variable \(\alpha N+1 (t)\) to the deep neural network that outputs \(N + 1\)-dimensional actions, and then use the softmax function for these \(N + 1\)-dimensional variables in the output layer to satisfy:

[0078]

[0079] Only take the first N actions; similarly, for each task i, the output action and Using the softmax function, we can achieve In this way, the action space is N+1+3N dimensional, and the dimension of the action space increases as the number of task types increases; similarly, by outputting the action of each task and Use the softmax function to achieve So the action space is 4N+1 dimensional;

[0080] The reward function r at time t t for:

[0081] r t (a t ,s t )=-F2(t) (38)

[0082] r t (a t ,s t ) indicates that in state s t Take action t After that, the environment gives the agent a reward feedback, expressed as π(a t |s t ) indicates that the agent is based on state s t Given the spatial distribution of actions taken, the expected long-term discounted return of the system is calculated as:

[0083]

[0084] Where γ∈[0,1] is a discount factor representing the agent’s focus on long-term or short-term rewards. The higher the value, the more the agent focuses on long-term rewards, and vice versa. a ,…) is the agent’s dependence on the action space distribution π(a t |s t )’s status and behavior trajectory.

[0085] Furthermore, the implementation process of step (8) is as follows:

[0086] The SAC algorithm introduces policy entropy while maximizing the optimization objective - F2(t) into the reward, the expected long-term discounted reward of the model is

[0087]

[0088] β tis the weight of the policy entropy, which balances between exploring feasible policies and maximizing the optimization objective; as the reward changes continuously, a fixed β t will affect the stability of the entire training, so it is very necessary to automatically adjust β during the training process t The optimization problem of reinforcement learning is transformed into:

[0089]

[0090] By setting a lower bound is to make β in formula (40) t H(π t (·|s t )) as large as possible; when the agent has not learned the best action, increase β t to explore more total space; conversely, if the best policy has been learned, then reduce β t to reduce exploration and accelerate the training of the model; Based on the Lagrange multiplier method, we can obtain

[0091]

[0092] Beneficial effects: Compared with the prior art, the beneficial effects of the present invention are as follows: The task offloading strategy and resource allocation scheme adopted by the present invention are superior to other offloading and resource allocation schemes in terms of reducing system utility, controlling the stability of the base station queue, and ensuring the task transmission delay requirements BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 is a scenario diagram of task offloading in a heterogeneous network vehicle edge computing

[0094] Figure 2 is a framework diagram of the server queue model

[0095] Figure 3 is a relationship diagram of the number of task types, task arrival rate, and upper bound of delay in the solution of the present invention

[0096] Figure 4 is a relationship diagram of the number of task types, task burstiness, and upper bound of delay in the solution of the present invention

[0097] Figure 5 is a relationship diagram of server computing resources, heterogeneous communication technologies, and delay violation probability in the solution of the present invention

[0098] Figure 6 is a relationship diagram of the number of task types, heterogeneous communication technologies, and delay violation probability in the solution of the present invention

[0099] Figure 7Relationship diagram of communication bandwidth resources, heterogeneous communication technologies, and latency violation probability for the solution of the present invention;

[0100] Figure 8 CCDF relationship diagram of tasks exceeding latency for the solution of the present invention;

[0101] Figure 9 Queue backlog comparison diagram of the solution of the present invention with the average allocation and average offloading strategy, random allocation and random offloading, and heterogeneous communication allocation strategy;

[0102] Figure 10 System utility comparison diagram of the solution of the present invention with the average allocation and average offloading strategy, random allocation and random offloading, and heterogeneous communication allocation strategy. Detailed implementation manners

[0103] The present invention will be further described in detail below with reference to the accompanying drawings.

[0104] In a vehicle edge computing network environment, aiming at the requirements of the 5G era for higher data rates, ultra-low latency, high reliability, and excellent user experience in vehicle networking applications, the present invention proposes a task offloading method based on heterogeneous communication technology ultra-reliable low-latency reinforcement learning, and a task offloading solution with ultra-reliable low-latency based on three heterogeneous communication technologies, namely millimeter wave, DSRC, and CV2I, to improve the network performance of vehicle edge computing. First, a vehicle edge computing task offloading model is proposed. Vehicle users can select the proportion of tasks offloaded by each communication technology among millimeter wave, DSRC, and CV2I to offload tasks to the edge server, thereby ensuring low latency of task offloading. And the server allocates CPU resources based on Lypunov technology to ensure the reliability of the base station queue. Secondly, the random network calculus theory is used to calculate the upper bound of the latency for offloading and processing based on different communication technologies, and this latency includes the communication transmission time and the server processing time. Finally, SoftActorCritic reinforcement learning is used to learn the offloading strategy for each task and the server CPU allocation strategy.

[0105] To ensure the stability of the base station queue, the task processing of a single task generated by a single vehicle is not considered. The total time is divided into T equal time slots, the interval of each time slot is △t, and the T time slots are represented by the set Indication. Considering the mobility of the vehicle, the channel state is time-varying. In the solution, it is assumed that the channel state information (CSI), the distances between the vehicle and the base station, and the vehicle and the RSU do not change within a time slot, but are different in different time slots. Consider that there are K types of tasks that the vehicle needs to process in the scenario, and task i represents the task of type i (i = 1, …, N). Assume that all vehicles can directly offload tasks to the cloud server connected to the base station for processing through three communication methods: millimeter wave, DSRC, and CV2I, or first offload the tasks to the RSU and then transmit the tasks to the cloud server connected to the base station through a wired method for processing. Use the set to represent the communication technologies that can be offloaded. When the communication technology is g, at this time, represents the offloading ratio of task i using communication technology g, where Since the time delay of wired transmission is small, the transmission time between the RSU and the base station is not considered in this solution. The base station has N queues. Assume that each queue is infinitely long. Queue i only stores task i. Use λ i to represent the number of vehicles generating task i within the time interval [s, t). The cloud server processes the tasks stored in the base station queues. Since the server is close to the base station, the transmission time from the base station to the server is ignored. The goal of this invention is to minimize the delay, including the task communication time and processing time, while meeting the requirements of power consumption, cost effectiveness, and queue stability. Specifically, it includes the following steps:

[0106] Step 1: Construct a vehicle edge computing scenario, as Figure 1 shown. The scenario consists of a base station connected to a server, multiple roadside units, and vehicles; construct a vehicle heterogeneous communication network composed of three communication technologies: millimeter wave, DSRC, and CV2I. Vehicles can offload tasks to the server for processing through the three communication technologies.

[0107] Step 2: Construct a bounded bursty traffic model based on the random network calculus theory; construct a dynamic change model of the base station queue to ensure the stability of the base station queue, as Figure 2 shown.

[0108] At the beginning of each t time slot, A i (t) is the cumulative amount of task data arriving at queue i within the time interval [t, t + 1). At the same time, given a time interval 0 ≤ s ≤ t, define the binary non-cumulative quantity A i (s, t) = A i (t - s) = A i (t) - A i as the cumulative task amount of the i-th arriving at queue i. Assume A i(s, t) is a bounded bursty model that satisfies a stationary non - negative stochastic process:

[0109] A i (s, t) = λ i [ρ i (t - s)+σ i (1)

[0110] where ρ i is the task arrival rate, σ i is the task burst size, both of which are constants, and λ i satisfies the Poisson distribution.

[0111] Consider that the server has N E CPU processing cores. The maximum CPU processing clock rate of the server is f E cycles (frequency) per second. And tasks of different application types that process each bit of data volume require different amounts of CPU clock cycles of the server's resources, that is, processing each bit of task i's data volume requires ω i units of CPU clock cycles. The server allocates α i (t) proportion of the CPU clock cycle resources to the i - th task, which is represented by the set α(t)=[α1(t),α2(t),...α3(t)]. Each element in the set needs to be in the feasible set A, that is In this scheme, the Lindley recursion is used to analyze the dynamic change of the queue length. Therefore, the queue length at the base station can be expressed as:

[0112]

[0113] where q i (t) is the queue length of the i - th task at the start time of time slot t, and [x] + = max(x, 0). The reliability of the server is achieved by ensuring the stability of each queue. Consider controlling the stability of all queues through the following definition:

[0114]

[0115] The left - hand side of the above formula describes the long - term time - average backlog of the queue. The above formula means that the strong stability of the queue corresponds to a finite average backlog with a finite average queuing delay.

[0116] Step 3: Based on the theory of stochastic network calculus, establish the communication transmission models for three communication technologies, namely millimeter wave, DSRC, and CV2I, and at the same time establish the computing and processing model of the CPU; according to the tandem theorem, perform min-plus convolution on the communication transmission model and the computing and processing model to obtain the system processing model; derive the upper bound of the delay probability for offloading and processing based on each communication technology; the delay includes the communication transmission time and the server computing and processing time.

[0117] Due to the propagation characteristics of the millimeter wave band, the small-scale fading in the millimeter wave channel is very weak, so the amplitude of the channel coefficient (in the millimeter wave band) is usually modeled as a random variable that follows the Nakagami-m distribution. For a given transmission distance l and path loss exponent δ, according to Shannon's formula, attempt to calculate the channel capacity of the millimeter wave band

[0118]

[0119] where represents the signal-to-noise ratio, B represents the bandwidth, the random variable ζ is the channel gain, which is independent and identically distributed with respect to time and follows the gamma distribution, i.e., ζ ∼ Γ(M, M -1 ), M is the Nakagami exponent, then the probability density function (p.d.f) of ζ is Use β mmw (s, t) to represent the total communication transmission volume that mmWave can provide in the time interval [s, t), and use C (q) to represent the channel capacity in the q time slot, and ζ (q) to represent the channel gain in the q time slot. Then according to the literature, the communication transmission volume of millimeter wave is

[0120]

[0121] where η = Blog2e, since the channel gain coefficients are independent and identically distributed Equation (5) can be further written as

[0122]

[0123] Use to represent the task volume of task i offloaded to the server via communication technology g in the time interval [s, t). According to the formula Since the offloaded through mmWave needs to compete for the millimeter wave communication bandwidth resources with other types of tasks. According to the remaining service theorem in stochastic network calculus, the mmWave communication transmission volume that can be obtained in the time interval [s, t) is the total mmWave communication transmission volume β in the time interval [s, t) mmw(s, t) subtracts the task volumes of all mmWave offloading-based tasks j ≠ i represents the proportion of task i that communicates and transmits via millimeter wave. Therefore, for task i, the communication transmission volume that millimeter wave can provide is

[0124]

[0125] Since the communication transmission volume provided by millimeter wave is usually greater than the data volume transmitted by the task, + is greater than 0 inside, so formula (7) can be sorted out as

[0126]

[0127] Let Formula (6) can be simplified to

[0128]

[0129] DSRC communication is based on the IEEE802.11p standard protocol. In the 802.11p protocol, the basic access mode of wireless communication is based on the distributed coordination function, that is, the retransmission of data packets after a collision adopts the exponential backoff algorithm. Therefore, the delay of DSRC communication is mainly the access delay after a collision, denoted by to represent the average access delay of collisions when transmitting data via DSRC. According to the literature u is a constant, is the Parato type tail exponent, R dsrc is the DSRC communication bandwidth. Let β dsrc (s, t) represents all the communication transmission volumes that DSRC can provide in the time interval [s, t). According to the network calculus theory, this type of communication transmission volume model can be established as a latency-rate service model, expressed as:

[0130]

[0131] Similarly to the method of establishing the millimeter wave communication transmission volume model, by the remaining service theorem, it can be obtained that for task i, the communication transmission volume that DSRC can provide is:

[0132]

[0133] Let The above formula can be simplified to

[0134]

[0135] CV2I is a communication technology based on cellular networks. In the 14th version of 3GPP, CV2X communication was standardized, and there are two scheduling-based modes: Mode 3 and Mode 4. Both of these modes are pre-allocated communication resources. Therefore, assuming that the basis of CV2I is the bandwidth reservation mode, that is, the bandwidth resources for communication are pre-scheduled and allocated by the base station to the i-th task. That is, the task i unloaded via CV2I does not need to compete for communication throughput with other tasks j that are also unloaded based on this communication technology. The cumulative communication throughput that the entire CV2I can provide to the i-th task in the time interval [s, t) is

[0136]

[0137] where R cv2i is the communication bandwidth reserved for the task. To facilitate the derivation of the upper bound of the calculation delay, Equation (13) is put into the same form as Equation (9) and Equation (12). Therefore, for task i, the communication throughput that CV2I can provide in the time interval [s, t) can be obtained as:

[0138]

[0139] It should be noted that here This is because the model based on CV2I reserves bandwidth resources for the transmitted tasks, and there is no competition for communication bandwidth resources between tasks.

[0140] Next, a CPU computing throughput model is established to obtain the system throughput model. Then, based on the established traffic model and system throughput model, the upper bound of the delay probability of task offloading and server processing based on each communication technology is derived using stochastic network calculus based on the moment generating function (MGF). Let represent the computing throughput that the server CPU can provide to task i unloaded to the server in the time interval [s, t). This computing throughput is equal to the amount of tasks processed by the CPU in Equation (2):

[0141]

[0142] Let represent the computing throughput that the server CPU provides to in the time interval [s, t). It should be noted that only processes the computing throughput of task i. Therefore, according to the residual service theorem in stochastic network calculus, needs to compete for the computing throughput allocated by the CPU with and the backlogged tasks q i (s) that have not been processed in base station queue i So the computing throughput that can be obtained can be calculated as:

[0143]

[0144] Let Then the above formula can be simplified to:

[0145]

[0146] Define the delay of the task as the time of communication transmission and the time of server CPU processing. Therefore, for task i, the total processing capacity that can be obtained by the system is the communication transmission volume and the computational processing volume of CPU processing. According to the tandem theorem in the theory of stochastic network calculus, the total processing capacity that the system can provide for task i unloaded based on communication technology g is the minimum-plus convolution of the communication technology g transmission and CPU computing processing :

[0147]

[0148] In the formula is the minimum-plus convolution operator, which is one of the most important operators in the theory of stochastic network calculus and has the following operation rules:

[0149]

[0150] Stochastic network calculus overcomes the problem that the deterministic envelope only considers the worst case, allows violating the envelope with a certain small probability, and makes full use of the statistical characteristics of the arriving data stream. The probabilistic delay upper bound represents that the probability that the task transmission and processing time W i g (t) exceeds is less than ε i , and is defined as follows:

[0151]

[0152] The inequality in the above formula is based on the Chernoff inequality P(x≥X)≤e -θX E[e θx Also define the moment generating function of x, that is, M x (θ) = E[e θx , then formula (20) can be converted to

[0153]

[0154] where is a binary function with respect to the positive parameter θ and . In order to obtain a more compact probabilistic upper bound, take the minimum as the equivalent violation probability, that is Then the inequality in the above formula can be transformed into

[0155]

[0156] Solving equation (22) can obtain the upper bound of delay. The solution of

[0157]

[0158] In the formula, can be calculated as:

[0159]

[0160] where To obtain the delay of task i, first, it is necessary to calculate Since the communication transmission volume model and CPU computing processing volume model of different communication technologies have been formally unified as described above, that is serve ∈ {comp, comm}, it is convenient to carry out formula derivation. The final result is as follows:

[0161]

[0162] Suppose and That is, the communication and computing resources that the offloaded tasks can obtain are much larger than the offloading rate of the tasks offloaded based on communication technology g. Substituting the result of equation (25) into equation (23) can obtain the closed-form solution of

[0163]

[0164] where and It can be seen from the above formula that the upper bound of delay is determined by many factors. The first term in the above formula indicates that the probability ε i beyond the delay bound determines the size of the delay, the terms in the numerator of the second term are all related to the task burst size, and the third term is jointly determined by communication technology g, the remaining resources of the server computing, and the task volume of task i offloaded based on communication technology g.

[0165] Step 4: Establish the utility of the vehicle edge computing system. The system utility consists of communication utility and computing utility. The utility of the system consists of two parts: communication utility and computing utility.

[0166] Communication utility: It is only assumed that the telecom operator charges for CV2I based on cellular communication, while assuming that DSRC and millimeter waves operate in unlicensed frequency bands. At the same time, the unit cost of transmitting each bit of data based on CV2I communication is defined as c. comm , and the amount of tasks offloaded through CV2I is Therefore, the communication utility F comm,i (t) generated by the system due to the allocation of the i-th task is:

[0167]

[0168] Calculating utility: The computing utility of the system refers to the power cost defined as generated by the server for processing tasks. The unit cost of power consumed by the server for each processing is c comp . To establish a more realistic computing environment, the Dynamic Voltage Frequency Scaling (DVFS) method is adopted to simulate the CPU power consumption. DVFS enables the system to operate at a lower frequency and correspondingly lower voltage under low-load or highly memory-constrained workloads, saving energy consumption with almost no loss of performance requirements. The server has allocated a total of CPU clock cycle resources for processing the tasks in the queue. Based on the DVFS assumption, the dynamic frequency of each CPU is Generally, the CPU power consumption is calculated as the cube of the frequency, so the power consumption of each CPU is where κ represents the effective switching capacitance parameter. Using F comp,i (t) to represent the computing utility of the CPU at time t:

[0169]

[0170] Finally, the utility function of the system can be obtained as:

[0171]

[0172] where and are the normalized weighting coefficients respectively to ensure that the orders of magnitude of the CPU computing utility and communication utility are consistent.

[0173] Step 5: Establish an optimization problem, where the optimization objective is to minimize the system utility while ensuring the task offloading delay and the stability of the base station queue.

[0174] Based on the obtained offloading processing delay and system utility function before, construct the heterogeneous communication offloading strategy and the optimal CPU resource allocation problem. The optimization goal is to minimize the average delay of all tasks in the model while satisfying the power consumption and communication cost constraints as well as the requirements of queue stability. The proposed optimal problem is as follows:

[0175]

[0176] where \(T\) i max is the maximum transmission and processing delay requirement for the \(i\)-th task. The control variable \(\alpha(t)=[\alpha_1(t),\alpha_2(t),....\alpha\) N (t)] allocates the CPU clock cycle resources, is the communication offloading strategy, where In the above constraints, condition C1 is to keep the queue in a stable state, and condition C2 ensures that the transmission and processing time of each type of task is within the maximum delay requirement. Since the tasks are offloaded through three different communication technologies, and the maximum value among the three is used as the upper bound of the transmission delay for the \(i\)-th task. Constraint (C3) ensures that the CPU clock cycles used to process all tasks cannot exceed the total available CPU computing resources on the server. Constraint (C4) ensures that each task selects mmWave, DSRC, or CV2I to execute the computing task.

[0177] It is not easy to solve P1 to obtain the optimal transmission offloading strategy and CPU resource allocation strategy. The main reason is that constraint condition (C1) is about the stability of the long-term time-average queue length, which has a great impact on the long-term stability of the short-term decision queue, and it is more desirable to make decisions without considering future information. Therefore, the Lypunov technique will be used first to solve this long-term stochastic constraint (C1).

[0178] The Lypunov function is an efficient framework for designing online control algorithms without any prior knowledge. Define the second-order Lypunov function \(L(t)\) and the 1-slot Lypunov drift \(\Delta L\) t :

[0179]

[0180]

[0181] where \(Q(t)=[q_1(t),q_2(t),...q\) N (t)]. Then, add the expected system utility to the drift to obtain the drift-plus-penalty term, that is where V is a non - negative parameter set by the system, used to trade - off between system utility and queue backlog. For any given control parameter V≥0, with respect to the offloading workload α i the drift - plus - penalty term can be obtained.

[0182]

[0183] where the original time - averaged long - term queue - length condition C1 is implicitly absorbed into the optimization objective. The optimization objective of problem P1 is transformed into F2(t):

[0184]

[0185] Although the original problem is simplified by means of Lyapunov's method, directly solving this problem is still far from easy. On the one hand, the optimization problem P2 is not a convex problem. On the other hand, due to the complexity of the vehicle environment, the solution of this problem is troubled by the curse of dimensionality. Therefore, a DRL framework including states, actions, and rewards is adopted to formulate the problem of computing resource allocation strategy and heterogeneous communication offloading strategy in VEC.

[0186] State space s t : Vehicles need to observe network resources and computing resources to determine the offloading rate of heterogeneous communication, while the server allocates CPU clock - cycle resources for each type of task arriving at the base station by observing the queue length. In this model, the size of the task is the most fundamental factor affecting task latency and queue length. The randomness of task arrival will affect system stability, task transmission latency, and cost. Therefore, first, the variable reflecting the data volume is taken as one of the states, that is the randomness of the task volume is reflected by the number of tasks λ i In addition, in order to achieve system stability, the queue length affects the server's CPU resource decision. Equation (2) shows the queue - length update method. Since A i (t) is a random variable, in the field of deep learning, it is known that an environment with random rewards is more difficult to learn than an environment with deterministic rewards. Using as the state of the queue length at time t, the randomness of A i (t) can be absorbed into the transition probability s t+1 ~P(s t+1 |s t ,a t ). In this way, a deterministic reward can be obtained from the environment. Therefore, using Q t =[q1(t)+A1(t),q2(t)+A2(t),...q N (t)+A N(t) represents the queue length status of all tasks. Another aspect to be considered is the The closed-form solution is determined by and which both reflect the available resources when task i competes with task j. Equation (26) shows that the latency is largely dominated by the relatively small remaining resources. Therefore, the available communication and computing resources of task i are also incorporated into the state, that is Here, only the case of DSRC is considered for the available communication resources because in CV2I, it is based on a bandwidth reservation model and there is no competition between tasks, and the abundance of millimeter-wave band resources does not need to be considered. Thus, the resource status of all tasks can be represented as ξ t =[ξ1(t), ξ2(t),... ξ N (t)]. Therefore, the state s at time t t can be defined as:

[0187]

[0188] Since and both have dimensions of N, while has a dimension of 4N, the dimension of the state space is 5N.

[0189] The action space a t : For each type of each type of task, allocate CPU clock cycle resources to process task i in the queue, α(t)=[α1(t), α2(t),.... α N (t)] determines how many tasks are offloaded through millimeter-wave, DSCR, and CV2I methods. Therefore, the action at time t is defined as:

[0190]

[0191] where α i (t) and both need to satisfy the constraints in Equation (30) and To more easily implement α i (t) at the output of the neural network to satisfy by adding a dummy variable α N+1 (t), a deep neural network that outputs N+1-dimensional actions, and then uses the softmax function on these N+1-dimensional variables in the output layer to satisfy:

[0192]

[0193] Only take the actions of the top N. Similarly, for the output actions of each task i and Using the softmax function, it can be achieved In this way, the action space is N + 1 + 3N dimensional, and the dimensionality of the action space increases with the increase in the number of task types. Similarly, by applying the softmax function to the output actions of each task and Using the softmax function, thus achieving So the action space is 4N + 1 dimensional.

[0194] Reward function r t : In this paper, the system aims to improve performance in terms of utility, stability, and latency. Since in reinforcement learning, the long-term reward is maximized while we want to minimize the optimization objective, the reward function of the model at time t is defined as the negative value of the optimization objective:

[0195] r t (a t , s t ) = -F2(t)(38)

[0196] r t (a t , s t ) indicates the reward feedback of the environment to the agent after taking action a t in state s t . Let π(a t | s t ) represent the action space distribution that the agent can take based on state s t . The expected long-term discounted return of this system is calculated as:

[0197]

[0198] where γ ∈ [0, 1] is the discount factor representing the agent's focus on long-term or short-term rewards. A higher value indicates that the agent pays more attention to long-term rewards, and vice versa for short-term rewards. τ = (s0, a0, s1, a a , …) is the state and action trajectory of the agent depending on the action space distribution π(a t | s t ).

[0199] Step 6: Use SoftActor Critic reinforcement learning to learn the offloading strategy for each task and the server CPU allocation strategy.

[0200] All the values in the state space and action space are continuous variables, while general reinforcement learning methods can only solve discrete variables and low-dimensional variables. For control tasks in high-dimensional, continuous state and action spaces, neural networks are used to approximate the variable values in the space. The Soft-Actor-Critic (SAC) algorithm is a reinforcement learning algorithm suitable for solving continuous state and action spaces. It introduces policy entropy into the reward, maximizing the reward while encouraging the agent to explore more feasible policies. Therefore, SAC has better robustness and stronger generalization ability.

[0201] The SAC algorithm introduces policy entropy while maximizing the optimization objective - F2(t) into the reward, and the expected long-term discounted reward of the model is

[0202]

[0203] β t is the weight of the policy entropy, which balances between exploring feasible policies and maximizing the optimization objective. As the reward changes continuously, a fixed β t will affect the stability of the entire training, so it is very necessary to automatically adjust β t during the training process. The optimization problem of reinforcement learning can be transformed into:

[0204]

[0205] By setting a lower bound to make β t H(π t (|s t )) in formula (40) as large as possible. When the agent has not learned the best action, increase β t to explore more total space; conversely, if the best policy has been learned, then reduce β t to reduce exploration and accelerate the training of the model. Based on the Lagrange multiplier method, we can obtain

[0206]

[0207] where is the best policy that SAC has learned to maximize the expected long-term discounted reward.

[0208] The reinforcement learning SAC training algorithm and test algorithm based on heterogeneous communication technology for low-latency and ultra-reliable task offloading are as follows:

[0209] Training phase algorithm:

[0210] The SAC algorithm is based on the actor-Critic network framework. The actor network is used for policy optimization, and the critic network is used to evaluate the policy. By continuously performing policy optimization and policy evaluation, the final policy π will converge to the optimal policy π * The actor network outputs the mean μ φ (t) and covariance ∑ φ (t) of the high-dimensional Gaussian distribution of the policy, where φ represents the neural network parameters of the actor network; the actor network samples from the high-dimensional Gaussian distribution of the policy and outputs the high-dimensional action a at the current state t The critic network outputs the approximate Q θ (s t ,a t ) for policy evaluation. Q θ (s t ,a t ) represents the action value function in state s t under action a t . Q θ (s t ,a t ) represents the expectation of the sum of future discounted rewards under the condition that action a t is selected at the current time t and the optimal action is taken afterwards:

[0211]

[0212] where θ represents the neural network parameters of the critic network. The SAC algorithm introduces two critic networks with the same network structure to ensure a reduction in the overestimation of Q θ (s t ,a t ), that is, it outputs the approximations and respectively. θ1 and θ2 represent the parameters of the two critic networks. In addition, to train more quickly and stably, two target critic networks with the same structure as the critic network are introduced, and their network parameters are and Next, the algorithm will be described in detail.

[0213] To study the trade-off between the utility and stability of the system under different values of the weighting factor V, it is necessary to train the SAC model under different values of V. First, initialize the network parameters φ, θ1, and θ2, while and are initialized to the same values as θ1 and θ2. A replay buffer with a sufficient spatial size is constructed Used to save the data collected during the training process. To reduce the impact of randomness on the stability of the SAC algorithm, the same millimeter-wave channel coefficient ζ is set. A total of K max rounds of training are carried out. At the beginning of each round, all queues are emptied, and the initial state ξ0 is set to the initial computing and communication resources, that is initially set to R dsrc , and and are both set to f E , while other initial states are set to 0.

[0214] Each round contains T max time steps. The algorithm starts from time step t = 0 to t = T max and begins. First, for each type of task, a random number of vehicles λ i is generated. Let the time slot be 1, that is, t - s = 1. According to A i (0) = λ i [ρ i (t - s) + σ i , the task volume arriving at the base station queue for each type of task is calculated, and thus the state and the state are obtained. In this way, a complete state s0 at t = 0 is obtained. The state s0 is sent into the actor network, and the mean μ φ (0) and variance ∑ φ (0) of the high-dimensional Gaussian distribution of the policy at t = 0 are output. Then, a sample is drawn from the Gaussian distribution to obtain an action space a0 containing N + 1 dimensions of α N+1 (0) and 3 * N dimensions of . A softmax operation is performed on α N+1 (0) to make it satisfy formula (37) and the first N values are taken to obtain the CPU offloading strategy. A softmax operation is performed on to make it satisfy . According to the obtained action a0 and formula (1), the length Q(0) of all queues is updated; based on formula (26), the delay of each task transmitted by different communication technologies is obtained . At the same time, the state ξ1 at the next moment is obtained. When the vehicle and the server take actions, feedback r0 is obtained from the environment. It should be noted that in order to obtain the of the state at the next moment, α(1) generated at t = 0 is required. Only then can the complete state at t = 1 be obtained. Then, the tuple (s0, a0, r0, s1) is stored in the replay buffer. When the number of samples collected in the replay buffer is less than At this time, the vehicle and the server will continue to send the next state s2 into the actor network to start the iteration of the next time step t = 2. When the number of samples in the replay buffer reaches At this time, the parameters φ, θ1, θ2 of the actor network, two critic networks, and two target critic networks are iteratively updated and and the relative entropy weight β to maximize the objective function J(π t ) of Equation (40). The model first randomly extracts a tuple (s t , a t , r t , s t+1 ) with a quantity of I from the replay buffer to form a mini-batch of data samples. All s t in the mini-batch of data samples are sent into the actor network to obtain a high-dimensional Gaussian distribution of the policy. According to the policy distribution, the action a new and the policy entropy The policy entropy weight β is updated according to the gradient of the policy entropy weight, that is, updated along the gradient The direction can reach the optimal Regarding it as the gradient of the relative entropy weight factor β is:

[0215]

[0216] Next, given the state s t and the action a new , the two critic networks output the state-value functions and The loss function of the actor network can be calculated as:

[0217]

[0218] Here, for the stability of training, In order to make the actor network differentiable, at the same time, the reparameterization trick technology for the action is used, that is, a t = f φ (ε t ; s t ), ε t is an input noise vector sampled from some fixed distribution. Simply put, first sample from the Gaussian distribution, multiply the sampled value by the covariance and then add the mean to get the final output action a t .

[0219] To obtain the loss function of the critic network, it is necessary to use s t and a tInput into two critic networks respectively to obtain the action-state value function based on state s t and action a t of the action-state value function and Send s in the tuple t+1 into the actor network to obtain the new policy entropy π' φ and sample to obtain action a next , and the two target critic networks will be based on s t+1 and a next to obtain the target action-state value function and Then the target value is:

[0220]

[0221] Then the loss functions of the critic networks are respectively:

[0222]

[0223]

[0224] Next, use the Adam optimizer to perform K u rounds of parameter updates. The relative entropy weight factor β and the actor network adopt the same learning rate of α A , and the learning rates of the two critic networks are α C . Finally, after K u rounds of updates, update the parameters of the two target critic networks:

[0225]

[0226]

[0227] where τ is a constant satisfying τ << 1.

[0228] After training, start the next time step. When t = T max , empty the queue and initialize the state again to start the next round. When reaching the maximum number of rounds K max , an optimal relative entropy weight factor β * , actor network parameters, critic network parameters, and target critic network parameters will be obtained, that is, an optimal policy will be obtained

[0229] Figure 3 Describes the upper bound of delay with respect to the arrival rate ρ under different numbers of task types iincreases with the increase. As the task arrival rate increases, the latency performance of C-V2I-based offloading is better than the other two offloading methods. The reason is that when transmitting through C-V2I, there is no competition for communication resources among tasks. Only when increasing the task arrival rate ρ i will it have a greater impact on the latency of C-V2I. This is because as the arrival rate increases, the server CPU cannot process the tasks in the base station in time, resulting in the increase of formula (14). From Figure 3 it can be seen that at low arrival rates, the latency of millimeter wave is even lower than that of C-V2I because the bandwidth of millimeter wave is very large. Another notable phenomenon in the figure is that when the task volume is relatively small (N = 5), there is a decreasing stage in the latency upper bounds of millimeter wave and CV2I. The reason for this result is that, as revealed by formula (26), at the beginning, the available communication resources for the offloaded tasks are greater than the available CPU computing resources, and the main factor affecting the latency is Then as the task arrival rate ρ i increases, the main determinant of task latency becomes

[0230] Figure 4 shows that the latency upper bound increases with the burstiness σ i In this simulation, the arrival rate ρ i of all tasks is set to 0.5 Mbps. When the burstiness σ i increases to about 5 Mbps, Figure 4 it can be seen from i that the burstiness σ

[0231] Figure 5 shows the violation probability of the latency upper bounds of different offloading communication technologies under different server CPU resources. The number of task categories is N = 5. From Figure 5 it can be clearly seen that increasing the CPU cycle resources can significantly reduce the probability latency. In addition, it is also observed that in the low-traffic task scenario (ρ = 0.5 Mbps), the performance of millimeter wave communication is better than the other two offloading communication technologies. The main reason is that in the low-traffic scenario, millimeter wave has a huge bandwidth advantage compared with the bandwidth reserved for tasks by C-V2I.

[0232] Figure 6 studied the influence of the number of different task categories N on the latency performance of different heterogeneous communication technologies. In this simulation, the arrival rate of all tasks is set to ρ = 0.8 Mbps, and the overall CPU processing capacity f E = 10 4GHz. From Figure 8 It can be seen that when N = 10, the probability delay of millimeter wave exceeds that of C-V2I, which is consistent with Figure 3 and Figure 4 indicating that the delay performance of millimeter wave is inferior to that of C-V2I under high traffic. On the other hand, as the traffic increases, the network performance of DSRC shows a trend of rapid deterioration.

[0233] Figure 7 shows the upper bound ε of the random delay probability of various communication resources of millimeter wave, DSRC and C-V2I. In this simulation, the same basic parameter settings are adopted for all communication technologies, that is, the burstiness σ and arrival rate ρ are set to 0 Mbps and 0.5 Mbps respectively, and the CPU cycle resources all have f E = 10 4 GHz, and there are N = 5 task offloads for each communication. From Figure 7 it can be seen that as the communication resources decrease, the delay performance of each communication technology deteriorates to varying degrees.

[0234] Figure 8 shows the complementary cumulative distribution function (CCDF) of the delay of three offloaded tasks with different V values under the SAC-based strategy. The CCDF reflects the probability that the task delay exceeds a certain threshold. Here, the task delay is defined as the maximum offload delay based on three different communication technologies. The probability that the task offload delay under the SAC-based strategy with different V values in the figure exceeds the delay requirement T = 50 ms is an extremely small probability event. This reflects the effectiveness of the SAC strategy proposed in the present invention and its ability to ensure the delay requirements of offloaded tasks.

[0235] Figure 9 represents the queue backlog corresponding to the offloaded 3D game under different weight coefficient V values. The reason for this consideration is that the 3D game requires the most CPU cycle resources, while for the other two tasks VR and AR, since the CPU cycle resources required for processing are relatively small, the server can generally process them in a timely manner. From Figure 9 it can be seen that the SAC-based strategy can ensure the stability of the queue. The larger V is, the smaller the final stable length of the queue is. Another phenomenon that can be obtained from the figure is that the queue lengths of the random-based strategy and the average-based strategy are both smaller than that of the SAC strategy, because these two strategies do not consider the system utility.

[0236] Figure 10 shows the system utility F(t) of the equal allocation equal offload (EAEO) strategy, random allocation random offload (RARO) strategy and SAC strategy with different V values. It can be seen from the figure that the system utility F(t) of the EAEO strategy and the RARO strategy is much larger than that of the SAC-based strategy, which indicates that Figure 9The EAEO strategy and the RARO strategy in [reference] verify the effectiveness of the SAC-based strategy by sacrificing system utility (i.e., CPU resources) to achieve the stability of the queue length. Figure 9 It also shows the difference in the utility F(t) based on the SAC strategy at different V values. As V increases, the system utility F(t) decreases instead because when V increases, the system will allocate more processing clock rate resources f E to ensure the stability of the system. The additional CPU cycle resources allocated enable the system to rely on millimeter-wave and DSRC communication technologies for offloading, thus making the overall rate of return r t minimal.

Claims

1. A task offloading method based on ultra-reliable and low-latency reinforcement learning of heterogeneous communication technology, characterized in that It includes the following steps: (1) Construct a vehicle edge computing scenario, which consists of a base station connected to a server, multiple roadside units, and vehicles; construct a vehicle heterogeneous communication network composed of three communication technologies: millimeter wave, DSRC, and CV2I. The vehicle offloads tasks to the server for processing through the three communication technologies; (2) Based on the theory of stochastic network calculus, construct a bounded bursty traffic model; (3) Construct a dynamic change model of the base station queue to ensure the stability of the base station queue; (4) Based on the theory of stochastic network calculus, establish communication transmission models for the three communication technologies of millimeter wave, DSRC, and CV2I, and at the same time establish a computing processing model of the CPU; by the tandem theorem, perform min-plus convolution on the communication transmission model and the computing processing model to obtain a system processing model; (5) Deduce the upper bound of the delay probability for offloading and processing based on each communication technology; the delay includes communication transmission time and server computing processing time; (6) Establish the utility of the vehicle edge computing system, which consists of communication utility and computing utility; (7) Establish an optimization problem, with the optimization goal of minimizing the system utility while ensuring task offloading delay and the stability of the base station queue; (8) Use Soft Actor Critic reinforcement learning to learn the offloading strategy for each task and the server CPU allocation strategy; The implementation process of step (7) is as follows: Among them, is the maximum transmission and processing delay requirement for the i-th task; the control variable α(t) = [α1(t), α2(t),....α N (t)] allocates the CPU clock cycle resources, is the communication offloading strategy, where Condition C1 is to keep the queue in a stable state; Condition C2 ensures that the transmission and processing times of each type of task are within the maximum delay requirement. Since the tasks are offloaded through three different communication technologies, and the maximum value among the three is used as the upper bound of the transmission delay for the i-th task; Constraint C3 ensures that the CPU clock cycles used to process all tasks do not exceed the total available CPU computing resources on the server; Constraint C4 ensures that each task selects mmwave, DSRC, or CV2I to execute the computing task; Adopt the Lypunov technique to solve this long-term stochastic constraint C1: Define the second-order Lyapunov function \(L(t)\) and the 1-slot Lyapunov drift \(\Delta L\) t : where \(Q(t)=[q_1(t),q_2(t),...q N (t)]\); then, the desired system utility is added to the drift to obtain the drift-plus-penalty term, i.e., \(\Delta L t + V\cdot E\{F(t)|Q(t)\}\), where \(V\) is a non-negative parameter set by the system, used to trade off between system utility and queue backlog; for any given control parameter \(V\geq0\) with respect to the offloading workload \(\alpha i \), the drift-plus-penalty term is obtained: Among them, The original time-averaged long-term queue length condition C1 is absorbed into the optimization objective in an implicit manner, and the optimization objective of problem P1 is converted to F2(t): Adopt a DRL framework including state, action, and reward to formulate the problem of computing resource allocation strategy and heterogeneous communication offloading strategy in VEC: The state space s at time t t is as follows: s t = [A t , Q t , ξ t (35) Since A t and are both N-dimensional, while is 4N-dimensional, the dimension of the state space is 5N-dimensional; The action space a at time t t is as follows: Among them, α i (t) and both need to satisfy the constraint conditions in formula (30) and Add a dummy variable α N+1 (t) output a deep neural network with N + 1 - dimensional actions, and then use the softmax function for these N + 1 - dimensional variables in the output layer to satisfy: Only take the actions of the top N; similarly, for the output actions of each task i and Use the softmax function to achieve In this way, the action space is N + 1 + 3N dimensional, and the dimension of the action space increases with the increase in the number of task types; similarly, by the output actions of each task and Use the softmax function to thus achieve So the action space is 4N + 1 dimensional; Reward function r at time t t is as follows: r t (a t ,s t ) = -F2(t) (38) r t (a t , s t ) indicates the reward feedback from the environment to the agent after taking action a t and π(a t | s t ) represents the action space distribution taken by the agent based on state s t . The expected long-term discounted return of the system is calculated as: t ​ Among them, $\gamma\in[0,1]$ is a discount factor representing the agent's attention to long-term or short-term rewards. The higher the value, the more the agent focuses on long-term rewards, and vice versa, it focuses on current short-term rewards; $\tau=(s_0,a_0,s_1,a$ a ,\ldots)$ is the state and action trajectory of the agent depending on the action space distribution $\pi(a$ t |s t ).

2. The task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to claim 1, characterized in that The implementation process of step (2) is as follows: Suppose the vehicle has K types of tasks to be processed. At the beginning of each t time slot, A i (t) is the amount of task data that has cumulatively arrived at queue i during the time interval [t, t + 1); meanwhile, given a time interval 0 ≤ s ≤ t, define the binary non-cumulative quantity A i (s, t) = A i (t - s) = A i (t) - A i (s) is the cumulative task volume that has arrived at queue i, and A i (s, t) is a bounded bursty traffic model that satisfies a stationary non-negative stochastic process: A i (s, t) = λ i [ρ i (t - s)+σ i (1) where ρ i is the task arrival rate, σ i is the task burst size, and both are constants. λ i follows a Poisson distribution, and λ i represents the number of vehicles that generate the i-th task within the time interval [s, t).

3. The task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to claim 1, characterized in that The implementation process of step (3) is as follows: The queue length of the base station is expressed as: where q i (t) is the queue length of the i-th task at the start time of time slot t, and A i (t) is the amount of task data accumulated in queue i during the time interval [t, t + 1), and f E is the maximum CPU processing clock rate of the server, and ω i represents the number of CPU clock cycles required for the server to process each bit of task i, and α i (t) represents the proportion of CPU clock cycles allocated by the server to the i-th task, and [x] + = ax(x, 0); the stability of all queues is controlled by the following definition: The left end of equation (3) describes the long-term time-average backlog of the queue; equation (3) means that the strong stability of the queue corresponds to a finite average backlog with a finite average queuing delay.

4. The task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to claim 1, characterized in that, The implementation process of step (4) is as follows: Using β mmw (s,t) represents the total communication transmission volume that mmWave can provide during the time interval [s,t), and C (q) represents the channel capacity of the q-th time slot, and ζ (q) represents the channel gain of the q-th time slot. represents the signal-to-noise ratio, B represents the bandwidth, l and δ represent the transmission distance and the path loss exponent respectively, and the total communication transmission volume that mmWave can provide is: where η = Blog2e, and represents the proportion of task i that communicates and transmits via millimeter waves; the communication transmission volume that millimeter waves can provide for the i-th task is: Among them Use the delay-rate model in the theory of stochastic network calculus to establish the total communication transmission volume that DSRC communication can provide within the time interval [s, t); R dsrc is the DSRC communication bandwidth, represents the average access delay of a collision occurring when transmitting data via DSRC; the communication transmission volume that DSRC can provide for the i-th task is: wherein represents the task ratio of task i for communication transmission via DSRC; Use to represent the communication transmission volume that DSRC can provide for the i-th task in the time interval [s, t): where R cv2i is the communication bandwidth reserved for the i-th task, where Use to represent the computing processing capacity that the server CPU can provide to task i unloaded to the server within the time interval [s, t), and this computing processing capacity is equal to the task amount processed by the CPU in formula (2): Let the set \(H = \{mmwave, dsrc, cv2i\}\) represent the communication technologies that can be offloaded; use to represent the amount of task \(i\) offloaded to the server via communication technology \(g\in H\) during the time interval \([s, t)\), to represent the proportion of tasks of task \(i\) communicated and transmitted via communication technology \(g\), \(q\) i \((s)\) represents the backlog tasks in base station queue \(i\) that have not been processed before time \(s\), use to represent the computing processing capacity provided by the server CPU to during the time interval \([s, t)\), The available computing processing capacity is calculated as: Among them The latency of a task is the sum of the communication transmission time and the server CPU processing time. For task i, the total processing capacity of the system is the sum of the communication transmission volume and the CPU computing processing volume. The total processing capacity that the system can provide for task i unloaded based on communication technology g is the communication transmission of communication technology g and CPU computing processing of the min-plus convolution: where is the min-plus convolution operator, which is one of the most important operators in the theory of stochastic network calculus and has the following operation rules:

5. The task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to claim 1, characterized in that The implementation process of step (5) is as follows: Use to represent the latency of task i offloaded based on communication technology g, and use to represent the upper bound of the probability that task i is transmitted via communication technology g, representing the task transmission and processing time exceeds with a probability less than ε i , which is defined as follows: Obtain the upper bound of the delay The solution is as follows: In the formula is as follows: Among them To obtain the upper bound of the probabilistic delay of task i First, it is necessary to calculate Hypothesis and That is, the communication and computing resources obtained by the offloaded task are much greater than the offloading rate of the task offloaded based on communication technology g Obtain Closed-form solution of Among them, and the upper bound of delay is determined by many factors. The probability ε beyond the delay bound i determines the magnitude of the delay. Each term in the numerator of the second term is related to the task burst size, and the third term is jointly determined by the communication technology g, the remaining resources of the server calculation, and the task volume of the task i offloaded based on the communication technology g to determine the upper bound of delay 6. The task offloading method based on heterogeneous communication technology for ultra-reliable and low-latency reinforcement learning according to claim 1, wherein The implementation process of step (8) is as follows: The SAC algorithm introduces policy entropy while maximizing the optimization objective - F2(t). into the reward, and the expected long-term discounted reward of this model is β t is the weight of the policy entropy, which balances between exploring feasible policies and maximizing the optimization objective; With the continuous change of the reward, the fixed β t will affect the stability of the entire training, so it is very necessary to automatically adjust β during the training process t The optimization problem of reinforcement learning is transformed into: By setting a lower bound is to make β in formula (40) t H(π t (·|s t )) as large as possible; when the agent has not learned the optimal action, increase β t to explore more of the total space; conversely, if the optimal strategy has been learned, then reduce β t to reduce exploration and accelerate the training of the model; obtained based on the Lagrange multiplier method

Citation Information

Patent Citations

  • Device for testing quality of service parameter of vehicle-road wireless communication network

    CN107820227A

  • Software-defined heterogeneous Internet of Vehicles access management and optimization method

    CN111601278A