Heterogeneous vehicle edge computing resource allocation method based on deep reinforcement learning
Through the resource allocation method based on deep reinforcement learning, the limitations of traditional heterogeneous vehicle edge computing systems in resource allocation are solved, the coordinated optimization of communication and computing resources is realized, the requirements of URLLC are met, and the system cost is reduced.
Patent Information
- Application Number
- CN202510280036.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional heterogeneous vehicle edge computing systems have limitations in resource allocation, cannot adapt to dynamically changing network environments and task requirements, and usually only focus on the optimization of communication or computing resources, while ignoring the synergy between the two.
The resource allocation method based on deep reinforcement learning is adopted, and the heterogeneous VEC system model is constructed, the optimization problem model is established, and the resource allocation process is modeled as a DRL process. The resource allocation strategy is optimized using improved SAC algorithms, and the allocation ratio of communication and computing resources is dynamically adjusted.
It significantly reduces system costs, including communication costs and computing costs, while meeting the strict requirements of ultra-reliable low-latency communication (URLLC), improving training convergence speed and network stability.
Smart Images

Figure CN120111583A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technology, and in particular relates to a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning. Background Art
[0002] Vehicle Edge Computing (VEC) is an emerging computing paradigm that aims to provide real-time computing support for vehicles by deploying computing resources on edge servers (such as roadside units RSU or base stations BS) near the vehicle. With the development of intelligent transportation systems and autonomous driving technologies, vehicles need to process a large amount of real-time data, such as high-definition video streaming, augmented reality / virtual reality (AR / VR), 3D games, and sensor data processing required for autonomous driving; these applications place extremely high demands on the real-time and low latency of data processing, while the vehicle's own computing and storage capabilities are limited, so it is necessary to offload tasks to nearby edge servers for processing.
[0003] Although VEC has significant advantages, it also faces some technical challenges in terms of communication resources, computing resources, and environment. Specifically, in terms of communication resources, traditional VEC relies on a single communication technology (such as cellular networks or dedicated short-range communications DSRC), which is difficult to meet the requirements of high data volume and low latency task offloading; in terms of computing resources, the computing power of vehicles and edge servers is limited, and computing resources need to be efficiently allocated to handle a large number of tasks; in terms of the environment, the mobility of vehicles causes the communication link status and task arrival rate to change continuously, and resource allocation strategies need to be adjusted dynamically. In order to solve the communication resource bottleneck of traditional VEC, heterogeneous vehicle edge computing (Heterogeneous VEC) was proposed, which integrates multiple communication technologies such as millimeter wave (mmWave), dedicated short-range communication (DSRC), and cellular network (C-V2I). By integrating these technologies, heterogeneous VEC can significantly improve communication and computing capabilities and better support autonomous driving and intelligent transportation systems. However, heterogeneous VEC systems involve multiple communication technologies and task types, and resource allocation issues become extremely complex.
[0004] In heterogeneous VEC environments, resource allocation is the key to achieving URLLC. However, traditional resource allocation methods have limitations, including: they cannot adapt to dynamically changing network environments and task requirements; they usually only focus on optimizing communication or computing resources, while ignoring the synergy between the two; heterogeneous VEC systems involve multiple communication technologies and task types, making resource allocation issues extremely complex. Summary of the invention
[0005] To solve the above problems, the present invention provides a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning, comprising the following steps:
[0006] S1. Construct a heterogeneous VEC system model including one VEC server, one base station, multiple roadside transceiver units and multiple vehicle units; the heterogeneous VEC system model also includes a data arrival model, a network service model and a computing service model; in the VEC system model, time is divided into equal-length time slots;
[0007] S2. Establish an optimization problem model with the goal of minimizing system utility under resource constraints and URLLC constraints;
[0008] S3. Model the unloading and resource allocation process as a DRL process, build a deep reinforcement learning framework, and obtain the state space, action space, and reward function;
[0009] S4. Based on the deep reinforcement learning framework, the resource allocation strategy is optimized through the improved SAC algorithm to obtain the optimal allocation ratio of communication resources and computing resources.
[0010] Beneficial effects of the present invention:
[0011] The present invention uses the technology of batch extracting experience samples with a higher priority during sample learning through the priority experience playback technology. Compared with uniform random sampling, it improves the ability of intelligent agents to learn and optimize quickly; the present invention significantly reduces system costs, including communication costs and computing costs, by optimizing resource allocation strategies, while meeting the strict requirements of URLLC. The present invention can improve the training convergence speed and network stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0013] Figure 2 It is a schematic diagram of the heterogeneous vehicle edge computing system model of the present invention;
[0014] Figure 3 This is an algorithm framework for allocating edge computing resources for heterogeneous vehicles based on deep reinforcement learning in the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0016] The present invention provides a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning, such as Figure 1 As shown, the following steps are included:
[0017] S1. Construct a heterogeneous VEC system model including 1 VEC server, 1 base station, multiple roadside transceiver units and multiple vehicle units; the heterogeneous VEC system model also includes a data arrival model, a network service model and a computing service model; in the VEC system model, time is divided into equal-length time slots.
[0018] Preferably, the data arrival in the heterogeneous VEC system model constructed by the present invention adopts Poisson distribution, and the heterogeneous VEC system model integrates multiple communication technologies, including dedicated short-range communication (DSRC), millimeter wave (mmWave), cellular-based vehicle-to-infrastructure (C-V2I) technology and local processing. Specifically, the VEC server communicates with the roadside transceiver unit through a wired link, and the base station and the vehicle unit adopt cellular-based vehicle-to-infrastructure communication; the roadside transceiver unit and the vehicle unit adopt dedicated short-range communication and millimeter wave communication, such as Figure 2 shown.
[0019] At the same time, in the embodiment of the present invention, the time is divided into equal length time slots. is a set of time slots, T is the number of time slots, and the duration of each time slot is τ. In the embodiment of the present invention, five task types are defined, namely, speech recognition, face recognition, language translation, 3D game processing, virtual reality (VR) and augmented reality (AR).
[0020] Specifically, in the embodiment of the present invention, the data arrival model is defined as the cumulative amount of tasks arriving at the task queue within a time interval; wherein, the time slot s is defined as 1 To time slot s 2 time, that is, the time interval [s 1 ,s 2 ) The cumulative amount of tasks of the i-th category in i (s 1 ,s 2 )for
[0021]
[0022] a i (t) = n i d i
[0023] Among them, a i (t) represents the cumulative amount of the i-th task in time slot t, n i represents the number of tasks of the i-th category following the Poisson distribution, d i represents the constant size of the i-th task.
[0024] In particular, the cumulative volume arriving at the task queue has a statistical envelope known as exponentially bounded burstiness (EBB), which is defined to provide a form of guarantee with a violation probability of:
[0025]
[0026] Among them, ρ i represents the long-term average arrival rate of the number of tasks of the i-th category, ρ i (s 1 ,s 2 ) represents the number of tasks of the i-th category in the time interval [s 1 ,s 2 ) within the long-term average arrival rate, σ i represents the burstiness of the number of tasks of the i-th category, and P[] represents the probability. The definition of EBB ensures that the number of tasks arriving will not be too sudden, and even if there is a sudden burst, the growth of the number of tasks is controlled and will not exceed an upper bound, and the probability of violating this upper bound will not exceed ε i a .
[0027] For different types of tasks, the frequency required to process each bit of data is different. Therefore, the cumulative amount of tasks Q of the i-th task queue in time slot t+1 is defined as i (t+1) is
[0028]
[0029] Among them, ω i represents the frequency required to process each bit of the i-th task, represents the number of tasks of the i-th type processed by the roadside transceiver unit in time slot t, represents the number of tasks of the i-th type processed by the vehicle unit in time slot t, f E Indicates the total CPU frequency of the VEC server, f v Indicates the total CPU frequency of the vehicle unit. Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, Indicates the CPU frequency ratio of the vehicle unit locally allocated to the i-th task; [x] + =max(x,0), x is the variable parameter.
[0030] It should be noted that the stability of the task queue directly affects the reliability of communication. If the task queue becomes unstable, the tasks arriving at the task queue may be discarded. Therefore, ultra-reliable communication is achieved by maintaining the strong stability of each task queue, which can be expressed as:
[0031]
[0032] While saving resources, not all CPU frequency is used to process tasks immediately, resulting in tasks being stored in a queue. The key is to ensure that the queue length remains stable over time, rather than constantly growing.
[0033] Specifically, in an embodiment of the present invention, based on the SNC theory, network service models of mmWave, DSRC and C-V2I are established respectively, and the network service model describes the maximum transmission capacity of each communication technology within a given time interval.
[0034] The network service of a communication technology refers to its transmission capacity, that is, the maximum number of all arriving tasks that can be provided after eliminating interference; therefore, in the mmWave network service model, the transmission rate of mmWave is calculated according to the Shannon theory, and the transmission rate of mmWave in time slot t is obtained:
[0035] C(t)=Blog 2 (1+ζ(t)γ sinr (t)l -δ )
[0036] Where C(t) represents the transmission rate of millimeter wave communication at time slot t; B represents the bandwidth of the heterogeneous VEC system; ζ(t) represents the amplitude of the millimeter wave communication channel gain coefficient, which follows the Nakagami-m distribution; γ sinr represents the signal-to-noise ratio; l represents the average transmission distance of the vehicle unit, and δ represents the path loss index;
[0037] Based on this, the time interval [s 1 ,s 2 ) within the millimeter wave communication network service β mmw (s 1 ,s 2 )for
[0038]
[0039] In the mmWave network service model, different task types will compete for network services, so the time interval [s 1 ,s 2 ) within the network service of millimeter wave communication for the i-th task for
[0040]
[0041] Where N represents the total number of task types. represents the resource ratio of the j-th task allocated to the millimeter wave communication technology in time slot t, a j (t) represents the cumulative amount of the j-th task in time slot t;
[0042] It is the network service of millimeter wave to other types of tasks, reflecting the interference of other types of tasks transmitted under the same spectrum resources.
[0043] In the network service model of DSCR, according to the IEEE 802.11p standard, access delay is the main delay of DSRC. Therefore, according to the classic delay rate service, it is defined in the time interval [s 1 ,s 2 ) within the network service for dedicated short-range communications β dsrc (s 1 ,s 2 )for
[0044]
[0045] Among them, R dsrc Indicates the maximum transmission rate of dedicated short-range communication, It means that it obeys the Pareto tail distribution with exponent -λ, that is u is a constant.
[0046] Similar to mmWave, it is defined in time interval [s 1 ,s 2 ) within the network service for dedicated short-range communication of the i-th task for
[0047]
[0048] in, represents the resource ratio of the j-th task allocated to the dedicated short-range communication technology in time slot t;
[0049] C-V2I is a communication technology that reserves bandwidth resources for different types of tasks in advance so that different types of tasks do not compete for network services. Therefore, in the network service model of C-V2I, the time interval [s 1 ,s 2 ), network services for cellular-based vehicle-to-infrastructure communications for the i-th task for
[0050]
[0051] in, Represents the maximum transmission rate for cellular-based vehicle-to-infrastructure communications.
[0052] Since local processing is performed directly on the vehicle that generates the task, the task transmission step is omitted. Therefore, for the local processing process, only the calculation of the service curve needs to be considered.
[0053] Specifically, in the embodiment of the present invention, based on the SNC theory, a computing service model is established, which describes the ability of the VEC server to process tasks within a given time interval. In the computing service model, the computing service volume of the i-th task is the computing service volume of the vehicle processor and the VEC server in the time interval [s 1 ,s 2 ] is the number of tasks of the i-th type that are processed in . To this end,
[0054] Defined in the time interval [s 1 ,s 2 ), is the computing service volume β of the i-th task i,comp (s 1 ,s 2 )for
[0055]
[0056] Among them, f E Indicates the total CPU frequency of the VEC server, f v Indicates the total CPU frequency of the vehicle unit. Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, It represents the proportion of CPU frequency allocated to the i-th task locally by the vehicle unit;
[0057] According to the residual service theorem, communication technology The offloaded tasks of type i will compete for computing services with tasks offloaded by other communication technologies, where Represents a set of communication technology types. The computing services of other communication technologies for the i-th type of task can be expressed as:
[0058]
[0059] represents the amount of computing services provided by the communication technology g for the i-th task, Indicates other communication technologies except communication technology g; It represents the resource ratio of the i-th task allocated to the communication technology g in time slot t.
[0060] Therefore, the time interval [s 1 ,s 2 ), the communication technology g offloads the computing service volume of the i-th task for
[0061]
[0062] in, represents the amount of computing services provided by the communication technology g for the i-th task, Indicates other communication technologies except communication technology g.
[0063] S2. An optimization problem model is established with the goal of minimizing system utility under resource constraints and URLLC constraints.
[0064] Specifically, in the embodiment of the present invention, the total system cost model F(t) is constructed as follows:
[0065]
[0066] represents the computational cost, C comp,i represents the communication cost, α 1 , α 2 represents the normalized weight factor used to balance the magnitude of communication cost and computational cost; N represents the total number of task types.
[0067] Calculate costs It reflects the energy consumption cost of VEC server when processing tasks, and its expression is:
[0068]
[0069] c comp represents the cost per unit of energy consumption, represents the energy consumption of the neth CPU core in time slot t, calculated based on the dynamic voltage frequency scaling (DVFS) model:
[0070]
[0071] κ represents the effective capacitance parameter related to hardware, f E Indicates the total CPU frequency of the VEC server, α j (t) represents the CPU frequency ratio assigned to the jth task, N E Indicates the number of CPU cores.
[0072] Communication cost C comp,i reflects the communication cost of offloading the task to the VEC server, and its expression is:
[0073]
[0074] c comm represents the unit communication cost of offloading tasks through C-V2I technology, represents the resource ratio of the i-th task allocated to C-V2I communication technology in time slot t, a i (t) represents the cumulative amount of the i-th task in time slot t.
[0075] Minimize the system utility under resource constraints and URLLC constraints. The optimization problem model is expressed as
[0076]
[0077]
[0078] in, represents the time slot set, t represents the time slot, C(t) represents the total cost of time slot t, represents the set of task types, α(t) represents the computing resource allocation ratio, including and Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, It represents the CPU frequency ratio of the vehicle unit locally allocated to the i-th task, represents the communication resource allocation ratio, which includes the resource ratio of the i-th task allocated to different communication technologies in time slot t, where represents the resource ratio of the i-th task allocated to the communication technology g in time slot t, T i max represents the maximum allowed delay, N represents the total number of task types; sup represents the supremum, which is the minimum upper bound value under all possible circumstances; Expressing expectation, Q i (m) represents the cumulative amount of tasks in the i-th task queue in time slot m, represents the upper bound of the delay of mmWave transmitting the i-th task in time slot t, represents the upper bound of the delay of DSRC transmitting the i-th task in time slot t, represents the upper bound of the delay of C-V2I transmitting the i-th task in time slot t, represents the upper bound of the delay of the vehicle locally processing the i-th task in time slot t, Represents a collection of communication technologies, which includes mmWave communication, DSRC communication, C-V2I communication, and vehicle local processing.
[0079] The goal of the optimization problem model is to minimize the use of communication and computing resources while meeting the requirements of ultra-reliable low-latency communication (URLLC). The ultra-reliability constraint requires that the average backlog of each task queue remains limited over a long period of time to ensure the stability of the system. The present invention uses the Lyapunov optimization method to transform the long-term ultra-reliability constraint into a short-term constraint, thereby simplifying the complexity of the problem.
[0080] The quadratic Lyapunov function L(Q(t)) is defined as:
[0081]
[0082] It means "defined as" and is used to define the quadratic Lyapunov function. Δ(Q(t)) represents the backlog increment of all queues from time slot t to t+1, that is, the Lyapunov drift function, which is specifically expressed as:
[0083]
[0084] in, represents the expectation. The upper bound of Δ(Q(t)) is
[0085]
[0086] in, represents expectation, Q(t)=[Q 1 (t),Q 2 (t),...Q N (t)].
[0087] set up for a i The upper bound of (t) is:
[0088]
[0089] Applying the chance expectation minimization technique, we get:
[0090]
[0091] In order to ensure the ultra-reliability of the system, the upper bound of the Lyapunov drift needs to be as small as possible. According to the Lyapunov optimization theory, if the upper bound of the Lyapunov drift can be controlled within a small range, the queue backlog of the system will remain limited, thus satisfying the ultra-reliability constraint.
[0092] Therefore, the short-term optimization goal can be expressed as:
[0093]
[0094] where α v (t) represents the proportion of vehicle local CPU frequency allocated to each task type at time slot t, α E (t) represents the CPU frequency ratio of edge servers allocated to each task type at time slot t.
[0095] In order to minimize the system cost while satisfying the ultra-reliability constraint, a weight factor V is introduced (to balance the ultra-reliability constraint and the system cost). The final optimization problem can be expressed as:
[0096]
[0097]
[0098] Among them, α(t) represents the computing resource allocation ratio, including and in Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, It represents the proportion of CPU frequency allocated to the i-th task locally by the vehicle unit; Indicates the communication resource allocation ratio, including It represents the resource ratio of the i-th task allocated to the communication technology g in time slot t.
[0099] Specifically, the calculation of the delay upper bound is a key step to ensure that the requirements of ultra-reliable low-latency communication (URLLC) are met. To this end, the present invention calculates the delay upper bound of the task offload by the MGF method and compares it with the maximum delay requirement of the task to ensure URLLC performance.
[0100] Since the arrival of tasks and network services are stable random processes, the probability delay upper bound is used to define the delay upper bound of the i-th task using communication technology g, using Indicates that:
[0101]
[0102] represents the system service of the i-th task through the communication technology g, which consists of network services and computing services; ε i represents the violation probability of the i-th task; W i g (t) represents the delay of the i-th task under communication technology g, which is calculated as
[0103]
[0104] in, It is min-pius deconvolution, inf represents the infimum, which is used to represent the minimum value of the set.
[0105] It is deduced from this that
[0106]
[0107] Among them, A i (s,t) represents the cumulative amount of the i-th task in the time interval [s,t), Indicates that the i-th task is in the time interval The system service through the communication technology g consists of network service and computing service; u' represents a time interval for calculating the cumulative probability; θ represents the parameter used in the moment generating function, yes The moment generating function (MGF) of , the arrival process and the service process are independent.
[0108] According to the series theorem of random network calculus (SNC), we get:
[0109]
[0110] is a min-pius convolution, based on the moment generating function (MGF) of the affine envelope model, and we get:
[0111]
[0112] is the moment generating function (MGF) of the time interval [s,z), is the time interval The moment generating function (MGF) of .
[0113] definition As the arrival process A i The moment generating function of (s, t) can be obtained as follows: where ρ i and σ i Represent the long-term average arrival rate of the i-th task and the burstiness of the arrival process, respectively, and we get:
[0114]
[0115] Since the interval time [s 1 ,s 2 ) is relatively small, so the time interval [s 1 ,s 2 ) is considered as a constant in the time interval, we get:
[0116]
[0117] According to the residual service theorem, we get:
[0118]
[0119] make:
[0120]
[0121] get:
[0122]
[0123] in,
[0124] Set the upper limit to ε i , except for local processing, the remaining communication technology g The closed form solution is:
[0125]
[0126] Among them, ln() represents logarithmic operation.
[0127] For the local processing process, only the service curve needs to be calculated. The failure probability inequality of local processing can be written as
[0128]
[0129] According to the residual service theory, the network service for local processing of the i-th task is obtained for:
[0130]
[0131] Get the upper bound of the latency for local processing:
[0132]
[0133] S3. The unloading and resource allocation process is modeled as a DRL process, a deep reinforcement learning framework is constructed, and the state space, action space and reward function are obtained.
[0134] Specifically, for complex and dynamic heterogeneous VEC environments, the optimization objective is non-convex and has a high dimensionality, so it is difficult to solve the problem based on traditional optimization methods such as convex optimization methods. In addition, traditional optimization methods are difficult to adapt to dynamic environments and cannot provide real-time solutions. In recent years, using DRL to solve non-convex optimization problems has become a trend in the research field. This is because it is robust and adaptable in solving complex non-convex problems in dynamic and high-dimensional environments. Therefore, in our research, DRL is an ideal choice to solve the problem.
[0135] The process of determining the state space, action space, and reward function includes:
[0136] Considering that the arrival of tasks is random and affects system performance, the number of arriving tasks in each time slot t is set to As the first element of the state; queue backlog is another key factor affecting the ultra-reliability of the system, especially at high vehicle density, so consider
[0137] As the second element of the state; since the delay is bounded by and Since mmWave has sufficient radio resources, the competition of mmWave network services in different types of tasks is not considered, and there is no competition of C-V2I network services between different types of tasks, so As the third element of the state, ξ t =[ξ 1 (t),ξ 2 (t),...ξ N (t)] represents the communication service capability and computing service capability of different communication technologies in time slot t, ξ i (t) represents the communication service capacity and computing service capacity of the i-th task in time slot t, represents the communication service capability of the i-th task via DSRC in time slot t, represents the computing service capacity of the i-th task through C-V2I in time slot t, represents the computing service capability of the i-th task through mmWave in time slot t, represents the computing service capacity of the i-th task through DSRC in time slot t, represents the computing service capacity of the i-th task in the vehicle. Therefore, in the state space s at time slot t t as follows:
[0138]
[0139] Define action a in time slot t t for
[0140] Among them, α v (t) represents the proportion of vehicle local CPU frequency allocated to each task type at time slot t; α E (t) represents the CPU frequency ratio of edge servers allocated to each task type at time slot t; Indicates the resource ratio of each task type allocated to different communication technologies (such as mmWave, DSRC, C-V2I, local) at time slot t;
[0141] The goal of DRL is to maximize the long-term discounted reward, while our goal is to minimize the target, so the reward of the DRL framework is expressed as the negative of the target, and the reward function r at time slot t t as follows:
[0142]
[0143] Among them, for the penalty of violating the delay constraint, we put the penalty term into the reward, P e represents the penalty weight, represents the upper limit of the delay of mm-Wave transmitting the i-th task in time slot t, represents the upper limit of the delay of DSRC transmitting the i-th task in time slot t, It represents the upper limit of the delay of C-V2I transmitting the i-th task in time slot t.
[0144] S4. Based on the deep reinforcement learning framework, the resource allocation strategy is optimized through the improved SAC algorithm, and the optimal allocation ratio of communication resources and computing resources is obtained by training the deep neural network.
[0145] The SAC algorithm is known for its stability and robustness, and SAC encourages exploration and maintenance of a more diverse range of actions, so SAC is suitable for optimizing policies in complex VECs with multiple types of tasks and communication technologies. Figure 3 As shown in the figure, the SAC algorithm mainly consists of five main network structures: a policy network (i.e., Actor network, with parameter φ) for approximating the optimal policy, which is used to output the action distribution under a given state; two value networks (i.e., Critic networks, with parameters φ respectively) for estimating the state-action value function. 1 and φ 2 ), used to estimate the value of the state-action pair; two target value networks (i.e., target Critic networks, with parameters and ), which is used to construct the target value to stabilize the training process.
[0146] According to SAC, the t and t The expected long-term discounted reward J under is as follows:
[0147]
[0148] Among them, π(·|s t ) is the strategy when taking all available actions under st, γ t-1 ∈[0, 1] is the discount factor, is the strategy entropy, β is the trade-off weight between exploring feasible strategies and maximizing returns, which can be called the entropy coefficient. It can be adjusted dynamically. The formula is as follows
[0149]
[0150] in, represents the dimension of the action space; π * (a t ∣s t ) is in a t and t The maximum expected return of the optimal strategy.
[0151] The specific process of the SAC algorithm is as follows:
[0152] Initialization phase:
[0153] Initialize the environment, including the state space and action space; initialize the policy network (with parameter φ), the two value networks (with parameters φ 1 and φ 2 ), two target value networks (with parameters and ), entropy coefficient β, experience replay buffer, learning rate, discount factor γ, etc.
[0154] Training cycle phase:
[0155] The specific process of each iteration includes:
[0156] S41. Strategy evaluation, which includes three parts: action selection, environment interaction, and experience storage. The main content is to sample an action for a given state based on the current strategy network, apply the action to the environment, obtain the next state, reward, and whether it ends; store the current experience (state, action, reward, next state) in the experience playback buffer.
[0157] In the embodiment of the present invention, in the first time slot, i.e., t=0, for each task type, the arrival rate of the task is obtained according to the Poisson distribution and the number of generated tasks, and the state s is initialized. 0 ; Set the state s 0 Input Actor network, output strategy π φ (a 0 ∣s 0 ), based on the strategy π φ (a 0 ∣s 0 ) Generate action a' 0 , which includes α′(0) and The action dimension is N+1, where N is the total number of task types. Applying the softmax function ensures that the following constraints are met: and Set α(0) to [a 1 (0),a 2 (0),...a N (0)], according to α(0) and Get action a 0 ; VEC server and vehicle take action locally a 0 , based on this action a 0 , calculate the upper bound of the delay and reward r t , and then update the state s of the next time slot 1 . Then the tuple (s 0 ,a 0 ,r 0 ,s 1 ) is stored in the buffer and the algorithm moves to the next time slot. The above process is iterated until the maximum number of time slots is reached.
[0158] S42. Randomly sample a batch of experience data from the experience playback buffer for training the network. This can improve data utilization, reduce the correlation between data, and help the convergence of the model.
[0159] In the embodiment of the present invention, M tuples are selected from the experience playback buffer to form a mini-batch of data, each tuple contains a state, action, reward and next state, a total of 4 elements. For tuple m = 1, 2, ..., M, its state s m Input the Actor network and obtain a new action a based on the output of the Actor network m,new and the corresponding strategy π(a m,new |s m ), and then calculate the loss function gradient of β
[0160] in, is the dimension of the action space, π * (a m,new ∣s m ) is in state s m The optimal strategy under .
[0161] Specifically, the improved SAC algorithm proposed in the present invention improves the basic SAC algorithm through the Priority Experience Replay (PER) technology. PER is a technology for batch extracting experience samples with a higher priority during sample learning. Compared with uniform random sampling, it improves the ability of the intelligent agent to learn and optimize quickly. The core idea is to use TD error to calculate the priority of the sample. The larger the TD error, the more worthy the experience of the batch is for the intelligent agent to learn. The main improvements of the present invention include
[0162] The absolute value of the TD error of the empirical sample in the SAC algorithm is defined as the average of the absolute values of the TD errors of the two value networks, expressed as
[0163]
[0164] Among them, |δ j | represents the absolute value of TD error. trepresents the immediate reward at time slot t, γ represents the discount factor, Represents the target value network in state s t+1 Take action a t+1 The value function estimate, Q φ q(s t ,a t ) indicates that the value network is in state s t Take action a t The value function estimate.
[0165] The priority calculation formula of the experience sample is:
[0166] p j =|δ j |+ε
[0167] Among them, p j represents the priority of the jth experience sample, and ω represents a small positive constant to avoid edge cases when TD-error is 0.
[0168] The priority determines the probability of experience being selected for replay during the learning process, so the probability of experience samples being sampled is
[0169]
[0170] Among them, ζ represents the adjustment factor of the priority. The larger the value, the stronger the priority playback. When ζ=0, the sampling method is no longer priority experience sampling, but random uniform sampling; when 0<ζ<1, it is partial priority sampling; when ζ=1, it is called full priority sampling. Indicates the priority under ζ.
[0171] S43. Calculation of loss function and update of various network parameters.
[0172] In the embodiment of the present invention, s m and a m,new Input into two value networks and output action value functions respectively and Q φ2 (s m ,a m,new ); Based on this, the gradient of the loss function of the value network is calculated for
[0173]
[0174] where ∈ is the noise sampled from a multivariate normal distribution, f(∈;s m ) is used to reparameterize action a m,new The function of Q(s m ,a m,new)yes and Q φ2 (s m ,a m,new ) is the minimum value between .
[0175] Calculate the target value based on the output of the current policy network and the two target value networks
[0176] π′ φ (a′ m ∣s′ m ) is in the next state s' m Next, according to the new action a' output by the policy network m probability; is the evaluation of the new state-action pair by the two target value networks. Next, the value function φ is calculated 1 and φ 2 The gradient of , the gradient is calculated as follows:
[0177]
[0178] in, Represents φ b The gradient of; M is the number of samples in the mini-batch; r m is in state s m Next, perform action a m Instant rewards received.
[0179] Use the Adam optimizer to update the β, φ, φ1, and φ2 parameters according to the calculated gradients. For the entropy coefficient β, use the gradient For the policy network parameter φ, use the gradient For the value network parameters φ1 and φ2, the gradients are used Every R t In the iteration, the parameters of the two target value networks are updated as follows:
[0180]
[0181] Among them, τ b is a small constant used to smoothly update the parameters of the target network. Repeat the above process until the maximum number of iterations or the maximum number of training rounds is reached, and the optimal parameter φ is finally obtained. * , indicating that the training is completed.
[0182] In summary, the present invention proposes a deep reinforcement learning solution for optimizing resource allocation in heterogeneous vehicle edge computing (VEC) systems to achieve ultra-reliable low-latency communication (URLLC) and minimize system costs. The method dynamically adjusts computing and communication resource allocation strategies by constructing a VEC system model that includes a variety of communication technologies and task types, and adopts an improved SAC algorithm. During the training phase, the algorithm interacts with the VEC environment to collect data, batch extracts empirical samples with higher priority during sample learning, and updates network parameters through policy gradient ascent and value function minimization to achieve policy optimization. This method effectively solves the limitations of traditional resource allocation strategies in dynamic adaptability and optimization, and provides an intelligent and adaptable resource management solution for VEC systems.
[0183] In the present invention, unless otherwise clearly stipulated and limited, the terms such as "installation", "setting", "connection", "fixation" and "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral one; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meanings of the above terms in the present invention according to the specific circumstances.
[0184] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Construct a heterogeneous VEC system model including one VEC server, one base station, multiple roadside transceiver units and multiple vehicle units; the heterogeneous VEC system model also includes a data arrival model, a network service model and a computing service model; in the VEC system model, time is divided into equal-length time slots; S2. Establish an optimization problem model with the goal of minimizing system utility under resource constraints and URLLC constraints; S3. Model the unloading and resource allocation process as a DRL process, build a deep reinforcement learning framework, and obtain the state space, action space, and reward function; S4. Based on the deep reinforcement learning framework, the resource allocation strategy is optimized through the improved SAC algorithm to obtain the optimal allocation ratio of communication resources and computing resources.
2. According to a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning in claim 1, it is characterized in that: In the heterogeneous VEC system model, the VEC server communicates with the roadside transceiver unit through a wired link, and the base station and the vehicle unit use cellular-based vehicle-to-infrastructure communication; dedicated short-range communication and millimeter wave communication are used between the roadside transceiver unit and the vehicle unit.
3. According to a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning in claim 1, it is characterized in that: In the data arrival model, Defined as the cumulative amount A of the i-th task in the time interval [s1, s2) between time slot s1 and time slot s2 i (s1,s2) is a i (t)=n i d i Among them, a i (t) represents the cumulative amount of the i-th task in time slot t, n i represents the number of tasks of the i-th category following the Poisson distribution, d i represents the constant size of the i-th task; Define the cumulative amount of tasks Q of the i-th task queue in time slot t+1 i (t+1) is Among them, ω i represents the frequency required to process each bit of the i-th task, represents the number of tasks of the i-th type processed by the roadside transceiver unit in time slot t, represents the number of tasks of the i-th type processed by the vehicle unit in time slot t, f E Indicates the total CPU frequency of the VEC server, f v Indicates the total CPU frequency of the vehicle unit. Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, Indicates the CPU frequency ratio of the vehicle unit locally allocated to the i-th task; [x] + =max(x,0), x is the variable parameter.
4. According to a method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 1, it is characterized in that: In the network service model, Define the network service β of mmWave communication within the time interval [s1, s2) mmw (s1,s2) is C(t)=Blog2(1+ζ(t)γ sinr (t)l -δ ) Where C(t) represents the transmission rate of millimeter wave communication at time slot t; B represents the bandwidth of the heterogeneous VEC system; ζ(t) represents the amplitude of the millimeter wave communication channel gain coefficient, which follows the Nakagami-m distribution; γ sinr represents the signal-to-noise ratio; l represents the average transmission distance of the vehicle unit, and δ represents the path loss index; Define the network service of mmWave communication for the i-th task in the time interval [s1, s2) for Where N represents the total number of task types. represents the resource ratio of the j-th task allocated to the millimeter wave communication technology in time slot t, a j (t) represents the cumulative amount of the j-th task in time slot t; Define the network service β of dedicated short-range communication within the time interval [s1, s2) dsrc (s1,s2) is Among them, R dsrc Indicates the maximum transmission rate of dedicated short-range communication, It means that it follows the Pareto tail distribution with exponent -λ; Define a network service for dedicated short-range communication of the i-th task in the time interval [s1, s2) for in, represents the resource ratio of the j-th task allocated to the dedicated short-range communication technology in time slot t; Define the network service for cellular-based vehicle-to-infrastructure communication for the i-th task in the time interval [s1, s2) for in, Represents the maximum transmission rate for cellular-based vehicle-to-infrastructure communications.
5. The method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 1 is characterized in that: In the computing service model, Defined as the computing service volume β of the i-th task in the time interval [s1, s2) i,comp (s1,s2) is Among them, f E Indicates the total CPU frequency of the VEC server, f v Indicates the total CPU frequency of the vehicle unit. Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, It represents the proportion of CPU frequency allocated to the i-th task locally by the vehicle unit; Defined as the amount of computing service that a communication technology can offload from the i-th task within the time interval [s1, s2) for in, represents the amount of computing services provided by the communication technology g for the i-th task, Indicates other communication technologies except communication technology g.
6. The method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 1 is characterized in that: The optimization problem model is expressed as in, represents the time slot set, t represents the time slot, F(t) represents the total cost of time slot t, represents the set of task types, α(t) represents the computing resource allocation ratio, including and Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, It represents the CPU frequency ratio of the vehicle unit locally allocated to the i-th task, Indicates the communication resource allocation ratio, represents the resource ratio of the i-th task allocated to the communication technology g in time slot t, T i max represents the maximum allowed delay, N represents the total number of task types, sup represents the supremum, Expressing expectation, Q i (m) represents the cumulative amount of tasks in the i-th task queue in time slot m, represents the upper bound of the delay of mm-Wave transmitting the i-th task in time slot t, represents the upper limit of the delay of DSRC transmitting the i-th task in time slot t, represents the upper limit of the delay of C-V2I transmitting the i-th task in time slot t, represents the upper limit of the delay of the vehicle locally processing the i-th task in time slot t, represents a collection of communication technologies; Using the Lyapunov optimization method, the weight factor V is introduced to transform the optimization problem into the following expression:
7. A method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 6, characterized in that: The total cost F(t) of time slot t is calculated as in, represents the computational cost, C comp,i represents the communication cost, α1 and α2 represent the normalized weight factors, and N E represents the number of CPU cores, N represents the total number of task types; c comp represents the cost per unit of energy consumption, represents the energy consumption of the neth CPU core in time slot t, κ represents the effective capacitance parameter related to the hardware, and f E Indicates the total CPU frequency of the VEC server, α j (t) represents the CPU frequency ratio assigned to the jth task, c comm represents the unit communication cost of offloading tasks through C-V2I technology, represents the resource ratio of the i-th task allocated to C-V2I communication technology in time slot t, a i (t) represents the cumulative amount of the i-th task in time slot t.
8. The method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 1 is characterized in that: Define the state space s at time slot t t for: in, represents the set of the number of tasks arriving at time slot t, a i (t) represents the cumulative amount of the i-th type of task in the t-th time slot, represents the backlog set of time slot t; ξ t =[ξ1(t),ξ2(t),...ξ N (t)] represents the communication service capability and computing service capability of different communication technologies in time slot t, represents the communication service capability and computing service capability of the i-th task in time slot t, represents the communication service capability of the i-th task via DSRC in time slot t, represents the computing service capacity of the i-th task through C-V2I in time slot t, represents the computing service capability of the i-th task through mmWave in time slot t, represents the computing service capacity of the i-th task through DSRC in time slot t, represents the computing service capability of the i-th task in the local vehicle; N represents the total number of task types; Define action a in time slot t t for Among them, α v (t) represents the proportion of vehicle local CPU frequency allocated to each task type at time slot t; α E (t) represents the CPU frequency ratio of edge servers allocated to each task type at time slot t; represents the resource ratio of each task type allocated to different communication technologies at time slot t; Define the reward function r for time slot t t for Among them, V represents the weight factor, Q i (t) represents the backlog of the i-th task queue at time slot t, f E Indicates the total CPU frequency of the VEC server, f v Indicates the total CPU frequency of the vehicle unit. Indicates the CPU frequency ratio of the VEC server allocated to the i-th task, represents the CPU frequency ratio of the vehicle unit locally allocated to the i-th task, ω i represents the frequency required to process each bit of the i-th task, F(t) represents the total cost of time slot t, and P e represents the penalty weight, T i max represents the maximum allowed delay, represents the upper limit of the delay of mm-Wave transmitting the i-th task in time slot t, represents the upper limit of the delay of DSRC transmitting the i-th task in time slot t, It represents the upper limit of the delay of C-V2I transmitting the i-th task in time slot t.
9. The method for allocating heterogeneous vehicle edge computing resources based on deep reinforcement learning according to claim 1, characterized in that: The improved SAC algorithm is based on the basic SAC algorithm. It introduces the priority experience playback technology and uses the TD error to calculate the priority of the experience sample. The absolute value of the TD error of the experience sample is defined as the average of the absolute values of the TD errors of the two value networks, expressed as Among them, |δ j | represents the absolute value of TD error, r t represents the immediate reward at time slot t, γ represents the discount factor, Represents the target value network in state s t+1 Take action a t+1 The value function estimate, Q φq (s t ,a t ) indicates that the value network is in state s t Take action a t The value function estimate.
Citation Information
Cited By
Dynamic data privacy protection method based on Lyapunov-SNC cooperative computing model
CN121727862A
A Dynamic Data Privacy Protection Method Based on the Lyapunov-SNC Collaborative Computing Model
CN121727862B