V2v communication multi-vehicle task offloading method based on deep reinforcement learning

The V2V communication multi-vehicle task offloading method based on deep reinforcement learning optimizes the trade-off between transmission latency and power consumption, solves the problem of resource waste in traditional algorithms, and achieves more efficient task offloading.

CN116017572BActive Publication Date: 2026-01-27JIANGSU EXPRESSWAY COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211680823.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2026-01-27
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

Traditional V2V communication task offloading algorithms fail to effectively balance transmission latency and power consumption, and fail to make full use of computing resources, resulting in resource waste.

Method used

A V2V communication multi-vehicle task offloading method based on deep reinforcement learning is adopted. By calculating the cooperation between the task vehicle and surrounding computing vehicles and relay vehicles, the transmission latency and interruption probability are optimized. The deep deterministic policy gradient algorithm (DDPG) is combined for task allocation and power optimization.

Benefits of technology

While ensuring a high success rate for task unloading, a better balance between transmission latency and power consumption is achieved, improving resource utilization and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116017572B_ABST
    Figure CN116017572B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of V2V communication multi-vehicle task unloading method based on deep reinforcement learning, comprising the following steps: the computing task of task vehicle HV is unloaded to the multiple vehicles of surrounding with computing resource far vehicle RV, while using relay vehicle REV to transmit task by multi-hop;Respectively, the transmission delay of single-hop V2V task unloading and the transmission delay of multi-hop V2V task unloading using relay vehicle are obtained;Transmission interruption probability of multi-hop V2V task unloading using relay vehicle is calculated;While guaranteeing the transmission success rate of task unloading, the power and transmission delay of task vehicle unloading process are optimized;According to optimization goal, use deep reinforcement learning DDPG algorithm to solve optimization problem.Compared with other task unloading schemes that trade off communication delay and power consumption, the present application can guarantee the transmission success rate of task unloading, and also has better performance in the trade-off between transmission delay and power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical fields:

[0001] This invention relates to the field of multi-vehicle task offloading in V2V communication, and particularly to a V2V communication multi-vehicle task offloading method based on deep reinforcement learning. Background technology:

[0002] With the rapid development of vehicle-to-everything (V2X) technology, traditional MEC (Multi-access Edge Computing) technology can no longer meet the needs of task offloading, especially in V2X applications such as autonomous driving, which require low latency and high reliability for task offloading. The focus of task offloading has gradually shifted from MEC to V2V (Vehicle-to-Vehicle).

[0003] Vehicle-to-vehicle (V2V) communication is the direct exchange of information between vehicles within effective communication range. It significantly reduces the burden on base stations, meeting the low latency and high reliability requirements of vehicle-to-everything (V2X) applications. Simultaneously, V2V greatly improves the utilization rate of surrounding idle computing resources. To meet the low latency and high reliability requirements of V2V communication systems, it is necessary to research and propose more reasonable and effective task offloading and enhancement schemes.

[0004] Traditional V2V communication task offloading algorithms fail to consider the dynamic nature of vehicular networks. The rapid movement of vehicles leads to continuous changes in network topology, and the unpredictability of traffic is a major challenge for V2V communication technology. Without considering the dynamic changes in real-world vehicular network scenarios, long-term system performance cannot be optimized. V2V communication places more stringent demands on latency and reliability, but most V2V communication task offloading algorithms cannot strike a good balance between communication latency and power consumption. A recent study considered combining deep reinforcement learning algorithms with the dynamic environment of vehicular networks, but further optimization of latency and reliability is needed. In particular, many computational resources remain unutilized, resulting in significant resource waste. Summary of the Invention:

[0005] The technical problem to be solved by the present invention is a task offloading method for V2V communication in the presence of relay vehicles. This method can ensure the success rate of task offloading while also achieving better performance in terms of the trade-off between transmission latency and power consumption.

[0006] This invention is achieved through the following technical solution:

[0007] A V2V communication multi-vehicle task offloading method based on deep reinforcement learning, the method includes the following steps:

[0008] (1) V2V multi-vehicle task offloading can offload the computing tasks of the task vehicle HV to multiple remote vehicles RV with computing resources in the vicinity, and can also use relay vehicles REV to perform multi-hop to transmit tasks.

[0009] (2) Calculate the transmission delay of single-hop V2V task offloading and the transmission delay of multi-hop V2V task offloading using relay vehicles respectively;

[0010] (3) Calculate the transmission interruption probability of multi-hop V2V task offloading using relay vehicles;

[0011] (4) Based on the transmission delay calculated in step (2) and the transmission interruption probability obtained in step (3), while ensuring the success rate of task unloading transmission, optimize the power and transmission delay of the task vehicle unloading process.

[0012] (5) Based on the optimization objective of step (4), the deep reinforcement learning DDPG algorithm is used to solve the optimization problem.

[0013] Step (1) refers to the fact that at each moment, the mission vehicle has a size of x. t The task is broken down into multiple subtasks through algorithm calculation, and can be offloaded to one or more computing vehicles through the PC5 interface, or it can be offloaded to computing vehicles outside the communication range through relay vehicles.

[0014] In step (2), the task vehicle can calculate the transmission delay of single-hop V2V task offloading and the transmission delay of multi-hop V2V task offloading using relay vehicles based on the task volume, computing resources of the computing vehicle, and channel status.

[0015] The transmission latency for single-hop V2V task offloading without selecting relay vehicles is expressed as follows:

[0016]

[0017] in,

[0018] x t This indicates the size of the task data generated at time t;

[0019] y t This indicates the amount of data returned after the task is uninstalled;

[0020] Indicates the uplink transmission rate of HV and RVn;

[0021] Indicates the downlink transmission rate of HV and RVn;

[0022] ω t This represents the calculated strength of RV;

[0023] f t,n This represents the CPU frequency allocated for processing HV tasks during time period t, where f t,n ∈[0,F n ];

[0024] F n This is the maximum CPU frequency of RVn;

[0025] The specific uplink and downlink transmission rates for HV and RVn are expressed as follows:

[0026]

[0027]

[0028] in, It is the transmission power allocated at each moment, W is the channel bandwidth, and σ is the transmission power allocated at each moment. 2 This represents noise power, where P is the fixed transmit power given by the RVn feedback result, and the wireless channel state. This represents the state between HV and RVn at time t. and It's an interference when the task is unloaded to RVn.

[0029] The mission offloading transmission delay from the mission vehicle to RVn via relay multi-hop transmission is expressed as follows:

[0030]

[0031] in,

[0032] d up (t,z) is the uplink wireless channel transmission delay at the z-th hop, and is the ratio of the transmission workload to the uplink transmission rate at the z-th hop. It is expressed as follows:

[0033]

[0034] d down (t,e) is the downlink radio channel transmission delay at the e-th hop, and is the ratio of the transmission workload to the uplink transmission rate at the e-th hop. It is expressed as follows:

[0035]

[0036] d com (t,n) represents the computation latency on the computing vehicle RV when the task is transmitted to the computing vehicle RV, as shown below:

[0037]

[0038] In step (3), there is a probability of interruption when transmitting tasks via multiple hops through relay vehicles. The vehicle obtains the distance between the multi-hop links to calculate the end-to-end interruption probability, as follows:

[0039] (4.1) For multi-hop V2V task offloading via relay vehicle REV, the vehicle-to-vehicle interruption probability is defined as the probability that the vehicle-to-vehicle signal-to-noise ratio is lower than the set limit value, i.e., the signal-to-noise ratio threshold. The vehicle-to-vehicle interruption probability is expressed as:

[0040]

[0041] (4.2) The amplitude of the signal received in the V2V channel follows a Weibull distribution, and the received signal-to-noise ratio f under Weibull fading is... SNR (a) is represented as:

[0042]

[0043] Where C represents the Weibull fading parameter, Γ represents the gamma function, and SNR... eq The average signal-to-noise ratio is expressed as follows:

[0044]

[0045] (4.3) Based on (4.1) and (4.2), the vehicle-to-vehicle interruption probability can be expressed as:

[0046]

[0047] (4.3) Probability of end-to-end interruption P outsum Represented as:

[0048]

[0049] in,

[0050] z indicates that there are z links in the communication process;

[0051] L i This represents the distance between the two ends of the i-th hop link;

[0052] P out (L i ) is the interruption probability of the i-th hop in the communication process.

[0053] In step (4), at each time step, in order to consider the trade-off between latency and power consumption, an optimization problem is proposed to minimize the overall consumption of the unloading process at the current time step:

[0054]

[0055] Simultaneously satisfy the constraints:

[0056] d max (t,n)≤d th

[0057]

[0058] ζ1,ζ1∈[0 ,1 ],ζ1+ζ1=1

[0059] in,

[0060] d th Threshold constraints representing latency;

[0061] P max Threshold constraints representing power consumption;

[0062] d max (t,n) represents the maximum transmission delay when multiple vehicles are selected;

[0063] ζ1 and ζ2 represent the weights of latency and power consumption, respectively.

[0064] In the continuous state of power allocation, the DDPG-based algorithm jointly allocates the task data volume and power in the vehicular network. The DDPG task offloading algorithm in step (5) is as follows:

[0065] (5.1) In the DDPG algorithm, the agent can generate an information set (s) by interacting with the environment. (t) ,a (t) ,r(s (t) ,a (t) ),s (t+1) The information is then stored in an experience replay buffer, and a small subset of samples is randomly selected from the buffer to train the neural network. This reduces the correlation between samples and improves learning efficiency. Therefore, the state space, action space, and reward function in the network model are defined.

[0066] During each learning process, the agent can obtain state input from the current environment. The state space in the DDPG network model is defined as:

[0067]

[0068] e (t) ={d (t) ,υ (t) ,z (t) ,h (t) ,l (t)}

[0069] in,

[0070] d (t) This indicates the amount of task data that needs to be unloaded from the front of the buffer.

[0071] υ (t) This refers to the set of vehicles with computing resources surrounding the main vehicle, which is the mission vehicle. z (t) This indicates the amount of computing resources available for surrounding vehicles to unload tasks.

[0072] h (t) Indicates the wireless channel status.

[0073] l (t) Indicates the distance between vehicles.

[0074] Use the previously learned state information from the previous φ times as the current state information.

[0075] The actions of an intelligent agent are determined by the amount of task data and the transmission power allocated by the onboard terminal device. The action space is defined as follows:

[0076]

[0077] in This represents the range of values ​​for the available transmission power. This represents the serial number of the selected vehicle with computing resources.

[0078] At the same time, the size of the task data volume needs to meet the following requirements:

[0079]

[0080] When the environment performs an action, the agent receives a corresponding reward or penalty. In this study of vehicle unloading, ζ2 in the objective function is denoted as α, and ζ1 is denoted as 1-α. The reward function is defined as:

[0081]

[0082] Where α represents the weight of the transmission power, This indicates that if the task unloading latency exceeds a certain threshold, a negative reward is awarded. This indicates that if the task unloading delay is within the delay limit, a positive reward will be given.

[0083] Meanwhile, the total reward must meet the following requirements:

[0084]

[0085] (5.2) In order to learn the allocation strategy, it is necessary to train a DDPG-based algorithm. The specific algorithm flow is as follows:

[0086] (5.2.1) Initialize the policy network, action network, and buffer, and initialize a stochastic process Ω for action learning;

[0087] (5.2.2) For each learning update period, observe the initial environment state s of the V2V task unloading. (1) ;

[0088] (5.2.3) Then generate an action a for each frame of each time period. (t) =μ(s) (t) |θ μ )+Ω (t) To determine the amount of tasks and transmission power assigned to the currently selected vehicle;

[0089] (5.2.4) Perform action a (t) Then receive the corresponding reward r(s) (t) ,a (t) ), updating observed new states s from the environment. (t+1) ;

[0090] (5.2.5) The information set (s) (t) ,a (t) ,r(s (t) ,a (t) ),s (t+1) ), stored in the experience replay buffer;

[0091] (5.2.6) Randomly sample a small subset of samples from the experience replay buffer to update the policy network:

[0092]

[0093] Where y (i) This represents the target value, y. (i) =r(s (i) ,a (i) )+γQ′(s (i+1) ,μ′(s (i+1) θ μ′ )|θ Q′ Q represents the evaluation network, Q′ represents the target evaluation network, and μ′ represents the target policy network. Q θ Q′ θ μ′ These represent the parameters of the corresponding networks;

[0094] (5.2.7) With the cooperation of the evaluation network, the policy network uses policy gradients to update parameters:

[0095]

[0096] (5.2.8) The target network uses a small constant τ to softly update the parameters:

[0097] θ Q′ =τθ Q +(1-τ)θQ′

[0098] θ μ′ =τθ μ +(1-τ)θ μ .

[0099] Compared with existing technologies, the technical solution adopted in this invention has the following technical effects: Compared with other task offloading schemes that balance communication latency and power consumption, this algorithm can achieve better performance in balancing transmission latency and power consumption while ensuring the success rate of task offloading. Attached image description:

[0100] Figure 1 This is a model diagram of a vehicle network system;

[0101] Figure 2 It represents the average success rate of V2V task unloading under different weighting coefficients α. Detailed implementation method:

[0102] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.

[0103] like Figure 1 The vehicle network shown includes one task vehicle (HV), one relay vehicle (REV), and multiple computing vehicles (RVs). First, 20 vehicles within 200m of the HV are randomly generated as computing vehicles. The communication range R of each vehicle is 200m, and each RV is allocated a CPU frequency of 20%F. n ~50% F n Between, F n Random values ​​are selected within the range of 3GHz to 8GHz. The distance between the vehicle (HV) and the RV varies within 50m at each moment, and the task size is uniformly distributed between 0.2 and 1 Mbits at each moment. The computational intensity ω0 is 1000 cycles / bit. The wireless channel state here follows an inverse power law h = A0l. -2 In this modeling, we assume A0 is -17.8 dB, l represents the distance between RV and HV, the path loss factor λ is 4, the channel bandwidth W is 10 MHz, and the noise power σ 2 10 -13 W, maximum transmission power P max The signal strength is 0.5W, the Weibull fading parameter C is 1.5, and the signal-to-noise ratio threshold is 15dB.

[0104] There is a task vehicle HV in the lane that needs to perform task calculations, and several distant vehicles RV with computing resources around the task vehicle HV. At each time step, HV has a value of x.t The task is calculated and broken down into multiple sub-tasks using an algorithm. These sub-tasks are offloaded to one or more RVs via the PC5 interface, or to RVs outside the communication range via a relay vehicle (REV). For single-hop V2V task offloading and multi-hop V2V task offloading using relay vehicles, the transmission delay for each offloading method is calculated based on the allocated task volume, channel state information, and allocated transmission power. Multi-hop task offloading requires calculating the end-to-end signal-to-noise ratio (SNR). The success probability of multi-hop task offloading is determined by comparing the end-to-end SNR with an SNR threshold. During task offloading, the task vehicle (HV) can choose the allocated task volume and transmission power. Based on the calculated transmission delay and success rate, the DDPG algorithm using reinforcement learning is used to obtain the optimal task offloading allocation and transmission power that meet the delay requirements at different times, while ensuring the task transmission success rate.

[0105] Figure 2 This represents the average success rate of V2V task offloading in the scheme when the weighting coefficient α is varied. The optimization problem balances latency and power by adjusting the control parameter α. Figure 2 As can be seen, for the three learning strategies DDPG, DDQN, and DQN, when the weight α of transmission power is 0, the average transmission power of the task is the maximum achievable value of 0.5W, and it decreases as the weight α of transmission power increases. The greater the weight of transmission power, the lower the average transmission power becomes, because task offloading decisions need to consider the trade-off between transmission power and latency; a higher weight for transmission power means a greater constraint on latency. Overall, the DDPG learning algorithm is more reliable in its trade-off between latency and power, and can obtain a better task offloading strategy.

[0106] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. A V2V communication multi-vehicle task offloading method based on deep reinforcement learning, characterized in that, Includes the following steps: (1) The computation task of the task vehicle HV is offloaded to multiple remote vehicles RV with computing resources in the vicinity, and the task is transmitted by multiple hops using the relay vehicle REV. (2) Calculate the transmission delay of single-hop V2V task offloading and the transmission delay of multi-hop V2V task offloading using relay vehicles respectively; (3) Calculate the transmission interruption probability of multi-hop V2V task offloading using relay vehicles; (4) Based on the transmission delay calculated in step (2) and the transmission interruption probability obtained in step (3), while ensuring the success rate of task unloading transmission, optimize the power and transmission delay of the task vehicle unloading process. (5) Based on the optimization objective of step (4), the deep reinforcement learning DDPG algorithm is used to solve the optimization problem; There is a probability of interruption when transmitting tasks via multi-hop relay vehicles. The end-to-end interruption probability is calculated by obtaining the distance between the multi-hop links of the vehicles, as follows: (4.1) For multi-hop V2V task offloading via relay vehicle REV, the vehicle-to-vehicle interruption probability is defined as the probability that the vehicle-to-vehicle signal-to-noise ratio is lower than the set limit value, i.e., the signal-to-noise ratio threshold. The vehicle-to-vehicle interruption probability is expressed as: (4.2) The amplitude of the signal received in the V2V channel follows a Weibull distribution, and the received signal-to-noise ratio f under Weibull fading is... SNR (a) is represented as: Where C represents the Weibull fading parameter, Γ represents the gamma function, and SNR... eq The vehicle-to-vehicle signal-to-noise ratio is expressed as follows: (4.3) Based on (4.1) and (4.2), the vehicle-to-vehicle interruption probability can be expressed as: (4.3) The probability of end-to-end interruption P outsum Represented as: in, σ 2 It is noise power; z indicates that there are z links in the communication process; L i This represents the distance between the two ends of the i-th hop link; P out (L i ) is the interruption probability of the i-th hop in the communication process.

2. The V2V communication multi-vehicle task offloading method based on deep reinforcement learning according to claim 1, characterized in that, In step (1), at each moment the mission vehicle has a size of x t The task is broken down into multiple subtasks through algorithm calculation, and then offloaded to one or more computing vehicles for calculation through the PC5 interface, or offloaded to computing vehicles outside the communication range through relay vehicles.

3. The V2V communication multi-vehicle task offloading method based on deep reinforcement learning according to claim 1 or 2, characterized in that, In step (2), the task vehicle can calculate the transmission delay of single-hop V2V task offloading and the transmission delay of multi-hop V2V task offloading using relay vehicles based on the task volume, computing resources of the computing vehicle, and channel status. The transmission latency for single-hop V2V task offloading without selecting relay vehicles is expressed as follows: in, x t This represents the size of the task data generated at time t; y t This indicates the amount of data returned after the task is uninstalled; Indicates the uplink transmission rate of HV and RVn; Indicates the downlink transmission rate of HV and RVn; ω t This represents the calculated strength of RV; f (t,n) This represents the CPU frequency allocated for processing HV tasks during time period t, where f (t,n) ∈[0,F n ]; F n This is the maximum CPU frequency of RVn; The specific uplink and downlink transmission rates for HV and RVn are expressed as follows: Among them, P t,n It is the transmission power allocated at each moment, W is the channel bandwidth, and σ is the transmission power allocated at each moment. 2 Where P is the noise power, P is the fixed transmit power given by the RVn feedback result, and h is the wireless channel state. t,n This represents the state between HV and RVn at time t. ) and It's an interference when the task is unloaded to RVn; The mission offloading transmission delay from the mission vehicle to RVn via relay multi-hop transmission is expressed as follows: in, d up (t,z) is the uplink wireless channel transmission delay at the z-th hop, and x is the transmission workload. t Compared to the uplink transmission rate of the z-th hop It is expressed as follows: d down (t,e) is the downlink wireless channel transmission delay at the e-th hop, and y is the data size. t Compared to the uplink transmission rate of the e-th hop It is expressed as follows: d com (t,n) represents the computation latency on the computing vehicle RV when the task is transmitted to the computing vehicle RV, as shown below:

4. The V2V communication multi-vehicle task offloading method based on deep reinforcement learning according to claim 1, characterized in that, In step (4), at each time step, in order to consider the trade-off between latency and power consumption, an optimization problem is proposed to minimize the overall consumption of the unloading process at the current time step: Simultaneously satisfy the constraints: d max (t,n)≤d th ζ1,ζ2∈[0,1],ζ1+ζ2=1 in, d th Threshold constraints representing latency; P max Threshold constraints representing power consumption; d max (t,n) represents the maximum transmission delay when multiple vehicles are selected; ζ1 and ζ2 represent the weights of latency and power consumption, respectively.