A Vehicle Edge Computing Task Sharing and Offloading Method Based on Deep Reinforcement Learning

By decomposing tasks into independent units in the Internet of Vehicles environment and utilizing the VEC server cache mechanism, combined with deep reinforcement learning to optimize the task unloading strategy, the problem of task sharing unloading in on-board edge computing is solved, achieving faster response speed and lower energy consumption.

CN115202757BActive Publication Date: 2025-07-22HUNAN INSTITUTE OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210834183.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2025-07-22
Estimated Expiration
2042-07-14

AI Technical Summary

Technical Problem

The existing on-board edge computing task unloading methods are difficult to achieve effective task sharing unloading in the Internet of Vehicles environment, resulting in network congestion, increased task response delay and excessive system energy consumption.

Method used

The task sharing and unloading method based on deep reinforcement learning is adopted. By decomposing the vehicle computing tasks into independent task units, the task result sharing is achieved using the VEC server's cache mechanism, and the task offloading strategy is optimized through deep reinforcement learning, and the optimal computing location is selected to reduce duplicate calculations and data transmission.

Benefits of technology

It effectively reduces the system response delay and energy consumption of task offloading, expands the operating space of task offloading, and improves the energy efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115202757B_ABST
    Figure CN115202757B_ABST
Patent Text Reader

Abstract

To reduce the computing load of edge servers and improve the system response speed, a task sharing and offloading method based on deep reinforcement learning is proposed. It includes: (1) decomposing in-vehicle tasks into relatively independent task units and assigning a complete task unit ID field; (2) caching all offloaded computing task units in the ID pool of the VEC server, and the ID pool updates cached tasks according to the effective time of task units and the least recently used principle; (3) establishing a matching mechanism between newly created task units and cached task units, and only task units that match successfully can share task calculation results; (4) establishing a task offloading utility function and its optimization model for computing latency, energy consumption, and computing cost, and equivalent the overall optimal task offloading scheme to a Markov process; (5) completing the selection of the optimal offloading scheme through deep reinforcement learning methods based on the task unit offloading state space and policy space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an Internet of Vehicles (IoVs) task offloading strategy, and in particular to a vehicular edge computing task sharing offloading method based on deep reinforcement learning. Background Art

[0002] With the continuous popularization of intelligent connected vehicles, the Internet of Vehicles (IoVs) undertakes a large number of application services, including autonomous driving, path planning, collision avoidance, in-vehicle entertainment, etc. The application services of IoVs are usually computationally and communication-intensive tasks, which require a large amount of computing and energy consumption overhead, and thus it is difficult to rely on the computing power of vehicles for local processing. With the continuous popularization of 5G networks, a large number of servers are deployed on the edge side close to network users to provide effective computing power support for IoVs application services, realizing vehicular edge computing (VEC). By offloading the computing tasks of vehicles to the edge server for execution, the computing load of vehicles is effectively reduced, thereby reducing the response latency of applications and extending the battery life of vehicles. In order to improve the efficiency of VEC server deployment, reduce system energy consumption and task response latency, the VEC server is deployed at the edge of the radio access network to bring rich computing power closer to the terminal vehicle users. VEC servers are usually deployed at fixed locations such as base stations (BS) and roadside units (RSU) to improve the flexibility of the VEC network.

[0003] Although vehicular edge computing can effectively expand the computing power of vehicles, how to reasonably and effectively offload the tasks of vehicles to the VEC server is a major problem. If a large number of tasks are offloaded to the VEC server for execution, it will cause network congestion, increase task response latency and system energy consumption. Both BS and RSU can provide network coverage and VEC server access services for vehicles. Therefore, how to effectively offload tasks in the network coverage environment of BS and RSU is a key issue of VEC. The existing VEC task offloading decision mechanism mainly constructs a global optimization model based on the current state of tasks and edge servers for task offloading decision. The problem of global optimization is generally a non-convex function, which does not satisfy the convex optimization conditions and is difficult to perform global optimization. For the distributed architecture of IoVs, global information is usually difficult to obtain, and even if global information is obtained, it may not have convex optimization conditions.

[0004] Currently, no existing work has proposed a VEC task sharing offloading method. Summary of the Invention

[0005] Based on this, it is necessary to provide a vehicular edge computing task sharing offloading method based on deep reinforcement learning for the above problems.

[0006] There are various application services in the vehicle networking, and computing tasks can be decomposed into relatively independent task units, and many task units are the same (for example, in the navigation path planning calculation, navigation calculation requests from many different users may all need to obtain traffic data between a certain two intersections). For the same task units from different vehicles, they actually have similar calculation results. At this time, only one calculation needs to be performed, and all the same task units can share the calculation results without having to perform repeated calculations for each task. Therefore, the task offloading problem in the case of having the same task units needs to be further explored. Based on this consideration, the present invention proposes a VEC task sharing and offloading method based on DRL (Deep Reinforcement Learning). The proposed method optimizes the task offloading of VEC with energy efficiency as the goal through DRL. For the same task units, the VEC server can directly retrieve the results from the cache and return them to the users, effectively expanding the optimization space for task unit offloading. The proposed method realizes the global task offloading optimization of VEC through DRL, without having to execute the entire calculation process for each computing task, with a faster response speed and lower energy consumption.

[0007] A VEC task sharing and offloading method based on DRL includes the following steps:

[0008] Step 1: Establishment of task units. The computing tasks generated by vehicles are decomposed into relatively independent task units according to functional modules, and each task unit is assigned a unique task ID. The task ID includes the task unit size d i , the required computing amount c i , the task execution time limit the task function ζ i , the generation time and the valid duration Δt of the task calculation result i and other six fields. The task unit size d i determines the communication overhead during task offloading. The required computing amount c i of the task unit determines the computing overhead of the task unit. The task execution time limit indicates that the task execution result must be returned within the time, otherwise the task execution fails. The task function ζ i is determined by the functional attributes of the task itself, and task units with the same ζ i are task units with the same function. The generation time is the generation moment of the task unit for the tasks generated by the vehicle, and for the task units cached in the VEC server, this field is the moment when the task unit enters the server cache. The valid duration Δt of the calculation result i indicates the valid time of the calculation result of this task unit.

[0009] Step 2: Execution method of task unit offloading. The task offloading method first needs to make a trade-off between local computing and VEC server computing. Vehicle computing tasks can be executed in a variety of different ways, including local execution and task offloading. When the in-vehicle computing capacity is insufficient, the task unit can be offloaded to the VEC server of the RSU or BS through a wireless communication link for computing. The IDs of all offloaded task units and their computing results will be saved in the ID pool of the VEC server. The ID pool of the VEC server of the BS is represented by IDStack BS and the ID pool of the VEC server of the RSU is represented by IDStack RSU When the ID pool data is full, the new task unit ID will overwrite the previous data. The overwrite mechanism uses the least recently used principle. The IDs that have not been matched for a long time are also less likely to be used in the future. When new IDs come in, these task units can be overwritten first.

[0010] If the ID function field ζ of the newly arrived task unit j is the same as that of the task unit i that has been cached in the VEC server, and the generation time of the new task unit j i is within the valid time Δt of the computing result of the cached task unit i, that is then the task unit matching is successful. The successful matching of the task unit means that the newly arrived task unit j can share the computing result of the task unit i that has been cached on the VEC server. At this time, if offloading is required, the VEC server can directly return the computing result of the task unit i to the requesting vehicle. Due to the limited coverage of the RSU and the high communication and computing costs of the BS, when the execution time limit i of the task unit is long, the vehicle can choose to complete the computing of the task locally. If the execution time limit of the task unit is very short, when the network load is large, it may not be possible to offload the task unit to the VEC server for execution in time, and the task unit will fail due to timeout. Different computing locations have different computing capabilities and computing costs. The comprehensive consideration of the computing capabilities, maximum latency, service costs, and offloading energy consumption required for each task determines the computing location of the task. of the task unit is long, the vehicle can choose to complete the computing of the task locally. If the execution time limit of the task unit is very short, when the network load is large, it may not be possible to offload the task unit to the VEC server for execution in time, and the task unit will fail due to timeout. Different computing locations have different computing capabilities and computing costs. The comprehensive consideration of the computing capabilities, maximum latency, service costs, and offloading energy consumption required for each task determines the computing location of the task.

[0011] Step 3: Establish a task unit offloading model. The present invention defines the utility function R i (t) for offloading and computing of the task unit i. R i (t) includes three major parts: the remaining delay of the task unit, the computing energy consumption of the task unit, and the offloading service cost. The longer the remaining delay of the task unit, the greater the utility, and when the computing delay exceeds the maximum delay There is a timeout penalty. The service cost includes the transmission cost of the task unit to the BS and the computing cost at the VEC server. The RSU communication is free of charge. R i (t) is in the following form:

[0012]

[0013] Among them, φ(i) is the computing time of task unit i, E i is the computing energy consumption of task unit i, L i is the normalized value of the service cost, u(·) is the step function, α1, α2, α3 are weight coefficients, and α1 + α2 + α3 = 1, κ1 and κ2 are constants close to 1, λ i is an arbitrary constant. R i The delay, energy consumption, and cost parameters in (t) all include the cost and benefit brought by shared offloading. The sum of the utility functions of each task unit in cycle t is defined as the reward function U(t) = ∑ i∈N R i (t). The policy space Π t of task offloading is defined as the set of all possible offloading strategies of task unit i in time period t. Then the optimal task offloading decision in cycle t can be modeled as maximizing the reward function of the task unit, and its constraints include: ① the lower limit of the transmission power when the vehicle communicates with the RSU; ② the lower limit of the transmission power when the vehicle communicates with the BS; ③ there can only be one offloading scheme for task unit i in cycle t; ④ if task unit i fails to match on the BS server, it cannot perform shared offloading on the BS server; ⑤ if task unit i fails to match on the RSU server, it cannot perform shared offloading on the RSU server. The failure to match in ④ and ⑤ includes two cases. One is that the newly arrived task is a brand-new task and there is no cache of this task in the VEC server, so shared offloading cannot be performed; the other is that the newly arrived task has exceeded the validity period of the task cache result in the VEC server, and shared offloading cannot be performed either. It can be seen from the optimization model that the optimization goal is to maximize the reward function. The proposed shared offloading scheme of task units effectively expands the decision space of the model, and the optimization of the large decision space model provides more suitable basic conditions for the application of deep reinforcement learning.

[0014] By maximizing the reward function, the optimal offloading strategy of the task unit in cycle t can be obtained Then the overall optimal offloading scheme P * in the process of task offloading time can be expressed as:

[0015]

[0016] where \(0 < \gamma < 1\) is the discount factor representing the impact of future long-term utility, \(T\) is the set of time periods, and \(t\) is the time period index. is the mathematical expectation.

[0017] Step 4: Initialize the task unit offloading strategy. At any time period \(t\in\{1,2,\ldots,T\}\), there are five offloading execution modes for task unit \(i\), and its offloading mode index is defined as indicating the offloading execution mode of task unit in period \(t\). means that task unit \(i\) is offloaded to the VEC server of the BS through V2B communication for computing. means that task unit \(i\) performs shared offloading computing on the VEC server of the BS. means that task unit \(i\) is offloaded to the VEC server of the RSU through V2R (Vehicle to RSU) communication for computing. means that task unit \(i\) performs shared offloading computing on the VEC server of the RSU. means that task unit \(i\) executes locally. Each task unit can only have one offloading computing mode, that is, only one of the values in can be 1.

[0018] Step 4: Before optimizing the model, first give the initial value of the offloading strategy according to the required transmission time and the maximum time limit of the task. If the task is only offloaded to the VEC server of the BS through V2B (Vehicle to BS) communication, the computing delay (the superscript "B" represents the base station BS, "co" represents computing) and the transmission delay (the superscript "B" represents the base station BS, "tr" represents transmission) of the task unit satisfy the task execution time limit that is and the computing time (the superscript "R" represents the roadside unit RSU, "co" represents computing) of the task unit on the VEC server of the RSU and the time required for the task unit to be transmitted to the RSU (the superscript "R" represents the roadside unit RSU, "tr" represents transmission) satisfy At this time, if the task unit ID does not match successfully, the offloading mode index is initialized as If the task unit ID matches successfully, the offloading mode index is initialized as Similarly, if the computing delay and the transmission delay of the task unit offloaded to the VEC server of the RSU satisfy the task execution time limit that is At this time, if the task unit ID does not match successfully, the offloading mode index is initialized. If the task unit ID matches successfully, the offloading mode metric is initialized. If the local computing time of the vehicle can meet the execution time limit of the task unit That is If the networks of the BS and RSU are unavailable, or the offloading of the VEC server cannot meet the execution time limit, then the local computing mode will be the only option, and the offloading mode metric is initialized. The specific initialization process is shown in Algorithm 1 in the detailed implementation manners.

[0019] Step Five: Offloading decision based on deep reinforcement learning. The best offloading strategy P * is selected depending on the current vehicle networking channel state, the historical execution of the task unit on the VEC server, the computing power of the VEC server, and the reward function U(t), etc.

[0020] First, define the state space S of the VEC task offloading t :

[0021]

[0022] Wherein and are the communication transmission rates from the vehicle to the BS and RSU respectively; f B , f R , f L are the CPU frequencies of the BS server, RSU server, and vehicle local computing unit respectively; is the computing waiting time of task unit i, IDStack BS and IDStack RSU are the ID pools of the BS server and RSU server respectively. The ID pool caches the task unit IDs that have been offloaded and executed on the server and their computing results, providing support for the shared offloading of subsequent task units; ID i represents the complete ID of task unit i.

[0023] The task offloading process is equivalent to a Markov decision process. Therefore, the value function Q(S t , π t ) is defined as the long-term expected value of the task offloading reward function based on the state space S t and the task unit offloading strategy π t :

[0024]

[0025] The present invention finds the best task offloading strategy for each task unit in the time series by updating the value function. The value function update process is expressed as the time difference method, in the following form:

[0026]

[0027] where β is the learning rate, and Q * (S t , π t ) is the optimal value of the value function.

[0028] The present invention uses a convolutional neural network to construct a target network and an evaluation network respectively. The target network calculates the value function of the computing offloading strategy π t of

[0029]

[0030] where θ t is the parameter of the convolutional neural network. The evaluation network finds the best offloading strategy according to the current state through the value function Q(S t , π t , θ t ). The present invention constructs a loss function Loss(θ t ) to measure the difference between the target network and the evaluation network:

[0031]

[0032] According to the loss function, the gradient descent method is used to update θ t of the evaluation network:

[0033]

[0034] Then, θ is updated according to the following formula:

[0035]

[0036] where is the scalar step size. Optimizing the best value function through the above iterative process is equivalent to finding the best overall task offloading scheme P * .

[0037] Beneficial effects of the proposed solution in this application: The VEC task sharing and offloading method proposed by the present invention adopts a DRL algorithm to construct an effective sharing and offloading solution in the dynamic environment of the vehicle network without a network prior model. Compared with existing work, the decision space for task offloading is larger, and it has strong adaptability to the complex dynamic environment of the vehicle network. The computing tasks from different vehicles can be decomposed into relatively independent task units. For task units with the same functional field, the computing results of previous task units can be shared within the task validity period. The proposed method adaptively selects three types of task computing modes according to tasks, VEC server computing resources, and network parameters: local execution, task offloading, and shared offloading. Compared with existing methods, it can effectively reduce the system response delay and energy consumption of task offloading.

[0038] The shared offloading method utilizes the commonality of vehicle task requests to avoid unnecessary repeated calculations and data transmissions. Compared with existing typical vehicle network task offloading solutions, the shared offloading mechanism effectively reduces the computing time and transmission time of task units. The method proposed by the present invention can effectively expand the operation space of task offloading and reduce the energy consumption and cost required for task offloading.

[0039] The present invention calls a deep neural network, which can more intelligently consider task offloading strategies and improve the overall energy efficiency of the system. Description of the Drawings

[0040] Figure 1 is for the vehicle network edge computing environment;

[0041] Figure 2 is for the task unit ID field structure and offloading process;

[0042] Figure 3 is for the task shared offloading mechanism;

[0043] Figure 4 is for the deep reinforcement learning task offloading process;

[0044] Figure 5 is for the total utility values of different methods;

[0045] Figure 6 is for the average reward function values of different methods;

[0046] Figure 7 is for the average remaining delay of different methods;

[0047] Figure 8 is for the average energy consumption of different methods;

[0048] Figure 9 is for the service costs of different methods;

[0049] Figure 10Relationship between the loss function and the training steps of the method in this paper. Detailed implementation

[0050] 1. Vehicular network edge computing environment

[0051] In the vehicular network computing scenario, in addition to the cellular network provided by the base station BS, there is also a DSRC (Dedicated Short-Range Communication) network provided by the roadside unit RSU. The cellular network and the DSRC network operate on different and non-overlapping spectrums. Compared with the cellular network with seamless coverage and high data transmission costs, DSRC has limited network coverage and low-cost access services. Both BS and RSU are equipped with their own VEC servers. Vehicles can offload tasks to the corresponding VEC servers through V2B (Vehicle to BS) and V2R (Vehicle to RSU) communication methods, where V2B is applicable to the task offloading scenario where vehicles use the cellular network, and V2R is applicable to the task offloading scenario where vehicles use the DSRC network. The vehicular network edge computing environment is as Figure 1 shown.

[0052] 2. Establishment of task units

[0053] Decomposing the computing tasks generated by vehicles into relatively independent task units according to functional modules is conducive to the formation of the task offloading decision space and the implementation of shared offloading. Each task unit is assigned a unique task ID. The task ID includes the task unit size d i , the computing amount c i required by the task unit, the task execution deadline task function ζ i , the generation time and the effective duration Δt i of the task computing result, etc. The six fields. The task unit size d i determines the communication overhead during task offloading. The computing amount c i required by the task unit determines the computing overhead of the task unit. The task execution deadline indicates that the task execution result must be returned within time, otherwise the task execution fails. The task function ζ i is determined by the functional attributes of the task itself. Task units with the same ζ i are task units with the same function. The generation time for the tasks generated by vehicles is the generation moment of the task unit, while for the task units cached in the VEC server, this field is the moment when the task unit enters the server cache. The computing result effective duration Δt i indicates the effective time of the computing result of this task unit. The task unit ID field structure is asFigure 2 as shown

[0054] 3. Execution Modes for Unloading Task Units

[0055] The shared unloading method of the present invention takes energy consumption, latency, and cost as the comprehensive optimization objectives. Both local computing and task unloading incur energy consumption and latency. The task unloading method first needs to make a trade-off between local computing and computing on the VEC server. Vehicle computing tasks can be executed in a variety of different ways, including local execution and task unloading. As Figure 3 shown, when the in-vehicle computing power is insufficient, task units can be unloaded via a wireless communication link to the VEC server of the RSU or BS for computing. The IDs of all unloaded task units and their computing results will be stored in the ID pool of the VEC server. The ID pool of the VEC server of the BS is represented by IDStack BS and the ID pool of the VEC server of the RSU is represented by IDStack RSU When the ID pool data is full, new task unit IDs will overwrite the previous data. The overwrite algorithm uses the least recently used principle. IDs that have not been matched for a long time are also less likely to be used in the future. When new IDs come in, these task units can be overwritten first.

[0056] If the ID function field ζ i of the newly arrived task unit j is the same as that of the task unit i that has been cached in the VEC server, and the generation time of the new task unit j is within the valid time Δt i of the computing result of the cached task unit i, that is then the task unit matching is successful. The successful matching of the task unit means that the newly arrived task unit j can share the computing result of the task unit i that has been cached on the VEC server. At this time, if unloading is required, the VEC server can directly return the computing result of the task unit i to the requesting vehicle. For convenience, the time consumption of the VEC server directly extracting the computing result from the cache is ignored.

[0057] Due to the limited coverage of the RSU and the high communication and computing costs of the BS, when the execution time limit of the task unit is long, the vehicle can choose to complete the computing of the task locally. If the execution time limit of the task unit is very short, when the network load is large, it may not be possible to unload the task unit to the VEC server for execution in time, and the task unit will fail due to timeout. Different computing locations have different computing capabilities and computing costs. The computing capabilities and maximum latency required for each task determine the computing location of the task.

[0058] 4. Initialization of Task Unloading Decision

[0059] At any time period \(t\in\{1,2,\ldots,T\}\), there are five offloading and execution modes for task unit \(i\), and the offloading mode index is defined as indicating the offloading and execution mode of the task unit in period \(t\). indicating that task unit \(i\) is offloaded to the VEC server of the BS through V2B communication for calculation, indicating that task unit \(i\) performs shared offloading calculation on the VEC server of the BS, indicating that task unit \(i\) is offloaded to the VEC server of the RSU through V2R communication for calculation, indicating that task unit \(i\) performs shared offloading calculation on the VEC server of the RSU, indicating that task unit \(i\) is executed locally. Each task unit can only have one offloading calculation mode, that is, only one of them can take the value of 1, and the corresponding constraint is expressed as follows:

[0060]

[0061] Before model optimization, the initial value of the offloading strategy is given according to the transmission time and maximum time limit required by the task. If the task is only offloaded to the VEC server of the BS through V2B communication, the computing delay and transmission delay of the task unit satisfy the task execution time limit that is, and the computing time of the task unit on the VEC server of the RSU and the time required for the task unit to be transmitted to the RSU satisfy At this time, if the task unit ID does not match successfully, the offloading mode index is initialized If the task unit ID matches successfully and shared offloading can be performed at the BS, the offloading mode index is initialized Similarly, if the computing delay and transmission delay of the task unit offloaded to the VEC server of the RSU satisfy the task execution time limit that is, At this time, if the task unit ID does not match successfully, the offloading mode index is initialized If the task unit ID matches successfully and shared offloading can be performed at the RSU, the offloading mode index is initialized If the local computing time of the vehicle can meet the execution time limit of the task unit that is, If the network of the BS and RSU is unavailable, or the offloading of the VEC server cannot meet the execution time limit, the local computing mode will be the only option, and the offloading mode metrics will be initialized.

[0062] Based on the initialization process, the present invention constructs an optimal offloading strategy for task units targeting energy consumption and latency through a deep learning method. The initialization process of the offloading decision of the VEC is shown in Algorithm 1.

[0063]

[0064]

[0065] 5. Formal Expressions of Latency and Energy Consumption

[0066] (1) Latency and Energy Consumption Generated by Task Local Computing

[0067] If task unit i is locally computed by the vehicle itself The required computing time Can be expressed as:

[0068]

[0069] Where c i Represents the required computing volume of task unit i, and f L Represents the CPU frequency of the vehicle-mounted computing platform.

[0070] The energy consumed by task unit i during local computing by the vehicle Is expressed as:

[0071]

[0072] Where ω L Is the computing power coefficient of the vehicle-mounted computing platform.

[0073] (2) Energy Consumption and Latency Generated by Task Offloading Computing

[0074] The vehicle offloads task unit to the VEC server for computing through the BS Or shares the computing results of historical tasks And offloads task unit to the VEC server for computing through the RSU Or shares the computing results of historical tasks All belong to the category of task offloading strategies. The transmission power of the vehicle in the V2B transmission mode is The transmission power in the V2R transmission mode is The cellular link spectra used by vehicles in the V2B transmission mode are orthogonal. According to Shannon's formula, the transmission rate when the vehicle offloads tasks to the BS can be expressed as:

[0075]

[0076] where W is the system communication bandwidth, is the channel gain for V2B transmission, and σ 2 is the noise power.

[0077] Vehicles adopt the DSRC communication method in the V2R mode. DSRC has a total of K service channels and uses a competitive access method. When the number of accessing users is no more than K, each channel does not interfere with each other. When the number of users is greater than K, channel collisions will occur, and the transmission behaviors of the extra users will cause interference at this time. The number of vehicle users communicating with the RSU in the current service area is M. Define the co-channel interference power as:

[0078]

[0079] where ρ is the average interference coefficient generated by over-capacity users, generally taking The transmission rate when the vehicle communicates with the RSU in the V2R mode in this case is:

[0080]

[0081] where σ 2 is the noise power, ρ k is the co-channel interference coefficient, is the channel gain of the k-th channel for V2R transmission.

[0082] Under the V2B and V2R offloading strategies, if the task unit is executed by the VEC server, the task unit ID and the task unit content need to be migrated to the VEC server as a whole. If the task unit shares the historical calculation results on the VEC server, only the task unit ID field needs to be transmitted to the VEC server. Therefore, the delay of the task unit i transmitted to the VEC server of the BS and the delay

[0083]

[0084]

[0085] where δ i is the length of the ID of the task unit i. Adding the marker variable and Used to distinguish between shared offloading and non-shared offloading scenarios. The computing delay of task unit i in the VEC server of the BS and the computing delay of the VEC server of the RSU are respectively expressed as:

[0086]

[0087]

[0088] Since shared offloading directly calls the task results cached in the VEC server ID pool, the computing delay of task units during shared offloading in the VEC server is not considered. The energy consumption generated when task units are offloaded to the VEC server of the BS through V2B communication and the energy consumption generated when offloaded to the VEC server of the RSU through V2R communication are respectively expressed as:

[0089]

[0090]

[0091] Among them, ω B and ω R respectively represent the computing power coefficients of the VEC server of the BS and the VEC server of the RSU, f B and f R respectively represent the CPU frequencies of the VEC server of the BS and the VEC server of the RSU. The marker variables and are introduced to distinguish the computing energy consumption of the server during shared offloading and non-shared offloading. If the task unit adopts the shared offloading method on the VEC server, only the transmission energy consumption is considered without considering the computing energy consumption. Different task offloading methods and task computing locations have different energy consumptions. For task unit i, its energy consumption E i can be expressed as the sum of the energy consumptions under five different offloading execution methods including shared offloading:

[0092]

[0093] Task unit i needs to queue and wait when offloaded to the VEC server. Therefore, the waiting time when the task unit starts to transmit is defined as:

[0094]

[0095] Actually, it includes the transmission and calculation delays of the previous tasks. Since there are different offloading execution methods such as shared offloading for the previous tasks, the execution delays are also different. Therefore, marker variables are introduced in the above formula for distinction. The total calculation time φ of task unit i during the offloading process i,t includes its own waiting time, transmission time, and calculation time under different calculation methods such as shared offloading, non-shared offloading, and local calculation, and can be expressed as:

[0096]

[0097] where represents the time for task unit i - 1 to perform calculations through the BS, represents the time for task unit i - 1 to perform calculations through the RSU.

[0098] 6. Deep Learning Method for Task Shared Offloading

[0099] The present invention defines the utility function R i (t) of task unit i for offloading calculations, which includes three major parts: the remaining delay of the task unit, the calculation energy consumption of the task unit, and the offloading service cost. The longer the remaining delay of the task unit, the greater the utility. And when the calculation delay exceeds the maximum delay of the task unit , there is an overtime penalty. The service cost includes the transmission cost of the task unit to the BS and the calculation cost on the VEC server. The RSU communication is free of charge. Different offloading calculation methods have different overheads. The shared offloading method does not need to transmit task data to the VEC server and thus has no calculation cost. Offloading to the VEC server of the BS requires task data transmission cost and server calculation cost. Offloading to the VEC server of the RSU does not require data transmission cost but has to pay server calculation cost. Therefore, the normalized value L of the service cost i is defined as

[0100]

[0101] where v B is the transmission charging unit of the BS, and v C is the charging unit for using the server.

[0102] Definition 1. Utility function. The utility function is the normalized weighted sum of the remaining delay, the calculation energy consumption of the task unit, and the offloading service cost during the offloading process of task unit i. The utility function is expressed as:

[0103]

[0104] where u(·) is the step function, α1, α2, and α3 are weight coefficients, and α1 + α2 + α3 = 1, κ1 and κ2 are certain constants close to 1, and λ iis an arbitrary constant. R i The delay, energy consumption, and cost parameters in R(t) all include the cost-benefit brought by shared offloading.

[0105] Based on the utility function R i (t), the overall utility value of offloading all tasks is defined as the reward function.

[0106] Definition 2. Reward function. The reward function is the overall utility value generated by offloading all task units, used to evaluate the quality of the current task unit offloading scheme. The reward function is expressed as:

[0107]

[0108] Definition 3. Policy space. The policy space of the task offloading method is the set of all possible offloading strategies of task unit i within time period t, expressed as:

[0109]

[0110] It can be seen from Definition 3 that due to the adoption of the shared offloading method, task units have a larger offloading policy space. Within period t, the goal of task offloading decision optimization is to select the best offloading execution plan for each task unit, so as to maximize the reward function of the task unit. The optimization model is established as follows:

[0111]

[0112] Among the constraint conditions of the above model, C1 represents the lower limit of the transmission power when the vehicle communicates with the RSU and C2 represents the lower limit of the transmission power when the vehicle communicates with the BS ; C3 means that task unit i can only have one offloading plan within period t; C4 means that if task unit i fails to match successfully on the BS server, it cannot perform shared offloading on the BS server; C5 means that if task unit i fails to match successfully on the RSU server, it cannot perform shared offloading on the RSU server. The unsuccessful matching in C4 and C5 includes two cases. One is that the newly arrived task is a brand-new task and there is no cache of this task in the VEC server, so shared offloading cannot be performed; the other is that the newly arrived task has exceeded the validity period of the task cache result in the VEC server and shared offloading cannot be performed either. It is not difficult to see from the optimization model that the optimization goal is to maximize the reward function. The proposed shared offloading scheme of task units in the present invention effectively expands the decision space of the model, and the optimization of the large decision space model provides a more suitable basic condition for the application of deep reinforcement learning.

[0113] By maximizing the optimization goal of the above model, the optimal offloading strategy of the task unit in period t can be obtained Then the overall optimal offloading scheme in the task offloading time process can be expressed as:

[0114]

[0115] Where \(0 < \gamma < 1\), \(\gamma\) is the discount factor representing the impact of future long-term utility, \(T\) is the set of time periods, \(t\) is the time period serial number, is the mathematical expectation.

[0116] The selection of the optimal offloading strategy \(P\) * depends on the current vehicle network channel state, the historical execution situation of the task unit on the VEC server, the computing power of the VEC server, and the reward function \(U(t)\), etc. Therefore, first define the state space and value function of VEC task offloading.

[0117] Definition 4. State space. The state space of the DRL shared offloading method is:

[0118]

[0119] Where, and are the communication transmission rates from the vehicle to the BS and RSU respectively; \(f\) B , \(f\) R , \(f\) L are the CPU frequencies of the BS server, RSU server, and vehicle local computing unit respectively; is the computing waiting time of task unit \(i\), IDStack BS and IDStack RSU are the ID pools of the BS server and RSU server respectively, used to cache the executed task IDs and computing results to support the shared offloading of subsequent tasks; ID i represents the complete ID of task unit \(i\).

[0120] Definition 5. Value function. The task offloading process is equivalent to a Markov decision process, so the value function is defined as the long-term expected value of the task offloading reward function based on the state space \(S\) t and the task unit offloading strategy \(\pi\) t :

[0121]

[0122] The present invention finds the optimal task offloading strategy of each task unit in the time series by updating the value function. Express the value function update process as the time difference method, in the following form:

[0123]

[0124] According to the above formula, the optimal value \(Q\) of the value function* (S t , π t ) is expressed as:

[0125]

[0126] According to the DRL algorithm, the update strategy of the value function can be expressed as:

[0127]

[0128] where β is the learning rate.

[0129] The present invention uses a convolutional neural network to construct a target network and an evaluation network respectively. The target network calculates the offloading strategy π t of the value function The evaluation network finds the optimal offloading strategy according to the current state through the value function Q(S t , π t , θ t ), where θ t are the parameters of the convolutional neural network. The present invention uses the loss function Loss(θ t ) to measure the difference between the target network and the evaluation network:

[0130]

[0131] where θ t are the parameters of the convolutional neural network. The present invention uses the gradient descent method to update θ of the evaluation network t :

[0132]

[0133] Then, θ is updated according to the following formula:

[0134]

[0135] where is the scalar step size.

[0136] The pseudo-code of the deep reinforcement learning algorithm is as follows:

[0137]

[0138]

[0139] 7. Experimental evaluation

[0140] To illustrate the performance of the proposed shared offloading strategy, the present invention considers an area of 120 square kilometers. Each cell has 1 base station and 0 - 3 RSUs. Each cell has 1 - 3 vehicle users performing computing tasks in each time slot. Other evaluation parameters are listed in Table 1. The present invention compares the proposed DRL - based shared offloading method with 5 other methods, namely: "Q - learning" represents the offloading method based on Q - learning for the scheme where all task units are executed; "Greedy" refers to the scheme where all task units execute the offloading method based on time limit or the local computing method; "Unshared offloading" refers to the non - shared offloading method based on DRL where all task units are executed; "Offloading only" refers to the method where all task units perform computing through offloading; "Local only" represents the method where all task units are executed locally.

[0141] Table 1 Main experimental parameters

[0142]

[0143]

[0144] The present invention modifies the DQN algorithm of DRL and applies it to the offloading optimization of each strategy. In the design of the shared offloading strategy, the learning rate of the neural network model is 0.1. The exploration degree is 0.9, that is, there is a 90% probability of selecting the best action and a 10% probability of selecting a random action, allowing the algorithm to explore all possibilities in the environment. There are two uncorrelated networks in the shared offloading strategy network, and each network has two - layer neural networks. Each layer is a fully - connected layer composed of 20 neurons and uses the Relu activation function. When the network selects an action - intensive action, the environment feeds back the energy consumption and the remaining delay of the action to the agent, and then updates the network in turn. The buffer size is 500 and the batch size is 32. In Q - learning, the learning rate is also 0.01 and the discount factor gamma is 0.9. Both the shared offloading strategy and Q - learning have 200 learning times.

[0145] Figure 5 It shows that with the accumulation of cycles, the reward of the method of the present invention gradually reaches the maximum value, and the final reward value is significantly higher than that of the methods based on Q - learning and Greedy. The non - intelligence of the Greedy method results in poor processing delay and energy consumption in each time period, which is a penalty for timeout. The present invention calls the deep neural network, which can more intelligently consider the task offloading strategy and improve the overall energy efficiency of the system.

[0146] Figure 6It shows that the reward value obtained by the present invention is significantly better than other methods. This indicates that the method of the present invention has achieved remarkable improvement in the comprehensive optimization of energy consumption, time delay and service cost. The reward values of Q-learning and Unshared offloading methods are only second to our method. Due to the learning mechanism, the reward values of the three methods all have a gradually increasing process. Due to the lack of a flexible decision-making mechanism, the reward values of only the Local only and Offloading only methods are the lowest.

[0147] Figure 7 It shows that after 3000 rounds, the remaining time of the task unit in the present invention is significantly greater than other methods. This is because the shared offloading mechanism effectively reduces the computing time and transmission time of the task unit. Due to limited computing power, the remaining time of the Local only method is the least. The Offloading only method needs to migrate all tasks to the VEC server, which will cause network congestion and further reduce the remaining time of the task unit. Q-learning can make comprehensive decisions on task offloading, and its performance is relatively close to our method.

[0148] Figure 8 It shows the comparison of the energy consumption of each method. After 3000 cycles, the energy consumption generated by the present invention is better than other methods. This is because the present invention reduces the data transmission volume and task computing volume. Since the Local only method has no data transmission, the energy consumption is relatively low. The Offloading only method requires a large amount of task transmission, so the energy consumption is high. Q-learning and Unshared offloading methods take energy consumption as one of the optimization objectives, and the energy consumption is also relatively low.

[0149] Figure 9 It shows the BS service charges of different methods. Since the Local only does not consume the offloading cost of the base station, the service fee of the base station is 0. However, a large number of task units cannot be executed due to timeout. The service fee for offloading all tasks to the VEC server is the highest. The Greedy method adopts the principle of local computing first, and the service fee is low, but the system energy consumption and delay are large. Since the service fee is only one of the system utility optimization objectives, the service fee of the present invention is larger than that of the Local only and the greedy offloading algorithm, but significantly lower than the similar Unshared Offloading and Q-learning methods.

[0150] Figure 10Shows the convergence of the present invention. It is not difficult to see from the figure that after 400 iterations of training, the value function of the present invention tends to be stable. This is because the algorithm uses the gradient descent method for the loss function, effectively improving the convergence performance. In addition, the fast convergence also improves the adaptability of the method.

Claims

1. A vehicle-mounted edge computing task sharing and offloading method based on deep reinforcement learning, characterized in that Including the following steps: Step 1: Task unit establishment; The computing tasks generated by the vehicle are decomposed into relatively independent task units according to functional modules, and each task unit is assigned a unique task unit ID; The task unit ID includes the task unit size d i , the required computing amount c i , the task execution time limit task function ζ i , generation time and the valid duration Δt of the task calculation result i six fields; Step 2: Execute task unit offloading; First, make a trade-off between local computing and VEC server computing; When the in-vehicle computing power is insufficient, the task unit is offloaded to the VEC server of the RSU or BS through a wireless communication link for calculation. The IDs of all offloaded task units and their calculation results will be stored in the ID pool of the VEC server. When the ID pool is full, the ID of the new task unit will overwrite the previous data. The overwrite algorithm uses the least recently used principle. The ID that has not been matched for a long time is less likely to be used in the future. When a new task unit comes in, these cached task units can be overwritten first. If the new task unit has the same functional field as the cached task unit and is within the valid duration Δt i of the calculation result of the cached task unit, the match is successful, and the new task unit can share the calculation result cached by the server; Step 3: Establish a task unit offloading model; define the utility function \(R i (t)\) for offloading calculation of task unit \(i\), and define the sum of the overall utility functions of task units as the reward function \(U(t)\). \(R i (t)\) includes three main parts: the remaining delay of the task unit, the computing energy consumption of the task unit, and the offloading service cost. Step 4: Initialize the task unit offloading strategy; task unit i has five execution modes: BS offloading execution, BS shared offloading execution, RSU offloading execution, RSU shared offloading execution, and local execution, which respectively correspond to the flag variables The flag variable taking a value of 1 indicates that task unit i adopts its corresponding execution mode; before model optimization, first compare the time required for task unit computing in the BS, RSU, and locally with the maximum time limit of the task unit, and at the same time give the initial value of the offloading strategy according to the possibility of task unit i sharing computing results in the VEC servers of the BS and RSU; Step 5: Offloading decision based on deep reinforcement learning; optimal offloading policy P * The selection factors of include the current vehicle network channel state, the historical execution of task units on the VEC server, the computing power of the VEC server, and the reward function U(t); the optimal task offloading policy for each task unit in the time series is found by updating the value function. During this process, a convolutional neural network is used to construct the target network and the evaluation network respectively, and the difference between the value functions calculated by the two is used as the loss function. Then, the optimal task offloading scheme is found by the gradient descent method of the loss function.

2. The method according to claim 1, wherein In the second step, the task unit size d i determines the communication overhead when the task is offloaded, and the amount of computation c required for the task unit i determines the computation overhead of the task unit, and the task execution time limit indicates that the task execution result must be returned within the time; otherwise, the task execution fails. The task function ζ i is determined by the functional attributes of the task unit itself. Task units with the same ζ i are task units with the same function. The generation time For tasks generated by the vehicle, it is the generation moment of the task unit. For task units cached in the VEC server, this field is the moment when the task unit enters the server cache. The effective duration Δt of the calculation result i represents the valid time of the calculation result of this task unit.

3. The method according to claim 1, characterized in that, In the third step described above, the longer the remaining time delay of the task unit, the greater the utility, and when the calculation delay exceeds the maximum delay of the task unit there is an overtime penalty; the service cost includes the transmission cost of the task unit to the BS and the calculation cost in the VEC server. The vehicle communicates with the RSU using the DSRC (Dedicated Short-Range Communication) communication method, which is not charged.

4. The method according to claim 1, wherein In the third step, the utility function R i (t) is in the following form: Among them, φ(i) is the computing time of task unit i, and E i is the computing energy consumption of task unit i, Li is the normalized value of service cost, u(·) is the step function, α1, α2, α3 are weight coefficients, and α1 + α2 + α3 = 1, κ1 and κ2 are certain constants close to 1, and λ i is an arbitrary constant; the sum of the utility functions of each task unit in period t is defined as the reward function U(t) = ∑ i∈N R i (t). The policy space Π t of task offloading is defined as the set of all possible offloading strategies of task unit i in time period t. Then, the optimal task offloading decision in period t can be modeled as maximizing the reward function U(t) of task unit i, and its constraints include: ① The lower limit of the transmission power when the vehicle communicates with the RSU ; ② The lower limit of the transmission power when the vehicle communicates with the BS ; ③ Task unit i can only select one offloading method including the shared offloading method in period t; ④ If task unit i fails to match on the BS server, it cannot perform shared offloading on the BS server; ⑤ If task unit i fails to match on the RSU server, it cannot perform shared offloading on the RSU server.

5. The method according to claim 4, wherein In step 3, by maximizing the reward function, the optimal offloading strategy of the task unit at cycle t can be obtained The overall best offloading scheme P in the task offloading time process * It is expressed as: where \(0 < \gamma < 1\) is the discount factor representing the impact of future long-term utility, \(T\) is the set of time periods, and \(t\) is the time period serial number. is the mathematical expectation.

6. The method according to claim 1, wherein In the fourth step, at any time period \(t\in\{1,2,\ldots,T\}\), there are five offloading execution modes for task unit \(i\), and the offloading mode index is defined as indicating the offloading execution mode of task unit at period \(t\); indicating that task unit \(i\) is offloaded to the VEC server of the BS through V2B communication for calculation, indicating that task unit \(i\) performs shared offloading calculation on the VEC server of the BS, indicating that task unit \(i\) is offloaded to the VEC server of the RSU through V2R communication for calculation, indicating that task unit \(i\) performs shared offloading calculation on the VEC server of the RSU, indicating that task unit \(i\) executes locally; each task unit can only have one offloading calculation mode, that is, only one of them can take the value of 1.

7. The method according to claim 1, wherein In step 4, if the task unit ID does not match successfully, the offloading mode metric is initialized If the task unit ID matches successfully, the offloading mode metric is initialized Similarly, if the computational latency and transmission latency of the task unit offloaded to the VEC server of the RSU satisfy the task execution time limit That is At this time, if the task unit ID does not match successfully, the offloading mode metric is initialized If the task unit ID matches successfully, the offloading mode metric is initialized If the vehicle local computing time can satisfy the execution time limit of the task unit That is If the networks of the BS and the RSU are unavailable, or the offloading to the VEC server cannot meet the execution time limit, then the local computing mode will be the only option, and the offloading mode metric is initialized 8. The method according to claim 5, characterized in that, In the fifth step, define the state space S of the VEC task offloading t : Among them, and are the communication transmission rates from the vehicle to the BS and RSU respectively; f B , f R , f L are the CPU frequencies of the BS server, RSU server, and vehicle local computing unit respectively; is the calculation waiting time of task unit i, IDStack BS and IDStack RSU are the ID pools of the BS server and RSU server respectively. The ID pool caches the ID of the task unit that has been unloaded and executed on the server and its calculation result, providing support for the shared offloading of subsequent tasks; ID i represents the complete ID of task unit i; The task offloading process is equivalent to a Markov decision process, and the value function Q(S t , π t ) is defined as the long-term expected value of the task offloading reward function based on the state space S t and the task unit offloading policy π t : Find the optimal task offloading strategy for each task unit in the time series by updating the value function, and express the value function update process as the temporal difference method, in the following form: where β is the learning rate, Q * (S t , π t ) is the optimal value of the value function.

9. The method according to claim 8, wherein In the fifth step, a convolutional neural network is used to construct a target network and an evaluation network respectively; the target network calculates the value function of the offloading policy π t ; where, θ t is a parameter of the convolutional neural network; the evaluation network finds the optimal offloading strategy according to the current state through the value function Q(S t , π t , θ t ); construct a loss function Loss(θ t ) to measure the difference between the target network and the evaluation network: Update θ of the evaluation network using gradient descent according to the loss function t : Then, update θ according to the following formula: wherein is a scalar step size.

Citation Information

Patent Citations

  • Low-delay cooperative task processing method and device for Internet of Vehicles in mobile environment

    CN111538583A

  • Joint optimization method and system for task unloading and service caching of Internet of Vehicles

    CN114143346A