Multilateral task unloading optimization method based on staged deep reinforcement learning
By applying a multilateral task offload optimization method with staged deep reinforcement learning in intelligent vehicles, the problem of task processing delay and high energy consumption caused by limited computing resources of intelligent vehicles is solved, and more efficient task processing and better user experience is achieved.
Patent Information
- Application Number
- CN202411958358.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In smart vehicles, due to limited computing resources, it is difficult to efficiently handle multiple complex on-board tasks at the same time, resulting in high task processing delays and energy consumption, which affects road safety and traffic efficiency.
A multilateral task offload optimization method based on staged deep reinforcement learning is proposed. By establishing a task model, resource model and calculation model, combined with reinforcement learning algorithm (such as DQN), the unloading decision of subtasks is optimized to minimize the total completion time of vehicle-mounted tasks in the overall fine-grained vehicle section.
It significantly reduces the overall completion time of vehicle-mounted tasks in the overall fine-grained vehicle-mounted tasks, improves task processing efficiency, reduces the waste of computing resources, and improves the performance and user experience of vehicle-mounted systems.
Smart Images

Figure CN120050625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multilateral task offloading optimization method based on staged deep reinforcement learning, and belongs to the field of Internet of Vehicles and edge computing. Background Art
[0002] The Internet of Vehicles (IOV) is a key part of realizing modern intelligent transportation systems. Through vehicle-to-vehicle (V2V) and vehicle-to-roadside infrastructure (V2I) communications, it can effectively reduce task processing delays and task execution energy consumption, while effectively improving road safety and traffic efficiency. With the rapid development of IOV technology, the intelligence level of modern intelligent vehicles has been significantly improved. New intelligent vehicles are equipped with a large number of advanced intelligent devices, such as higher-resolution cameras, various complex sensors, and other human-computer interaction devices. With the increase in the number of these intelligent devices, more and more new intelligent vehicle applications have gradually emerged, including intelligent navigation, unmanned driving, intelligent voice assistants, and various multimedia applications, which greatly enrich the functions and user experience of the vehicle system. These vehicle applications will generate various types of vehicle tasks. When intelligent vehicles process these complex tasks, they need huge computing resources to support them. However, the computing resources of the vehicle itself are very limited, and it is difficult to simultaneously meet the processing of all tasks with high computing resource requirements.
[0003] The emergence of Mobile Edge Computing (MEC) provides a highly potential solution to the problem of vehicle-borne task offloading. MEC deploys edge servers in roadside infrastructure close to vehicles, allowing smart vehicles to offload their own tasks to edge servers for processing, and then quickly return the calculation results to the vehicles. This mechanism greatly alleviates the dilemma of smart vehicles being unable to efficiently handle complex vehicle-borne tasks due to their limited computing power. Compared with traditional cloud computing architecture, the computing resources of the MEC architecture are closer to the data generation source, thereby overcoming the large transmission delay problem caused by the long physical distance between the remote cloud server and the smart vehicle. In addition, with the rapid development of smart vehicles, idle smart vehicles on the road section can also serve as one of the vehicle-borne task offloading nodes with their sufficient computing power, receiving computing tasks from other smart vehicles carrying tasks, so as to alleviate the problem of excessive load on the roadside edge server caused by road congestion, and improve the overall vehicle-borne task completion efficiency of the road section. To this end, this patent proposes a task offloading optimization method with the goal of minimizing the total completion time consumption of fine-grained vehicle tasks as a whole on the road section. Summary of the invention
[0004] The problem to be solved by the present invention is: for a designated road section, the road section includes several task vehicles carrying on-board tasks, several idle vehicles not carrying on-board tasks and several edge servers covering the entire road section. One task vehicle carries one on-board task, and one on-board task consists of a group of subtasks with data dependencies. Specifically, before a subtask starts execution, it often needs to wait for other related subtasks to be completed and provide processing results. In the same group of tasks, some subtasks can be processed in parallel, while other subtasks can only be processed strictly serially because they depend on the results of other tasks. Roadside edge servers and idle vehicles can serve as unloading nodes for on-board tasks to help task vehicles perform task calculations. In this scenario, a task unloading scheme is determined to minimize the total completion time of the overall fine-grained on-board tasks of the road section. The specific solution is as follows:
[0005] 1. Establish a task model
[0006] In a fixed period of time, each task vehicle on the road section will carry a vehicle-borne task. 1 , t 2 , t 3 , ..., t, n} to represent the vehicle-borne task set on the road section, n represents the number of vehicle-borne tasks on this road section during this period, and the present invention uses a binary t i =(sn i , d i ) to represent the vehicle-mounted task, where sn i Represents task t i The number of subtasks, d i Represents task t i Each vehicle-mounted task consists of a group of subtasks with data dependencies. Before a subsequent subtask can start executing, it must receive the execution result data of all previous subtasks. A directed acyclic graph (DAG) is established based on the data dependencies between subtasks. The vertices in the graph represent a specific subtask, and the edges represent the dependencies between subtasks. To represent the task t i The present invention uses an octet st i,j =(b i,j , in i,j , c i,j , o i,j ,pre i,j , suc i,j , rk i,j ,ly i,j ) to represent the subtask, where b i,j Indicates subtask st i,j The tasks in i,jIndicates subtask st i,j The amount of input data, c i,j Indicates subtask st i,j The computing resources required, o i,j Indicates subtask st i,j The amount of output data after processing is completed, pre i,j Indicates subtask st i,j The preceding subtask set, suc i,j Indicates subtask st i,j The subsequent subtask set, rk i,j Indicates subtask st i,j The priority value, ly i,j Indicates subtask st i,j The number of levels.
[0007] 2. Establish resource model
[0008] On a fixed road section, there is a road side unit (RSU), and one RSU has multiple edge servers, which can be used as unloading nodes. In addition, there will be some idle vehicles on the road section that do not carry on-board tasks. Their computing resources are also sufficient and can also be used as unloading nodes. Due to the distance limitation of the remote cloud, they are not considered as unloading nodes for the time being. Therefore, the computing resources of the present invention include task vehicle resources, edge server resources and idle vehicle resources. The present invention uses V = {v 1 , v 2 , v 3 ,...,v n} to represent the task vehicle set on the road section, where n represents the number of task vehicles on the road section. The present invention uses a triple v i =(vcf i ,vtp i ,veit i ) to represent the mission vehicle resources, where vcf i Represents the mission vehicle v i The calculation frequency of vtp i Represents the mission vehicle v i The transmission power, veit i Represents the mission vehicle v i The earliest idle time of 1 ,s 2 ,s 3 , ..., s m} to represent the edge server set on the road segment, where m represents the number of edge servers on the road segment. The present invention uses a triplet s j =(scf j , stp j , seitj ) to represent edge server resources, where scf j Represents edge servers j The calculation frequency of stp j Represents edge servers j The transmission power, j Represents edge servers j The earliest idle time of IL. 1 ,il 2 ,il 3 ,...,il p} to represent the idle vehicle set on the road section, where p represents the number of idle vehicles on the road section. The present invention uses a triple il k =(ilcf k ,iltp k ,ileit k ) to represent idle vehicle resources, where ilcf k Indicates idle vehicle il k The calculation frequency of iltp k Indicates idle vehicle il k The transmission power, ileit k Indicates idle vehicle il k Earliest free time.
[0009] 3. Build a computational model
[0010] Vehicles on the road communicate with each other and with RSUs through the 5G cellular network. The transmission rate between different devices is expressed by formula (1):
[0011]
[0012] In formula (1), start represents the transmitting device, start∈{V, S, IL}, end represents the receiving device, dest∈{V, S, IL}, b represents the channel bandwidth between the communicating parties, tp represents the transmission power of the transmitting device, h represents the channel gain between the communicating parties, and n represents the noise power between the communicating parties.
[0013] There are three options for subtasks before calculation: calculate locally, offload to edge servers, and offload to idle vehicles.
[0014] When the subtask is chosen to be calculated locally, its calculation time consumption is expressed by formula (2):
[0015]
[0016] When the subtask is offloaded to the edge server for calculation, its transmission time and calculation time are expressed by formula (3) and formula (4) respectively:
[0017]
[0018] When the subtask is offloaded to an idle vehicle for calculation, its transmission time and calculation time are expressed by formula (5) and formula (6) respectively:
[0019]
[0020] The completion time of the non-end subtask is expressed by formula (7)-formula (10):
[0021] tft i,j =tst i,j +ct i,j (7)
[0022] tst i,j =max(tt i,j ,tpt i,j , eit dest(i,j) ) (8)
[0023]
[0024] In formula (7), tft i,j Indicates non-end subtask st i,j The completion time, tst i,j Indicates non-end subtask st i,j The execution time of the start. In formula (8), tpt i,j Indicates non-end subtask st i,j The latest arrival time of the calculation result of the preceding subtask, eit dest(i,j) Indicates non-end subtask st i,j The earliest idle time of the execution location. The subtask can start execution only when the current subtask arrives at the execution location, the calculation results of all previous subtasks of the current subtask arrive at the execution location of the current subtask, and the execution location of the current subtask is idle. In formula (9), Indicates subtask st i,k The calculation results from the subtask st i,k The execution location is transferred to the subtask st i,j The length of time it takes for data to be transferred to the execution location.
[0025] The final completion time of a vehicle-borne task is expressed by formula (11):
[0026]
[0027] In formula (11), the final completion time of the vehicle-mounted task is the completion time of the final subtask of this group of subtasks plus the time when the result of the final subtask is transmitted back to the local.
[0028] 4. Define the optimization problem
[0029] The present invention uses a one-dimensional vector To represent the subtask st i,j Uninstall decision, use dest i,j To represent the subtask st i,j The uninstall destination, Indicates subtask st i,j Calculate locally, then dest i,j =0, Indicates subtask st i,j Offload to the edge server for calculation, then dest i,j ∈{1, 2, 3, ..., m}, Indicates subtask st i,j Unload to an idle vehicle for calculation, then dest i,j ∈{m+1, m+2, m+3, ..., m+q}. A subtask can only choose one execution location for calculation.
[0030] In the present invention, the total completion time of the entire vehicle-mounted task of the road section is used to represent the final optimization result. The smaller the time consumption, the better the unloading scheme. The objective function is expressed by formula (12):
[0031]
[0032] In addition, the constraints are: C1: C2: tft i <d i ;
[0033] Based on the above description, data preprocessing is performed on tasks with data dependencies across the entire system. The specific solution is as follows:
[0034] Since the execution of tasks with data dependencies needs to follow strict sequential rules, the starting subtask of all vehicle-mounted tasks of the system is first taken as the first level, and then the subsequent subtasks are hierarchically divided in turn. Except for the starting task, each subtask has several preceding subtasks. For each subtask, the maximum level of its preceding subtask is first calculated, and the level of this subtask is the maximum level of its preceding subtask plus one, and finally the task sequence of each level can be obtained. For each level of task sequence, it is prioritized according to certain rules. The specific rules are as follows: First, the computing frequency of an edge server on the road section is obtained and used as the basic frequency. For each subtask, its estimated computing time is calculated based on its required computing resources and basic frequency, and its estimated transmission rate is calculated based on the transmission power, channel bandwidth, channel gain and noise power of the vehicle to which it belongs. Its estimated transmission time is calculated based on its input data volume and estimated transmission rate. Its estimated computing time is subtracted from the estimated transmission time to obtain its task priority value. The specific algorithm design is as follows:
[0035]
[0036] The task offloading process is modeled as a Markov decision process. Reinforcement learning includes three elements: state, action, and reward. For the model to be solved in this patent, the specific design of the three elements of reinforcement learning is as follows:
[0037] 1) State space: The state is used to describe the real-time state in the current environment. The state space includes three aspects: task state, resource state, and communication state. The task state is expressed as T = {in, c, o, pt}, where in represents the amount of task input data, c represents the computing resources required for the task, o represents the amount of task output data, and pt represents the maximum completion time of the previous subtask of the task. The resource state is expressed as R = {vcf, scf, ilcf, veit, seit, ileit}, where vcf represents the calculation frequency of the task vehicle, scf represents the calculation frequency of all edge servers, ilcf represents the calculation frequency of all idle vehicles, veit represents the earliest idle time of the task vehicle, seit represents the earliest idle time of all edge servers, and ileit represents the earliest idle time of all idle vehicles. The communication state is expressed as E = {b, h, n}, where b represents the channel bandwidth, h represents the channel gain, and n represents the noise power. The overall state space is State = {T, R, E}.
[0038] 2) Action space: Action is used to represent the offloading decision of the current task. The task can be calculated locally, offloaded to the edge server, or offloaded to an idle vehicle. The action space is 1 plus the number of edge servers plus the number of idle vehicles, so the action space is Action = {0, 1, 2, ..., m, ..., m+q}.
[0039] 3) Reward function: After the system agent performs an action in each state, there should be a reward value to measure the quality of the decision made in the current step. According to our optimization goal, we need to minimize the time consumption for task completion, so the reward function is defined as the negative value of the time consumption for task completion in the current step.
[0040] The most traditional reinforcement learning algorithm is the Q-learning algorithm, which uses table storage to represent Q values and performs well in scenarios where all states can be clearly listed. However, the model of this patent has a large-scale state space, so DQN is used to solve the optimization problem. DQN uses a deep neural network to approximate the Q function. On the basis of Q-learning, a neural network is used to replace the Q table. On this basis, in order to avoid falling into the local optimal solution, DQN will use a greedy strategy for training. In the early stages of training, the agent decision will be more inclined to use a random strategy. As the number of training increases, the agent decision gradually tends to use existing experience to make decisions. In addition, DQN has an experience playback mechanism. After each decision, it will represent the current state, current action, current step reward, and next state as a four-tuple (s t , a t , r t ,s i+1 ) is stored in the experience pool. During training, DQN will randomly sample data from the playback pool for learning to improve the stability of the algorithm. DQN uses the target network to calculate the target Q value of each sample, and the calculation formula is: t =r t +γmax a′ Q′(s t+1 ,a′;θ - ) After each fixed number of iterations, the parameters of the Q network are copied to the target network. DQN usually uses mean square error as the loss function, which is calculated as: The specific algorithm design is as follows:
[0041]
[0042] After learning and training with a large amount of data, the Q network will have a strong generalization ability and can effectively adapt to different highly dynamic and complex task offloading scenarios. When a new task arrives, the agent can use the current system environment state to derive a better offloading strategy in real time. When using the trained Q network to solve the optimization problem, first perform data preprocessing to obtain a subtask dictionary after hierarchical segmentation and sorting, and then use the Q network to make decisions on the subtask sequence of each level in turn, and finally come to a better offloading decision. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1Flowchart of the optimization method for multi-task offloading based on staged deep reinforcement learning
[0044] Figure 2 System architecture diagram for task offloading
[0045] Figure 3 Comparison of the effects of different algorithms under different numbers of mission vehicles DETAILED DESCRIPTION
[0046] In order to more clearly illustrate the content of the present invention, an example is used for specific description below.
[0047] The scenario of this experiment is set in a highway environment. The system architecture is as follows: Figure 2 As shown in the figure, an RSU is set up in the center of a two-way four-lane road section. There are 10 edge servers under the RSU, and the coverage range of the RSU is 500m. There are 70 task vehicles and 10 idle vehicles on the road section. Each task vehicle carries a vehicle task, and each vehicle task consists of 5-10 subtasks. The input data volume of each subtask is randomly selected between 100KB and 200KB, and the required computing resources are 5*10 7 To 1*10 8 In real scenarios, the amount of output data is usually very small and its transmission time consumption is almost negligible. Therefore, for the convenience of calculation, the output data volume is set to 0 in this example. The computing frequency of the task vehicle is randomly selected between 0.5GHz and 1GHz, and the transmission power is fixed at 1W. The computing frequency of the edge server is fixed at 4GHz, and the transmission power is fixed at 4W. The computing frequency of the idle vehicle is randomly selected between 2GHz and 2.5GHz, and the transmission power is fixed at 2W. For specific information, see Table 1. The channel bandwidth on the road section is fixed at 20MHz, the channel gain is fixed at 1, and the noise power is fixed at 1*10 - 9 W.
[0048] Since the experimental task vehicles and task data space are relatively large, the task set data carried by the first five task vehicles are used to illustrate the hierarchical segmentation and priority value algorithm. The specific information is shown in Table 2.
[0049] Table 1 Idle vehicle set information
[0050] Idle Vehicle ID Calculation frequency (Hz) Transmission power(W) 1 2429679107 2 2 2186048955 2 3 2080231071 2 4 2323114117 2 5 2284330735 2 6 2108151281 2 7 2168045714 2 8 2474650528 2 9 2383486085 2 10 2387107690 2
[0051] Table 2 Task set information carried by the first five task vehicles
[0052]
[0053]
[0054] First, the tasks are divided into levels. 1 , t 7 , t 15 , t 21 , t 29 The number of levels is set to 1, and then the number of levels of subsequent subtasks is calculated in topological order, such as t 8 The only previous subtask is t 7 , t 7 The number of levels is t 8 The maximum number of levels of the preceding subtask, so t 8 The number of levels is t 7 The number of levels is increased by 1, that is, the number of levels is 2, t 9 The preceding subtasks of 7 and t 8 , we need to find the maximum number of levels of the preceding subtasks, t 8 The number of levels is 1, which is the largest number of levels among the previous subtasks, so t 9 The number of levels is t 8 The number of levels is 1 plus 1, that is, the number of levels is 2, and so on, the number of levels of each subtask can be obtained. Then, according to the different levels of subtasks, arrays are created for subtasks with the same number of levels. Finally, 7 arrays can be obtained, namely [t 1 ,t 7 , t 15 , t 21 , t 29 ],[t 2 , t 4 , t 8 , t 10 , t 16 , t 22 , t 30 , t 31 ],[t 3 , t 9 , t 11 , t 17 , t 23 , t 32 ],[t 5 , t 12 , t 13 , t 18 , t 24 , t 26 , t 33 , t 36 ],[t 6 , t 14 , t 19 , t 25 , t 34 ],[t 20 , t 27 , t 35],[t 28 , t 37 ]. Create a dictionary with the key as the number of levels and the value as the array corresponding to the number of levels. Then loop through each subtask in each array in the dictionary, calculate the priority value of each subtask according to the formula, and finally sort each array in the dictionary according to the priority value. Finally, a complete subtask dictionary after data preprocessing is obtained. Then use the trained Q network to solve each layer of subtask sequence in turn, and finally get a better task offloading decision.
[0055] In this example, three algorithms were used as comparison algorithms to reflect the superiority of the algorithm proposed in the present invention. Using the local algorithm, that is, all tasks are chosen to be calculated locally, the total completion time obtained is 52.95s. Due to limited local computing resources, the local algorithm has the situation that the task cannot be completed within the deadline. The total completion time obtained by using the genetic algorithm is 33.64s. Using the greedy algorithm, that is, selecting the unloading node with the lowest current processing time for each subtask in sequence, the total completion time obtained is 29.85s. The total completion time obtained by using the multilateral task unloading optimization method based on staged deep reinforcement learning proposed in the present invention is 28.40s. The experimental results show that the optimization effect of the method proposed in the present invention is the best. In addition, in order to reflect the optimization ability of the method proposed in the present invention under different situations, the comparison is further made under different numbers of task vehicles (such as Figure 3 As shown in the figure, curve 1 represents the local algorithm, curve 2 represents the genetic algorithm, curve 3 represents the greedy algorithm, and curve 4 represents the algorithm of the present invention), the optimization effects under different idle vehicle numbers, the optimization effect of the method proposed by the present invention is the best. The method proposed by the present invention significantly reduces the total completion time of the fine-grained vehicle tasks of the entire road section, improves the task processing efficiency, and brings a better vehicle experience to users.
Claims
1. A multi-task offloading optimization method based on staged deep reinforcement learning, characterized in that A mathematical model is established for the combinatorial optimization problem in a specific environment. With the goal of reducing the total completion time of the overall fine-grained vehicle tasks on a specified road section as much as possible, an efficient task offloading optimization method is proposed to obtain a task offloading decision that can achieve a better target result. The method includes three steps: 1) Establishing the task model, resource model and calculation model in the Internet of Vehicles environment and defining the optimization problem; 2) All fine-grained on-board tasks on the road section are hierarchically segmented and their priority values are calculated; 3) The task offloading problem is modeled as a Markov decision process and solved using a deep Q-network.
2. The method for optimizing multilateral task offloading based on staged deep reinforcement learning according to claim 1, characterized in that: In step 1), a task model in the Internet of Vehicles environment is established, using a binary t i =(sn i ,d i ) to represent the vehicle-mounted task, where sn i Represents task t i The number of subtasks, d i Represents task t i The deadline. Use an octet st i,j =(b i,j ,in i,j , c i,j ,o i,j ,pre i,j ,suc i,j , rk i,j ,ly i,j ) to represent the subtask, where b i,j Indicates subtask st i,j The tasks in i,j Indicates subtask st i,j The amount of input data, c i,j Indicates subtask st i,j The computing resources required, o i,j Indicates subtask st i,j The amount of output data after processing is completed, pre i,j Indicates subtask st i,j The preceding subtask set, suc i,j Indicates subtask st i,j The subsequent subtask set, rk i,j Indicates subtask st i,j The priority value, ly i,j Indicates subtask st i,j The resource model in the Internet of Vehicles environment is established, using a triplet v i =(vcf i ,vtp i ,veit i ) to represent the mission vehicle resources, where vcf i Represents the mission vehicle v i The calculation frequency of vtp i Represents the mission vehicle v i The transmission power, veit i Represents the mission vehicle v i The earliest idle time. Use a triple s j =(scf j ,stp j ,seit j ) to represent edge server resources, where scf j Represents edge servers j The calculation frequency of stp j Represents edge servers j The transmission power, j Represents edge servers j The earliest idle time. Use a triple il k =(ilcf k ,iltp k ,ileit k ) to represent idle vehicle resources, where ilcf k Indicates idle vehicle il k The calculation frequency of iltp k Indicates idle vehicle il k The transmission power, ileit k Indicates idle vehicle il k The earliest idle time. A calculation model in the Internet of Vehicles environment was established. Vehicles communicate with each other and with RSUs through 5G cellular networks. The transmission rate between different devices is: start represents the sending device, start∈{V,S,IL}, end represents the receiving device, dest∈{V,S,IL}, b represents the channel bandwidth between the communicating parties, tp represents the transmission power of the sending device, h represents the channel gain between the communicating parties, and n represents the noise power between the communicating parties. When the subtask is chosen to be calculated locally, its calculation time is: When the subtask is offloaded to the edge server for calculation, its transmission time and calculation time are: When the subtask is offloaded to an idle vehicle for calculation, its transmission time and calculation time are: The calculation formula for the completion time of non-end subtasks is: tft i,j =tst i,j +ct i,j tst i,j =max(tt i,j ,tpt i,j ,it dest(i,j) ) tft i,j Indicates non-end subtask st i,j The completion time, tst i,j Indicates non-end subtask st i,j The time when execution starts. i,j Indicates non-end subtask st i,j The latest arrival time of the calculation result of the preceding subtask, eit dest(i,j) Indicates non-end subtask st i,j The earliest idle time of the execution location. The subtask can start execution only when the current subtask arrives at the execution location, the calculation results of all the previous subtasks of the current subtask arrive at the execution location of the current subtask, and the execution location of the current subtask is idle. Indicates subtask st i,k The calculation results from the subtask st i,k The execution location is transferred to the subtask st i,j The length of time it takes for data to be transferred to the execution location. The final completion time of a vehicle task is: Define the optimization problem as a one-dimensional vector To represent the subtask st i,j Uninstall decision, use dest i,j To represent the subtask st i,j The uninstall destination, Indicates subtask st i,j Calculate locally, then dest i,j =0, Indicates subtask st i,j Offload to the edge server for calculation, then dest i,j ∈{1, 2, 3, ..., m}, Indicates subtask st i,j Unload to an idle vehicle for calculation, then dest i,j ∈{m+1, m+2, m+3, ..., m+q}. A subtask can only choose one execution location for calculation. The total completion time of the entire vehicle-mounted task on the road section is used to represent the final optimization result. The smaller the time consumption, the better the unloading solution. The objective function is: In addition, the constraints are: C1: C2: tft i <d i ; 3. The method for optimizing multilateral task offloading based on staged deep reinforcement learning according to claim 1, characterized in that: In step 2), all vehicle-borne tasks on the road section are hierarchically segmented and their priority values are calculated. Specifically: First, the starting subtask of all vehicle-mounted tasks of the system is taken as the first level, and then the subsequent subtasks are hierarchically divided in turn. Except for the starting task, each subtask has several preceding subtasks. For each subtask, the maximum level of its preceding subtask is first calculated, and the level of this subtask is the maximum level of its preceding subtask plus one, and finally the task sequence of each level can be obtained. For each level of task sequence, it is prioritized according to certain rules. The specific rules are as follows: First, the computing frequency of an edge server on the road section is obtained and used as the basic frequency. For each subtask, its estimated computing time is calculated according to its required computing resources and basic frequency, and its estimated transmission rate is calculated according to the transmission power, channel bandwidth, channel gain and noise power of the vehicle to which it belongs. Its estimated transmission time is calculated according to its input data volume and estimated transmission rate. Its estimated computing time is subtracted from the estimated transmission time to obtain its task priority value.
4. The method for optimizing multilateral task offloading based on staged deep reinforcement learning according to claim 1, characterized in that: In step 3), the task offloading problem is modeled as a Markov decision process and solved using a deep Q network. Specifically: Design three elements of reinforcement learning: 1) State space: The state is used to describe the real-time state in the current environment. The state space includes three aspects: task state, resource state, and communication state. The task state is expressed as T = {in, c, o, pt}, where in represents the amount of task input data, c represents the computing resources required for the task, o represents the amount of task output data, and pt represents the maximum completion time of the previous subtask of the task. The resource state is expressed as R = {vcf, scf, ilcf, veit, seit, ileit}, where vcf represents the calculation frequency of the task vehicle, scf represents the calculation frequency of all edge servers, ilcf represents the calculation frequency of all idle vehicles, veit represents the earliest idle time of the task vehicle, seit represents the earliest idle time of all edge servers, and ileit represents the earliest idle time of all idle vehicles. The communication state is expressed as E = {b, h, n}, where b represents the channel bandwidth, h represents the channel gain, and n represents the noise power. The overall state space is State = {T, R, E}. 2) Action space: Action is used to represent the offloading decision of the current task. The task can be calculated locally, offloaded to the edge server, or offloaded to an idle vehicle. The action space is 1 plus the number of edge servers plus the number of idle vehicles, so the action space is Action = {0, 1, 2, ..., m, ..., m+q}. 3) Reward function: After the system agent performs an action in each state, there should be a reward value to measure the quality of the decision made in the current step. According to our optimization goal, we need to make the task completion time as low as possible, so the reward function is defined as the negative value of the task completion time of the current step. A large amount of data is used for learning and training to obtain a Q network with excellent generalization ability. The Q network is used to make decisions on the sub-task sequence of each level in turn, and finally a better offloading decision is obtained.
Citation Information
Patent Citations
Calculation unloading method considering task priority for mobile edge computing network
CN114302456A
Internet of vehicles time delay sensitive application unloading method based on federated learning
CN115767634A
Task scheduling and resource allocation method based on federal reinforcement learning in Internet of Vehicles
CN116709378A
Dependent task unloading method in V2V scene, terminal and storage medium
CN117939535A
Road-section-based task unloading global optimization distribution method in Internet of Vehicles environment
CN118337824A