RSU auxiliary multi-hop task unloading method based on asynchronous deep reinforcement learning

By adopting the RSU-assisted multi-hop task unloading method with asynchronous deep reinforcement learning in the Internet of Vehicles environment, and using the A3C algorithm to optimize the task unloading strategy, it solves the problem that traditional single-hop strategy is difficult to cope with the dynamic network environment, and realizes efficient task processing and resource utilization.

CN119996446APending Publication Date: 2025-05-13WUXI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510091288.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the Internet of Vehicles environment, traditional single-hop task offloading strategies are difficult to meet the dynamically changing network environment and complex computing tasks requirements, resulting in high task processing delays and uneven resource utilization.

Method used

The RSU-assisted multi-hop task unloading method based on asynchronous deep reinforcement learning is adopted. The task unloading strategy is optimized under the framework of Markov decision-making process through the A3C algorithm, and the optimal multi-hop communication path is dynamically selected to maximize task processing efficiency and minimize transmission delay.

Benefits of technology

It realizes effective response to network topology changes in dynamic multi-hop task offload scenarios, improves the accuracy and efficiency of task offload decisions, improves the system's resource utilization rate and task completion rate, and adapts to different network conditions and user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996446A_ABST
    Figure CN119996446A_ABST
Patent Text Reader

Abstract

The invention discloses an RSU auxiliary multi-hop task unloading method based on asynchronous deep reinforcement learning, relates to the technical field of edge unloading and deep reinforcement, and accurately evaluates the connection stability between nodes by establishing a vehicle movement model and a communication model; modeling a task unloading problem into a high-dimensional MDP, and solving the problem by adopting an A3C algorithm; according to the A3C algorithm, a plurality of agents are asynchronously trained, so that the optimization problem in a large-scale action space is effectively solved, and a better task unloading strategy is realized in a complex and changeable car networking environment; compared with an existing task unloading strategy, the method has the advantages that the performance is remarkably improved in the aspects of reducing task processing delay and improving the computing resource utilization rate, and particularly in a multi-RSU and high-dynamic scene, the task completion rate and the resource utilization efficiency of the system are remarkably improved; in addition, the method also has good expansibility, and can adapt to more complex application requirements in the future Internet of Vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of edge offloading and deep reinforcement technology, and in particular to an RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning. Background Art

[0002] In the modern Internet of Vehicles (IoV), with the continuous development of autonomous driving technology and the widespread application of intelligent transportation systems, the data exchange between vehicles and infrastructure in the network has shown an exponential growth. This surge in data traffic has brought tremendous pressure to the existing communication network, especially during peak traffic hours and dense urban environments. The traditional single-hop communication method is difficult to meet the needs of real-time task processing, resulting in higher task processing delays and uneven resource allocation, which in turn affects the overall network performance and user experience.

[0003] To meet this challenge, multi-hop task offloading technology has gradually become an important means to optimize the utilization of computing resources in the Internet of Vehicles. The core idea of ​​multi-hop task offloading is to introduce multiple forwarding nodes in the network, through which tasks are transferred from the generation location to other nodes with computing capabilities (such as cloud servers, roadside units RSUs or other vehicles) to fully utilize the computing resources in the network. However, due to the high-speed mobility of vehicles and the dynamic changes in network topology, multi-hop task offloading faces problems such as unstable communication connections and complex path selection. These problems make it a huge challenge in practical applications to select the optimal task offloading path and ensure low-latency processing of tasks.

[0004] Traditional task offloading strategies usually rely on single-hop communication, that is, tasks are directly transmitted from roadside units (RSUs) to vehicles within coverage or uploaded to cloud servers for processing. This strategy may work well when network nodes are sparse or the communication environment is stable, but in environments with dense and highly dynamic nodes, due to limited coverage and uneven resource utilization, some node resources are often overloaded while other node resources are idle, seriously affecting the efficiency of the overall network. To overcome these shortcomings, researchers have proposed a multi-hop task offloading strategy, hoping to achieve better task allocation and resource utilization by making full use of all available network resources during task forwarding.

[0005] To this end, in recent years, deep reinforcement learning (DRL) technology has gradually been applied to multi-hop task offloading problems to cope with complex network environments and high-dimensional decision spaces. In particular, the task offloading strategy based on the asynchronous advantage actor-critic (A3C) algorithm performs well in dealing with large-scale action spaces and dynamic environments. The A3C algorithm can converge to the optimal strategy faster through asynchronous training of multiple agents, and shows significant advantages in the allocation of computing resources and dynamic offloading of tasks.

[0006] In the specific implementation, the A3C algorithm first establishes a vehicle mobility model and a communication model to accurately evaluate the connection stability between nodes to ensure the reliability of data transmission and the rationality of path selection during multi-hop task offloading. Then, the task offloading problem is modeled as a Markov decision process (MDP), and the A3C algorithm is used to train and optimize the decision process to maximize the processing efficiency of the task and minimize the transmission delay. The A3C algorithm dynamically adjusts the task offloading strategy by exploring multiple task paths in parallel, so that tasks can be quickly and efficiently assigned to the most suitable computing nodes in the complex and changing Internet of Vehicles environment.

[0007] However, although the application of the A3C algorithm in multi-hop task offloading has shown great potential, its computational overhead and training complexity in practical deployment are still issues that need further study. In addition, with the continuous development of the Internet of Vehicles, the number and heterogeneity of nodes in the network are also increasing, which puts higher requirements on the multi-hop task offloading strategy. Therefore, future research directions will include further optimizing the efficiency of the A3C algorithm, reducing its computational cost, and exploring the possibility of combining it with other advanced technologies (such as edge computing, 6G networks) to achieve a smarter and more efficient Internet of Vehicles task offloading system. Summary of the invention

[0008] In order to solve the above technical problems, the present invention provides an RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning, comprising the following steps:

[0009] S1. Establish an edge unloading system, which includes a cloud server, several roadside units and several vehicles;

[0010] S2. In each time slot, the roadside unit senses the current road environment and generates a computing task;

[0011] S3: The roadside unit finds candidate vehicles that can receive computing tasks by broadcasting service requests. The candidate vehicles feed back their status information to the roadside unit, including the task queue and computing resource availability.

[0012] S4, the roadside unit calculates the transmission path of each candidate node based on the received information and selects the optimal transmission path. The candidate nodes include the vehicle, other roadside units and cloud servers; then the total processing delay after the task is offloaded to the vehicle, roadside unit and cloud server is calculated;

[0013] S5. Through the A3C algorithm, the task offloading strategy is optimized under the Markov decision process framework to maximize the negative reward of task processing. The optimal offloading strategy is obtained by continuously interacting with the environment to update the parameters and minimize the total processing delay.

[0014] The beneficial effects of the present invention are:

[0015] (1) The present invention can cope with dynamic multi-hop task offloading scenarios. Traditional task offloading algorithms usually assume that the communication path and network conditions are static, while the present invention adopts the asynchronous deep reinforcement learning (A3C) algorithm, which can effectively cope with the frequent changes of network topology in a dynamically changing Internet of Vehicles environment and has more practical application value.

[0016] (2) In the present invention, the A3C algorithm is an efficient algorithm in current deep reinforcement learning. Through the parallel training of multiple agents, the present invention realizes the rapid search and optimization of large-scale action space, improves the accuracy and efficiency of task offloading decisions, and improves the resource utilization of the system, and also ensures the training stability and global optimality of the results in complex scenarios.

[0017] (3) In the present invention, adaptive selection of multi-hop communication paths is achieved. By constructing a vehicle mobility model and a communication model and dynamically evaluating the connection stability between nodes, the optimal multi-hop communication path can be adaptively selected, which reduces data transmission delay and improves the overall efficiency of task processing.

[0018] (4) In the present invention, the Actor-Critic structure of the A3C algorithm is used to achieve efficient task offloading strategy training. The synergy between the strategy network and the value network ensures the stable update of the strategy, significantly improving the success rate of task offloading and the overall performance of the system.

[0019] (5) In the present invention, resource utilization and task completion rate are significantly improved. In multi-wayside units and high dynamic scenarios, by optimizing the task offloading strategy, the utilization of system resources and the task completion rate are significantly improved. Compared with the traditional task offloading strategy, the present invention can better cope with the changing network environment and ensure efficient task execution;

[0020] (6) The present invention has high flexibility and scalability, and can adapt to different network conditions and user needs by dynamically adjusting the unloading strategy; at the same time, the system can work in conjunction with cloud computing, edge computing and other technologies to further improve the overall system performance, providing solid technical support for the future development of the Internet of Vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0022] Figure 2 A schematic diagram of a model of an edge unloading system in an embodiment of the present invention;

[0023] Figure 3 is a block diagram based on the A3C algorithm in an embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the RSU assisted unloading process algorithm in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] This embodiment provides an RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning, such as Figure 1 As shown, the following steps are included:

[0026] S1. Establish an edge unloading system, which includes a cloud server, several road side units (RSUs) and several vehicles;

[0027] S2. In each time slot, the roadside unit senses the current road environment and generates a computing task;

[0028] S3: The roadside unit finds candidate vehicles that can receive computing tasks by broadcasting service requests. The candidate vehicles feed back their status information to the roadside unit, including the task queue and computing resource availability.

[0029] S4, the roadside unit calculates the transmission path of each candidate node based on the received information and selects the optimal transmission path. The candidate nodes include the vehicle, other roadside units and cloud servers; then the total processing delay after the task is offloaded to the vehicle, roadside unit and cloud server is calculated;

[0030] S5. Through the Asynchronous Advantage Actor-Critic (A3C) algorithm, the task offloading strategy is optimized under the Markov decision process framework to maximize the negative reward of task processing. The optimal offloading strategy is obtained by continuously interacting with the environment to update parameters and minimize the total processing delay.

[0031] like Figure 2 As shown, the method of this embodiment considers a road environment consisting of a cloud server, multiple roadside units (RSUs) and multiple vehicles. The cloud server, roadside units and vehicles are regarded as three different types of computing nodes. The edge unloading system operates in a time slot mode. The time is divided into multiple time slots. The vehicle may move to different locations in each time slot.

[0032] The set R = {1, 2, ..., r, ..., R} is used to represent the uniformly distributed roadside units, and the set V = {1, 2, ..., v, ..., V} is used to represent the vehicles. Considering that the edge unloading system operates in a time slot-by-time slot manner, the time is divided into multiple time slots, which are represented by the set {1, 2, ..., τ, ...}. At time slot τ, the vehicles in the coverage area of ​​the roadside unit r are represented by the set Represents; in this environment, vehicles i and j that cannot communicate directly establish communication through vehicle v, and vehicle v is defined as a forwarding vehicle; adjacent roadside units can communicate with each other or directly with the cloud server; in addition, roadside units can not only communicate with vehicles within their coverage, but also communicate with vehicles outside their coverage through forwarding vehicles.

[0033] At time slot τ, the roadside unit r senses the road environment and generates a computation task (s r , c r , t r ), where s r Indicates the task size, c r represents the computing resources required to process the task, t r Indicates the execution time limit of the task; each task generated by the roadside unit can be offloaded to the cloud server, roadside unit or vehicle for execution, using x r,n represents an uninstall decision, where and n∈{1,...,V,V+1,...,V+R,V+R+1}.

[0034] Due to the movement of vehicles, the risk of communication interruption is high, which brings risks to data transmission; the connectivity between nodes can be represented by the link connection time, which refers to the duration for which two nodes maintain communication; during the link connection time, the two nodes can continue to communicate, otherwise the communication will be interrupted.

[0035] For simplicity, this embodiment assumes that the vehicle is moving at a constant speed only along the x-axis, which can be easily extended to a three-dimensional area, where the speed in the right direction is positive and the speed in the left direction is negative; for simplicity, the communication range between vehicles is represented by a constant denoted by , and the communication range between the vehicle and the roadside unit is denoted by ψ.

[0036] For any adjacent vehicles i and j, the position coordinates are expressed as (x i ,y i ) and (x j ,y j ), the speed is represented by v i and v j , if the speed is greater than 0, it means that the vehicle is moving in the correct direction, otherwise, it means that the vehicle is moving in the opposite direction of the correct direction.

[0037] When v i >0, v j > 0, the link connection time l between vehicles i and j i,j It is expressed as:

[0038]

[0039] When v i <0, v j <0, the link connection time l between vehicles i and j i,j It is expressed as:

[0040]

[0041] When v i >0, v j <0, the link connection time l between vehicles i and j i,j It is expressed as:

[0042]

[0043] When v i <0, v j > 0, the link connection time l between vehicles i and j i,j It is expressed as:

[0044]

[0045] The numerator represents the total driving distance when vehicle i communicates with vehicle j, and the denominator represents the relative driving speed of vehicle i to vehicle j.

[0046] For any vehicle i and its associated roadside unit r, the location coordinates are expressed as (x i ,y i ) and (x r ,y r ), the speed of vehicle i is represented by v i .

[0047] When v i > 0, the link connection time l between vehicle i and roadside unit r r,i It is expressed as:

[0048]

[0049] When v i When <0, the link connection time l between vehicle i and roadside unit r r,i It is expressed as:

[0050]

[0051] Among them, the numerator represents the total driving distance when vehicle i communicates with the roadside unit r, and the denominator represents the driving speed of vehicle i; it is necessary to calculate the link connection time only when the roadside unit decides to offload the task to the candidate vehicle for processing; there will be multiple forwarding nodes between the roadside unit that generates the task and the candidate vehicle. In this embodiment, it is only necessary to calculate the connection time between the roadside unit, the forwarding node and the candidate vehicle.

[0052] In this case, the roadside unit can communicate not only with vehicles within its coverage area, but also with vehicles outside its coverage area through forwarding vehicles; the transmission rate of adjacent vehicles i and j in time slot τ is It is expressed as:

[0053]

[0054] Among them, B V2V represents bandwidth, P V represents the transmission power of the vehicle, K represents the fixed loss, ω represents the noise power, and D i,j (τ) represents the distance between vehicle i and vehicle j, and σ represents the path loss factor.

[0055] Due to the movement of vehicles, the distance between vehicles i and j will change over time, resulting in fluctuations in the transmission rate; the average transmission rate is defined according to the above formula and is expressed as:

[0056]

[0057] Similarly, the transmission rate of vehicle i and roadside unit r in time slot τ is expressed as:

[0058]

[0059] According to the above definition, the average value of the transmission rate is:

[0060]

[0061] in, represents the transmission rate of vehicle i and roadside unit r in time slot τ.

[0062] Computing tasks on vehicle v: The computational process of offloading the tasks generated by the roadside unit r to vehicle v includes three steps: finding candidate vehicles, determining the transmission path, and executing the offloading decision. In the process of finding candidate vehicles, the roadside unit will send service requests to the vehicles within its coverage area. After receiving the request, each vehicle will further forward the request to the next-hop vehicle with which it communicates. However, the task has not actually been transmitted at this time, and whether the task is transmitted depends on the final offloading decision. After receiving the request, each vehicle records all the forwarding nodes that the request passes through before reaching the vehicle, and then the vehicle sends this information and the current status details back to the roadside unit.

[0063] After finding the candidate vehicles, we get all the vehicles that can communicate with the RSU, called candidate vehicles, thus creating the candidate vehicle set Θ τ ; Candidate vehicle set Θ τ The vehicles in the vehicle have one or more paths to the roadside unit. Assume that there are H paths from the roadside unit r to the candidate vehicle v, where a given path h is given by u h A single-hop link consists of: roadside unit r→vehicle h1→vehicle h2→…→vehicle v, and h∈H; the link connection time of path h It is the minimum value of the connection time of all single-hop links in path h, expressed as:

[0064]

[0065] The optimal transmission path p from the roadside unit r to the candidate vehicle v r,v It is the path with the longest connection time among all paths, expressed as:

[0066]

[0067] The minimum link connection time between all adjacent nodes is selected as the link connection time of path h to ensure the reliability of communication between all adjacent nodes in the path; the path with the largest link connection time is selected as the optimal path to prevent task transmission failure due to unstable communication.

[0068] After obtaining the candidate vehicles and the optimal transmission path, the roadside unit makes an unloading decision. Based on the unloading decision, the task is unloaded to the appropriate vehicle for calculation, and finally the calculation result of the vehicle is transmitted back to the roadside unit.

[0069] The total processing delay includes data transmission delay, queuing delay, calculation delay and result transmission delay; the data transmission delay of the unloading task from the roadside unit r to the vehicle v through the transmission path is: It is expressed as:

[0070]

[0071] If the vehicle v is within the coverage of the roadside unit r, data transmission is performed via single-hop transmission; otherwise, multi-hop transmission is performed.

[0072] The queue delay is expressed as The calculation of queue delay is divided into two parts. The first part consists of It is used to calculate the time required to process the tasks that were previously unloaded to vehicle v but have not been processed; the subsequent segment reflects the processing time of tasks assigned by other roadside units in the current time slot, whose indices are lower than that of roadside unit r:

[0073]

[0074] Among them, the task list v represents the tasks that have not been calculated in the task queue previously unloaded to vehicle v, and f v Represents the computing power of vehicle v.

[0075] Calculate the delay express:

[0076]

[0077] The processing results are very small compared to the data size, so the result transmission delay can be ignored.

[0078] The total processing delay from the roadside unit r to the vehicle v unloading task is expressed as:

[0079]

[0080] If the unloading decision selects a vehicle v that can communicate with the roadside unit r, the total processing delay is calculated according to the above formula, otherwise the total processing delay is set to 0.

[0081] In this embodiment, it is assumed that adjacent roadside units can communicate directly. The communication delay between any two roadside units is proportional to the number of communication hops. The total processing delay of the calculation offloaded to the roadside unit k includes data transmission delay, queuing delay, calculation delay and result transmission delay. The data transmission delay is expressed as:

[0082]

[0083] Among them, h r,k represents the number of hops required for data transmission between two roadside units, and λ represents the time required for single-hop transmission.

[0084] As before, the queue delay also consists of two parts, expressed as:

[0085]

[0086] Among them, the task list (k) represents the tasks that were previously unloaded to the roadside unit k but have not yet been calculated, and f k Represents the computing power of roadside unit k.

[0087] The computation delay is expressed as:

[0088]

[0089] Therefore, according to the above formula, the total processing delay of task offloading from roadside unit r to roadside unit k is expressed as:

[0090]

[0091] This embodiment assumes that the roadside unit can communicate directly with the cloud server. The data transmission delay required to offload the task generated by the roadside unit r to the cloud server is expressed as:

[0092]

[0093] Among them, r R2C Indicates the transmission rate between the roadside unit and the cloud server.

[0094] Since cloud servers usually have a large amount of computing resources, the queuing delay and computational delay in the computing process on the cloud server can be ignored. Similarly, the result transmission delay can also be ignored. Therefore, the total processing delay of the task offloaded from the roadside unit r to the cloud server is expressed as:

[0095]

[0096] In each time slot τ, each roadside unit generates a task. The goal of this embodiment is to optimize the offloading decision and minimize the total processing delay of all tasks generated by the RSU. The total processing delay required for the tasks generated by the roadside unit r is defined as:

[0097]

[0098] Therefore, the optimization objective is expressed as:

[0099]

[0100] Constraints C1 and C2 indicate that each task can only be unloaded and executed on one node; constraint C3 ensures that the task is completed within the task constraint time, while constraint C4 ensures that task processing will not fail due to communication interruption; constraint C5 ensures that the task will not be unloaded to a vehicle that cannot communicate with the roadside unit that generates the task; constraints C6 and C7 represent the effectiveness of data transmission; the optimization problem is a typical 0-1 problem, which is converted into a Markov decision process (MDP) problem in this embodiment and solved by the DRL method.

[0101] The method of this embodiment is a centralized solution. A "central server" is used as an agent to solve the model problem and provide task offloading decisions. S is designated as the state space and A is designated as the action space. At the beginning of each time slot τ, each roadside unit actively perceives the current road environment. The agent collects information from all computing nodes. The state is represented by s τ (s τ ∈S); the agent then takes action a τ (a τ ∈A), unload all tasks to the corresponding nodes; execute a τ After that, the system status is updated to s τ+1 , the agent gets reward r τ ,like Figure 3 shown.

[0102] State space: In each time slot τ, the agent collects task queue information from all computing nodes, so the state information is expressed as:

[0103]

[0104] Where C represents the total computing resources required to process all tasks in the task queue.

[0105] Action Space: Perception State s τ After that, the agent makes an unloading decision and the action space is expressed as:

[0106] a τ ={x1, x2, ..., x r , ..., x R}

[0107] Among them, x r represents the task unloading decision generated by roadside unit r, x R ∈{1,...,V,V+1,...,V+R,V+R+1}.

[0108] Reward function: The agent is in state s τ Next execution actions τ After that, you can get the reward immediately τ ; MDP is usually used to maximize the long-term return of the system. However, the goal of the optimization problem is to minimize the total processing delay of the task. Therefore, after converting the optimization problem into an MDP problem, the goal is to maximize the negative value of the objective function, which is expressed as:

[0109]

[0110] When the constraints C1 to C7 are met, the immediate reward r τis the negative of the total processing delay of all computational tasks generated in time slot τ; this penalty encourages the agent to choose behaviors that satisfy constraints C1 to C7.

[0111] The A3C algorithm is used because it evolves and performs asynchronous training with multiple agents that interact with the environment, and is suitable for environments with large action spaces. Each agent in A3C consists of an actor network and a critic network, and usually shares the same parameters in the non-output layers of the actor network and the critic network.

[0112] In time slot τ, state s τ is used as the input of the proxy network. After passing through the shared dual neural network (DNN), the actor network outputs π(a τ |s τ ; θ), while the critic network generates V(s) through a linear layer τ θ V ), where a τ represents an action, θ and θ V Denotes the parameters of the dual neural network at time slot τ.

[0113] In status τ Take action a τ , the agent gets the output s τ+1 and reward r τ , the agent repeats this reasoning process until τ reaches a given maximum step size T; the training goal is to minimize the loss of the actor network and the critic network, and the loss of the actor network is defined as:

[0114] L π (θ) = -logπ(a τ |s τ ;θ)(R τ -V(s τ θ V ))

[0115] Among them, R τ is the sum of discounted returns, expressed as:

[0116]

[0117] Among them, γ is a constant representing the reward discount factor.

[0118] The loss of the critic network is defined as:

[0119] L V (θ V )=(R τ -V(s τ θ V )) 2

[0120] After repeated reasoning and training, a well-trained agent is obtained, which can make the best decision to maximize long-term benefits. The algorithm process is as follows Figure 4 shown.

[0121] This embodiment provides an RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning, which is used to optimize the computing resource utilization efficiency in the RSU-assisted Internet of Vehicles (IoV). With the development of intelligent transportation systems, the data interaction and computing requirements between vehicles and road infrastructure are increasing. Traditional task offloading strategies are difficult to cope with dynamically changing network environments and complex computing tasks. Especially in the IoV scenario, how to efficiently use distributed computing resources (including cloud servers, RSUs and vehicles) to complete multi-hop task offloading while ensuring communication stability has become an important challenge.

[0122] To this end, this embodiment proposes an innovative multi-hop task offloading mechanism, which accurately evaluates the connection stability between nodes by establishing a vehicle mobility model and a communication model, and designs a task offloading decision process based on the A3C algorithm; specifically, this embodiment models the task offloading problem as a high-dimensional Markov decision process (MDP) and uses the A3C algorithm to solve the problem.

[0123] The A3C algorithm effectively handles the optimization problem in a large-scale action space by asynchronously training multiple agents, thereby realizing a better task offloading strategy in a complex and changeable Internet of Vehicles environment. Experimental results show that compared with the existing task offloading strategy, the method of this embodiment shows significant performance improvement in reducing task processing delay and improving computing resource utilization, especially in multi-RSU and high dynamic scenarios, which significantly improves the system's task completion rate and resource utilization efficiency. In addition, this embodiment also has good scalability and can adapt to more complex application requirements in future Internet of Vehicles. By collaborating with cloud computing, edge computing and other resources, the task offloading strategy of this embodiment provides important technical support for the next generation of intelligent transportation systems.

[0124] In addition to the above embodiments, the present invention may also have other implementation modes. Any technical solution formed by equivalent replacement or equivalent transformation falls within the protection scope required by the present invention.

Claims

1. An RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning, characterized in that: The following steps are involved: S1. Establish an edge unloading system, which includes a cloud server, several roadside units and several vehicles; S2. In each time slot, the roadside unit senses the current road environment and generates a computing task; S3: The roadside unit finds candidate vehicles that can receive computing tasks by broadcasting service requests. The candidate vehicles feed back their status information to the roadside unit, including the task queue and computing resource availability. S4, the roadside unit calculates the transmission path of each candidate node based on the received information and selects the optimal transmission path. The candidate nodes include the vehicle, other roadside units and cloud servers; then the total processing delay after the task is offloaded to the vehicle, roadside unit and cloud server is calculated; S5. Through the A3C algorithm, the task offloading strategy is optimized under the Markov decision process framework to maximize the negative reward of task processing. The optimal offloading strategy is obtained by continuously interacting with the environment to update the parameters and minimize the total processing delay.

2. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 1 is characterized in that: In step S1, the cloud server, roadside unit and vehicle are regarded as three different types of computing nodes, and the set R = {1, 2, ..., r, ..., R} is used to represent the evenly distributed roadside unit, and the set V = {1, 2, ..., v, ..., V} is used to represent the vehicle; the edge unloading system operates in a time slot-by-time slot manner, dividing the time into multiple time slots, which are represented by the set {1, 2, ..., τ, ...}; at time slot τ, the vehicles in the coverage area of ​​the roadside unit r are represented by the set Represents; In this edge unloading system, vehicles i and j that cannot communicate directly establish communication through vehicle v, and vehicle v is defined as a forwarding vehicle; adjacent roadside units communicate with each other or directly with the cloud server; roadside units communicate not only with vehicles within their coverage, but also with vehicles outside their coverage through forwarding vehicles.

3. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 2 is characterized in that: In step S2, at time slot τ, the roadside unit r senses the road environment and generates a computing task (s r , c r , t r ), where s r Indicates the task size, c r represents the computing resources required to process the task, t r Indicates the execution time limit of the task; each task generated by the roadside unit is offloaded to the cloud server, roadside unit or vehicle for execution, and x r,n represents an uninstall decision, where and n∈{1,...,V,V+1,...,V+R,V+R+1}.

4. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 3 is characterized in that: In step S3, it is assumed that the vehicle is traveling at a constant speed only along the x-axis, the speed in the right direction is a positive value, and the speed in the left direction is a negative value, and the communication range between the vehicles is represented by a constant denoted by , while the communication range between the vehicle and the roadside unit is denoted by ψ; For any adjacent vehicles i and j, the position coordinates are expressed as (x i ,y i ) and (x j ,y j ), the speed is represented by v i and v j , if the speed is greater than 0, it means that the vehicle is moving in the correct direction, otherwise, it means that the vehicle is moving in the opposite direction of the correct direction; When v i >0, v j > 0, the link connection time l between vehicles i and j i,j It is expressed as: When v i <0, v j <0, the link connection time l between vehicles i and j i,j It is expressed as: When v i >0, v j <0, the link connection time l between vehicles i and j i,j It is expressed as: When v i <0, v j > 0, the link connection time l between vehicles i and j i,j It is expressed as: The numerator represents the total driving distance when vehicle i communicates with vehicle j, and the denominator represents the relative driving speed of vehicle i to vehicle j.

5. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 4 is characterized in that: In step S3, for any vehicle i and its associated roadside unit r, the position coordinates are expressed as (x i ,y i ) and (x r ,y r ), the speed of vehicle i is represented by v i ; When v i > 0, the link connection time l between vehicle i and roadside unit r r,i It is expressed as: When v i When <0, the link connection time l between vehicle i and roadside unit r r,i It is expressed as: Among them, the numerator represents the total driving distance when vehicle i communicates with the roadside unit r, and the denominator represents the driving speed of vehicle i; the link connection time is only calculated when the roadside unit decides to offload the task to the candidate vehicle for processing; there are multiple forwarding nodes between the roadside unit that generates the task and the candidate vehicle, and only the connection time between the roadside unit, the forwarding node and the candidate vehicle needs to be calculated.

6. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 5 is characterized in that: In step S3, the roadside unit not only communicates with vehicles within its coverage area, but also communicates with vehicles outside its coverage area through forwarding vehicles; the transmission rate of adjacent vehicles i and j in time slot τ It is expressed as: Among them, B V2V represents bandwidth, P V represents the transmission power of the vehicle, K represents the fixed loss, ω represents the noise power, and D i,j (τ) represents the distance between vehicle i and vehicle j, σ represents the path loss factor; Due to the movement of vehicles, the distance between vehicles i and j changes over time. The average transmission rate is defined according to the above formula and is expressed as: The transmission rate of vehicle i and roadside unit r in time slot τ is expressed as: According to the above definition, the average value of the transmission rate is: in, represents the transmission rate of vehicle i and roadside unit r in time slot τ.

7. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 6 is characterized in that: In step S3, the calculation process of offloading the task generated by the roadside unit r to the vehicle v includes three steps: finding candidate vehicles, determining the transmission path, and executing the offloading decision. In the process of finding candidate vehicles, the roadside unit sends a service request to the vehicles within its coverage area. After receiving the request, each vehicle further forwards the request to the next-hop vehicle communicating with it. Whether to transmit the task depends on the final offloading decision. After receiving the request, each vehicle records all forwarding nodes that the request passes through before reaching the vehicle, and then the vehicle sends this information and current status details back to the roadside unit.

8. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 7 is characterized in that: In step S4, all vehicles that can communicate with the roadside unit are called candidate vehicles, thereby creating a candidate vehicle set θ τ ; Candidate vehicle set Θ τ The vehicles in the vehicle have one or more paths to the roadside unit. Assume that there are H paths from the roadside unit r to the candidate vehicle v, where a given path h is given by u h A single-hop link consists of: roadside unit r→vehicle h1→vehicle h2→…→vehicle v, and h∈H; the link connection time of path h It is the minimum value of the connection time of all single-hop links in path h, expressed as: The optimal transmission path p from the roadside unit r to the candidate vehicle v r,v It is the path with the longest connection time among all paths, expressed as: The minimum link connection time between all adjacent nodes is selected as the link connection time of path h, and the path with the maximum link connection time is selected as the optimal path.

9. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 8 is characterized in that: In step S4, after obtaining the candidate vehicle and the optimal transmission path, the roadside unit makes an unloading decision. According to the unloading decision, the task is unloaded to the candidate vehicle for calculation, and finally the calculation result of the vehicle is transmitted back to the roadside unit; The total processing delay includes data transmission delay, queuing delay, calculation delay and result transmission delay; the data transmission delay of the unloading task from the roadside unit r to the vehicle v through the transmission path: roadside unit It is expressed as: If vehicle v is within the coverage of roadside unit r, data transmission is performed via single-hop transmission; otherwise, multi-hop transmission is performed; the queue delay is expressed as The calculation of queue delay is divided into two parts. The first part consists of It is used to calculate the time required to process the tasks that were previously unloaded to vehicle v but have not been processed; the subsequent segment reflects the processing time of tasks assigned by other roadside units in the current time slot, whose indices are lower than that of roadside unit r: Among them, the task list v represents the tasks that have not been calculated in the task queue previously unloaded to vehicle v, and f v represents the computing power of vehicle v, and the computing delay is express: Ignoring the result transmission delay, the total processing delay from the roadside unit r to the vehicle v unloading task is expressed as: If the unloading decision selects a vehicle v that can communicate with the roadside unit r, the total processing delay is calculated according to the above formula, otherwise the total processing delay is set to 0; Assuming that adjacent roadside units communicate directly, the communication delay between any two roadside units is proportional to the number of communication hops. The total processing delay of the calculation offloaded to roadside unit k includes data transmission delay, queuing delay, calculation delay and result transmission delay; the data transmission delay is expressed as: Among them, h r,k represents the number of hops required for data transmission between two roadside units, and λ represents the time required for single-hop transmission; the queue delay also consists of two parts, expressed as: Among them, the task list (k) represents the tasks that were previously unloaded to the roadside unit k but have not yet been calculated, and f k represents the computing power of roadside unit k; the computing delay is expressed as: The total processing delay of task offloading from roadside unit r to roadside unit k is expressed as: Assuming that the roadside unit communicates directly with the cloud server, the data transmission delay required to offload the task generated by the roadside unit r to the cloud server is expressed as: Among them, r R2C represents the transmission rate between the roadside unit and the cloud server; the queuing delay, calculation delay and transmission delay in the calculation process on the cloud server are all negligible, and the total processing delay of the task offloaded from the roadside unit r to the cloud server is expressed as: In each time slot τ, each roadside unit generates a task, and the total processing delay required for the tasks generated by roadside unit r is defined as: Therefore, the optimization objective is expressed as: C3:T r ≤t r ,r∈R C5:T r ≠0,r∈R Constraints C1 and C2 indicate that each task can only be offloaded and executed on one node; constraint C3 ensures that the task is completed within the task constraint time, while constraint C4 ensures that task processing will not fail due to communication interruption; constraint C5 ensures that the task will not be offloaded to a vehicle that cannot communicate with the roadside unit that generated the task; constraints C6 and C7 represent the effectiveness of data transmission.

10. The RSU-assisted multi-hop task offloading method based on asynchronous deep reinforcement learning according to claim 9, characterized in that: In step S5, a central server is used as an agent to solve the model problem and provide task offloading decisions. S is designated as the state space and A is designated as the action space. At the beginning of each time slot τ, each roadside unit actively perceives the current road environment. The agent collects information from all computing nodes. The state is represented by s τ (s τ ∈S); the agent then takes action a τ (a τ ∈A), unload all tasks to the corresponding nodes; execute a τ After that, the system status is updated to s τ+1 , the agent gets reward r τ ; State space: In each time slot τ, the agent collects task queue information from all computing nodes, so the state information is expressed as: Where C represents the total computing resources required to process all tasks in the task queue; Action Space: Perception State s τ After that, the agent makes an unloading decision and the action space is expressed as: fluent τ {x1,x2,...,x r ,...,x R } Among them, x r represents the task unloading decision generated by roadside unit r, x R ∈{1,...,V,V+1,...,V+R,V+R+1}; Reward function: The agent is in state s τ Next execution actions τ After that, you can get the reward immediately τ , after converting the optimization problem into a Markov decision process problem, the goal is to maximize the negative value of the objective function, which is expressed as: When the constraints C1 to C7 are met, the immediate reward r τ It is the negative value of the total processing delay of all computing tasks generated in the time slot τ; Each agent in the A3C algorithm consists of an actor network and a critic network. In the time slot τ, the state s τ is used as the input of the proxy network. After passing through the shared dual neural network, the actor network outputs π(a τ |s τ ; θ), while the critic network generates V(s) through a linear layer τ θ V ), where a τ represents an action, θ and θ V represents the parameters of the dual neural network at time slot τ; In status τ Take action a τ , the agent gets the output s τ+1 and reward r τ , the agent repeats this reasoning process until τ reaches a given maximum step size T; the training goal is to minimize the loss of the actor network and the critic network, and the loss of the actor network is defined as: L π (θ)=-logπ(a τ |s τ ;θ)(R τ -V(s τ ;θ V )) Among them, R τ is the sum of discounted returns, expressed as: Among them, γ is a constant, representing the reward discount factor; The loss of the critic network is defined as: L V (i V )=(R τ -V(s τ ;θ V )) 2 After repeated reasoning and training, the best decision is made.

Citation Information

Cited By

  • Optimization method for edge computing task unloading and resource scheduling of Internet of Vehicles

    CN120916201A

  • Optimization method for task offloading and resource scheduling of edge computing in internet of vehicles

    CN120916201B