A "zero interruption" task offloading method based on hierarchical reinforcement learning
By adopting a hierarchical reinforcement learning method based on road dynamic topology environment, the problem of latency and stability of task "0 interrupts" offloading is solved, and high-stability task offloading and network throughput are achieved under low latency conditions.
Patent Information
- Application Number
- CN202510152979.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-12
AI Technical Summary
In the dynamic road topology environment, the uninstallation of the task "0 interrupt" faces delay and stability problems, especially due to the rapid changes in vehicle location and speed, resulting in frequent interruptions in node connections, threatening the continuity and real-time nature of data processing.
Using a method based on hierarchical reinforcement learning, a road vehicle communication model and mobility model is constructed, a task transmission link stability entropy value model is established, and the task offload optimization problem is decomposed into two sub-optimization goals of task allocation and 0 interrupt routing path planning. Using a hierarchical reinforcement learning algorithm, the path selection problem is modeled as partially observable Markov processes, and design high-level policy networks and low-level policy networks, as well as their corresponding action space, state space and reward functions.
On the premise of meeting the delay requirements, the task offloading of "0 interrupts" is ensured, and network throughput is maximized, which improves the stability and reliability of data transmission.
Smart Images

Figure CN119629671B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of road edge computing network technology, and in particular to a task "0 interruption" unloading method based on hierarchical reinforcement learning. Background Art
[0002] With the rapid development of information technology, Vehicle Edge Computing (VEC), as a key component of Intelligent Transportation Systems (ITS), has the significant advantages of low latency, high bandwidth, and high reliability by migrating computing tasks from traditional cloud servers to edge nodes around vehicles. With the development of the times, emerging businesses have put forward unprecedented stringent requirements on data transmission conditions - that is, the pursuit of the ideal state of "0 interruption" to ensure the ultimate stability, reliability, and ultra-low latency of data transmission.
[0003] However, the rapid changes in the position and speed of road vehicles not only increase the complexity and uncertainty of task offloading, but also may cause frequent interruptions in the connection between nodes, seriously threatening the continuity and real-time performance of data processing. In addition, the deployment of road side units (RSUs) is often restricted by multiple factors such as the environment and cost, showing the characteristics of uneven distribution, which further aggravates the instability of task "zero interruption" offloading, making it difficult to meet the latency and stability of task offloading in a dynamic road topology environment. Summary of the invention
[0004] The purpose of the present invention is to provide a task "0 interruption" offloading method based on hierarchical reinforcement learning to solve the problems existing in the background technology.
[0005] To achieve the above object, the present invention provides a task "0 interruption" offloading method based on hierarchical reinforcement learning, comprising the following steps:
[0006] S1. Establish a road vehicle communication model and a mobility model, and construct a transmission link stability entropy value model of tasks in the VEC architecture through the road vehicle communication model and the mobility model;
[0007] S2, decompose the multi-task cooperative offloading optimization problem with dynamic resource coordination and intelligent scheduling mechanism into two sub-optimization goals: differentiated task allocation and zero-interrupt routing path planning;
[0008] S3. Based on the hierarchical reinforcement learning algorithm, the process of finding a zero-interruption route in the dynamic road topology is modeled as a partially observable Markov process, and the high-level policy network and the low-level policy network and their corresponding action space, state space and reward function are established hierarchically.
[0009] Preferably, the content of S1 is as follows:
[0010] S11. Construct road vehicle communication model, vehicle node With communication node Data transfer rate between for:
[0011] ; (1)
[0012] in and represents the channel bandwidth and channel gain between node i and node j; Represents its transmission power; represents the background noise power, which is assumed to be Gaussian white noise; Represents the signal interference between the two;
[0013] S12. Construct a road vehicle mobility model, using express and The duration of the direct link interruption between and to represent the vehicles at time t With vehicle The coordinates of and To represent the speed of the vehicle at time t, the distance between vehicles can be simplified as ;
[0014] S13, build a communication link stability entropy model, using To represent the stability entropy value of the link, The larger the value, the worse the stability of the transmission link. and , and the calculated connection time To calculate :
[0015] ; (2)
[0016] in, A stability entropy calculation model for the transmission link between vehicles traveling in the same direction; It indicates the situation when the vehicle is traveling in the opposite direction to the vehicle; It indicates the communication status between the vehicle and the roadside unit RSU. represents the vehicle standard speed deviation evaluation parameter, Represents the connection time evaluation parameters, represents the speed consistency evaluation parameter, Represents the weight coefficient of each parameter, Indicates the maximum speed of the vehicle.
[0017] Preferably, the content of S11 is as follows:
[0018] S111. Construct a time delay model for the task. Generated tasks , the total time of the uninstallation execution can be The total time is represented by the transmission time and calculate execution time composition:
[0019] ; (3)
[0020] in The task k generated by vehicle v is determined by Node devices are used for processing.
[0021] S112, any task Execution time The calculation formula is constructed as:
[0022] ; (4)
[0023] in Used to represent nodes The computing power of the device, Represents the computational complexity of task k. Binary variable To indicate the offloading decision of the task. If the task is decided to be processed locally, that is, ,but ,otherwise ; Represents a collection of road nodes.
[0024] S113. Task transmission time , when the vehicle The generated k-type tasks determine the local vehicle When performing calculations, its task transfer time .
[0025] When you need to offload tasks to other service nodes When performing task calculation, the task transmission time is:
[0026] ; (5)
[0027] in, represents the task size of task k; Indicates the size of data returned after task k is completed; Represents the downlink transmission rate between nodes ij.
[0028] Preferably, the content of S12 is as follows:
[0029] S121. When two vehicles are traveling in the same direction, When the two vehicles Link time between It can be calculated as:
[0030] ; (6)
[0031] in, Representation Node The communication radius; Indicates the node at time t The distance between
[0032] when When the two vehicles Connection time between It can be calculated as:
[0033] ; (7)
[0034] S122. When two vehicles are traveling in opposite directions, Connection time between It can be calculated as:
[0035] (8)
[0036] S123, Vehicle Connection time with roadside unit RSU-j It can be calculated as:
[0037] ; (9)
[0038] Represents the horizontal and vertical coordinates of the roadside unit RSU-j; Indicates vehicle The horizontal and vertical coordinates at time t.
[0039] Preferably, the content of S13 is as follows:
[0040] S131, vehicle standard speed deviation evaluation parameters , which is evaluated by the difference between the vehicle speed and the road reference speed, using To represent the prescribed speed limit of the road, the speed deviation is calculated as:
[0041] ; (10)
[0042] When the vehicle speed and When the speed difference from the prescribed speed limit is large, If the value is large, it indicates that the vehicle's driving conditions may change in the future, and the instability will increase. In order to ensure the consistency of the magnitude of the evaluation model, Perform normalization:
[0043] ; (11)
[0044] S132. Connection time evaluation parameters The calculation is mainly based on the calculated connection time , the longer the connection time, the better the potential stability of the link:
[0045] ; (12)
[0046] because It is an inverse proportional function, and its maximum and minimum values cannot be estimated. Therefore, we consider using nonlinear function transformation to perform normalization:
[0047] ; (13)
[0048] S133, to consider the impact of the speed relationship between vehicles on the link stability, an additional parameter is introduced To measure the relative difference in the speeds of the two cars:
[0049] ; (14)
[0050] This parameter can more accurately assess link stability in situations where the speed gap is small but the gap from the reference speed is large.
[0051] S134. For any multi-hop path , and use To represent the path An exponential growth strategy is used to calculate the total stability entropy of Multi-hop path with relay nodes The total stability entropy .
[0052] ; (15)
[0053] in, Representative path Middle p The stability entropy value calculated by the segment link, Stability Index .
[0054] Preferably, the content of S2 is as follows:
[0055] S21. The task offloading optimization objective is expressed as maximizing the average number of completed tasks. The total number of tasks generated in each time slot t is represented by N, and the task completion rate is:
[0056] ; (16)
[0057] Therefore, the optimization goal of the task offloading problem is expressed as:
[0058] ; (17)
[0059] Indicates the number of time slots in statistics;
[0060] S22. Decompose the optimization objective of the task offloading problem into two more focused optimization sub-problems: the classic task and service device matching problem and the path planning problem under the constraints of path connectivity and stability.
[0061] Preferably, the content of S22 is as follows:
[0062] S221. For the matching problem between tasks and service devices, the total execution time of tasks in time slot t is defined as the longest execution time among all tasks, and express:
[0063] ; (18)
[0064] in Represents the task allocation matrix.
[0065] because The formula consists of two parts: calculation execution time and transmission time. However, the accurate routing transmission time cannot be obtained under the current time scale, so it is used in this problem. To indicate the starting node With the target node Estimated transfer time between:
[0066] ; (19)
[0067] in is the median transmission rate.
[0068] Therefore It is expressed as:
[0069] ; (20)
[0070] The optimization objective of this sub-problem is:
[0071] ;(twenty one)
[0072] S222: For the path planning problem under the constraints of path connectivity and stability, N task allocation schemes can be obtained through S221. The successful implementation of each allocation scheme (obtaining a path that satisfies the connectivity and stability constraints of task k) ) will receive corresponding rewards Rk, so the optimization problem can be expressed as a reward maximization problem:
[0073] ; (twenty two)
[0074] in, Represents the routing path selection for task k.
[0075] Preferably, the content of S3 is as follows:
[0076] S31. The path selection problem in a dynamic environment is modeled as a POMDP, which includes the observation (O), state (S), action (A), and reward (R) of the environment. A deep reinforcement learning (DRL) algorithm is used to solve its optimization challenges. In order to solve the dynamic topology problem, a cluster-assisted hierarchical reinforcement learning CAHRL algorithm is designed.
[0077] S32. At time t, the road cloud platform collects road information within the observed range, including vehicles (V) and roadside units (RSU). The definition is as follows:
[0078] ;(twenty three)
[0079] in, Represents the vehicle index list within the current road segment, Represents a list of vehicle speed information. and Represents a list of location information for each vehicle. RSU observation information The definition is as follows:
[0080] ;(twenty four)
[0081] in Represents the processing queue list of RSU; and Represents the location information list of RSU; Indicates the index list of RSU.
[0082] S33, the length is The road grid is divided into several blocks to characterize the road conditions. Due to the hierarchical decision-making design framework, the state information based on high-level decisions and low-level decisions is also different. is defined as follows:
[0083] ; (25)
[0084] in, represents the task type of task k; Indicates the current node information; Indicates the target node information; Indicates the number of nodes that can be selected in the current grid block n; Indicates the current hop count;
[0085] Low level status is defined as follows:
[0086] ; (26)
[0087] in, Indicates node speed information; and Indicates node location information;
[0088] S34. The design of the action space is also divided into two layers, which are defined as follows:
[0089] High-level actions Select n block options:
[0090] ; (27)
[0091] in, Indicates an optional block within the road;
[0092] Low-level actions Select m nodes in the selected block:
[0093] ; (28)
[0094] in, Indicates optional nodes in the block;
[0095] S35. In the reward part design, the reward formula for low-level decision making is defined as:
[0096] ; (29)
[0097] ; (30)
[0098] in is the reward discount coefficient, which serves to unify the reward scale; Represents the reward for step n; Indicates reward for success; Indicates the basic step reward.
[0099] The reward formula for the high-level strategy is defined as:
[0100] ; (31)
[0101] in, Indicates whether nodes ij can communicate with each other.
[0102] Therefore, the present invention adopts the above-mentioned task "0 interruption" offloading method based on hierarchical reinforcement learning, which can ensure "0 interruption" task offloading while meeting the delay requirements and maximize network throughput.
[0103] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] Figure 1 It is a flowchart of a method for unloading a task with “0 interruption” based on hierarchical reinforcement learning according to the present invention;
[0105] Figure 2 This is a road unloading network scenario diagram of a task "0 interruption" unloading method based on hierarchical reinforcement learning in the present invention. DETAILED DESCRIPTION
[0106] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0107] See also Figure 1-Figure 2 , a task "0 interruption" offloading method based on hierarchical reinforcement learning, comprising the following steps:
[0108] S1. Establish a road vehicle communication model and a mobility model, and build a transmission link stability entropy value model of the tasks in the VEC architecture through the road vehicle communication model and the mobility model.
[0109] S11: Construct a road vehicle communication model and set vehicle nodes With communication node Data transfer rate between for:
[0110] ; (1)
[0111] node It can be the vehicle at the current time t Neighboring vehicles within the communication range can also be vehicles i A roadside unit capable of communication. and Representative Node With Node The channel bandwidth and channel gain between them; Represents its transmission power; represents the background noise power, which is assumed to be Gaussian white noise; Represents the signal interference between the two.
[0112] S11 includes the following steps:
[0113] S111. Construct a time delay model for the task. Generated tasks , the total time of the uninstallation execution can be The total time is represented by the transmission time and calculate execution time composition:
[0114] ; (3)
[0115] in It means that the task k generated by vehicle v is decided to be processed by the jth node device.
[0116] S112, any task Execution time The calculation formula is constructed as:
[0117] ; (4)
[0118] in Used to represent the computing power of the device of node j, Represents the computational complexity of task k. Binary variable To indicate the offloading decision of the task. If the task is decided to be processed locally, that is, ,but ,otherwise ; Represents a collection of road nodes.
[0119] S113. Task transmission time , when the vehicle The generated k-type tasks determine the local vehicle When performing calculations, its task transfer time When you need to offload tasks to other service nodes j When performing task calculations:
[0120] ; (5)
[0121] in, represents the task size of task k; Indicates the size of data returned after task k is completed; Representation Node The downlink transmission rate between.
[0122] S12. Construct a road vehicle mobility model. For a pair of nodes and If the distance between them does not exceed the communication range, they are considered connected. express and The duration of the direct link interruption between the two. and to represent the vehicles at time t With vehicle Use and to represent the speed of the vehicle at time t. The distance between vehicles at the current moment can be simplified as .
[0123] S12 includes the following steps:
[0124] S121. When two vehicles are traveling in the same direction, When the two vehicles Link time between It can be calculated as:
[0125] ; (6)
[0126] in, Representation Node The communication radius; Indicates the node at time t The distance between
[0127] when When the two vehicles Connection time between It can be calculated as:
[0128] ; (7)
[0129] S122, when two vehicles are traveling in opposite directions, the connection time between the two vehicles ij It can be calculated as:
[0130] ; (8)
[0131] S123, connection time between vehicle i and roadside unit RSU-j It can be calculated as:
[0132] ; (9)
[0133] in, Represents the horizontal and vertical coordinates of the roadside unit RSU-j; Indicates vehicle The horizontal and vertical coordinates at time t.
[0134] S13, construct a communication link stability entropy model, for the single-hop transmission link between the passing node i and the node j, use To represent the stability entropy value of the link, The larger the value, the worse the stability of the transmission link. and , and the calculated connection time To calculate :
[0135] ; (2)
[0136] in, A stability entropy calculation model for the transmission link between vehicles traveling in the same direction; It indicates the situation when the vehicle is traveling in the opposite direction to the vehicle; It indicates the communication status between the vehicle and the roadside unit RSU. When it indicates RSU, The latter two cases can use the same entropy calculation model. represents the vehicle standard speed deviation evaluation parameter, Represents the connection time evaluation parameters, represents the speed consistency evaluation parameter, Represents the weight coefficient of each parameter, Indicates the maximum speed of the vehicle.
[0137] S13 includes the following steps:
[0138] S131, vehicle standard speed deviation evaluation parameters , which is evaluated by the difference between the vehicle speed and the road reference speed, using To represent the prescribed speed limit of the road, the speed deviation is calculated as:
[0139] ; (10)
[0140] When the vehicle speed and When the speed difference from the prescribed speed limit is large, If the value is large, it indicates that the vehicle's driving conditions may change in the future, and the instability will increase. In order to ensure the consistency of the magnitude of the evaluation model, Perform normalization:
[0141] ; (11)
[0142] S132. Connection time evaluation parameters The calculation is mainly based on the calculated connection time , the longer the connection time, the better the potential stability of the link:
[0143] ; (12)
[0144] because It is an inverse proportional function, and its maximum and minimum values cannot be estimated. Therefore, we consider using nonlinear function transformation to perform normalization:
[0145] ; (13)
[0146] S133, to consider the impact of the speed relationship between vehicles on the link stability, an additional parameter is introduced To measure the relative difference in the speeds of the two cars:
[0147] ; (14)
[0148] This parameter can more accurately assess link stability in situations where the speed gap is small but the gap from the reference speed is large.
[0149] S134. For any multi-hop path , and use To represent the path An exponential growth strategy is used to calculate the total stability entropy of Multi-hop path with relay nodes The total stability entropy .
[0150] ; (15)
[0151] in, Representative path Middle p The stability entropy value calculated by the segment link, represents the stability index of the road segment The lower the value, the more stable the environment.
[0152] S2. Decompose the multi-task collaborative offloading optimization problem with dynamic resource coordination and intelligent scheduling mechanism into two sub-optimization goals: differentiated task allocation and zero-interruption routing path planning.
[0153] S2 includes the following steps:
[0154] S21, the task offloading optimization goal is expressed as maximizing the average number of completed tasks. The total number of tasks generated in each time slot t is represented by N, and the task completion rate is
[0155] ; (16)
[0156] Therefore, the optimization objective of the task offloading problem is expressed as
[0157] ; (17)
[0158] Represents the number of time slots in the statistics; the meaning of the overall formula is: to maximize the average number of tasks completed within the total time slot T.
[0159] S22. Decompose the optimization objective of the task offloading problem into two more focused optimization sub-problems: the classic task and service device matching problem and the path planning problem under the constraints of path connectivity and stability.
[0160] S22 includes the following steps:
[0161] S221. For the matching problem between tasks and service devices, the total execution time of tasks in time slot t is defined as the longest execution time among all tasks, and express.
[0162] ; (18)
[0163] in Represents the task allocation matrix.
[0164] because The formula consists of two parts: calculation execution time and transmission time. However, the accurate routing transmission time cannot be obtained under the current time scale, so it is used in this problem. To indicate the starting node With the target node Estimated transfer time between
[0165] ; (19)
[0166] in is the median transmission rate.
[0167] Therefore Expressed as
[0168] ; (20)
[0169] Therefore, subproblem P2 can be formulated as
[0170] ;(twenty one)
[0171] S222: For the path planning problem under the constraints of path connectivity and stability, the goal is to find a routing path from the task vehicle node to the target node that satisfies both connectivity and stability requirements based on the node topology information (including device location and speed) of the current time slot to ensure successful unloading of the task. Through step S2021, N task allocation schemes can be obtained. The successful implementation of each allocation scheme (obtaining a path h that satisfies the connectivity and stability constraints of task k) will receive a corresponding reward Rk, so the optimization problem can be expressed as a reward maximization problem.
[0172] ; (twenty two)
[0173] in, Represents the routing path selection for task k.
[0174] S3. Based on the hierarchical reinforcement learning algorithm, the process of finding a zero-interruption route in the dynamic road topology is modeled as a partially observable Markov process, and the high-level policy network and the low-level policy network and their corresponding action space, state space and reward function are established hierarchically.
[0175] S31. The path selection problem in a dynamic environment is modeled as a POMDP, which includes the observation (O), state (S), action (A), and reward (R) of the environment. A deep reinforcement learning (DRL) algorithm is used to solve its optimization challenges. In order to solve the dynamic topology problem, a CAHRL algorithm is designed. The algorithm effectively copes with the challenges brought by large-scale discrete state-action space through clustering, improves the stability and decision-making efficiency of the algorithm; at the same time, the complex multi-dimensional optimization objectives are hierarchically decided, which reduces the overall complexity of the problem and solves the sparse reward problem, improving the generalization ability and interpretability of the algorithm.
[0176] S32. At time t, the road cloud platform collects road information within the observed range, including vehicles (V) and roadside units (RSU). The definition is as follows:
[0177] ;(twenty three)
[0178] in, Represents the vehicle index list within the current road segment, Represents a list of vehicle speed information. and Represents a list of location information for each vehicle. RSU observation information The definitions are as follows:
[0179] ;(twenty four)
[0180] in Represents the processing queue list of RSU; and Represents the location information list of RSU; Indicates the index list of RSU.
[0181] S33, the length is The road is gridded into several blocks to characterize the road conditions. Since this embodiment adopts a hierarchical decision-making design framework, the state information based on the high-level decision and the low-level decision is also different. The specific design is as follows:
[0182] The high-level state at time t is defined as follows:
[0183] ; (25)
[0184] in, represents the task type of task k; Indicates the current node information; Indicates the target node information; Indicates the number of optional nodes in the current grid block n; Indicates the current hop count.
[0185] Low level status is defined as follows:
[0186] ; (26)
[0187] in, Indicates node speed information; and Represents node location information.
[0188] Mainly represents the status of nodes in the selected block.
[0189] Set the task type The purpose of adding states is to expect the neural network to be able to Situations, learning to obtain differentiated action strategies to maximize rewards.
[0190] S34. The design of the action space is also divided into two layers, which are defined as follows:
[0191] High-level actions Select n block options,
[0192] ; (27)
[0193] in, Represents an optional block within a road.
[0194] Low-level actions Select m nodes in the selected block,
[0195] ; (28)
[0196] in, Indicates an optional node in a block.
[0197] Note that the size of m in the action space is related to the set block option n, where m is the theoretical maximum number of options available under given conditions.
[0198] S35. In the reward part design, the reward formula for low-level decision making is defined as:
[0199] ; (29)
[0200] ; (30)
[0201] in is the reward discount factor; Represents the reward for step n; Indicates reward for success; Indicates the basic step reward. Its function is to unify the reward scale. Since there are differences in requirements between the two types of tasks AB, Different from the reward discount coefficient in the formula A differentiated design is also carried out so that it can reflect the sensitivity of different types of tasks to stability parameters and provide a training basis for the differentiated strategy of the neural network.
[0202] The reward formula for the high-level strategy is defined as
[0203] ; (31)
[0204] in, Characterization Node Whether they can communicate with each other. For the indicator function The main purpose of the design is to guide the high-level network to choose (communicable) action, so the indicator function will be based on The situation provides a basic reward of appropriate scale, while the differences between different actions are affected by the low-level action rewards. Please note that scale adaptation in reward design is very important and will directly affect learning efficiency and algorithm stability.
[0205] Therefore, the present invention adopts the above-mentioned task "0 interruption" offloading method based on hierarchical reinforcement learning, which can ensure "0 interruption" task offloading while meeting the delay requirements and maximize network throughput.
[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A "0-interruption" task offloading method based on hierarchical reinforcement learning, characterized in that: The following steps are involved: S1. Establish a road vehicle communication model and a mobility model, and construct a transmission link stability entropy value model of tasks in the VEC architecture through the road vehicle communication model and the mobility model; S11. Construct road vehicle communication model, vehicle node With communication node Data transfer rate between for: ;(1) in and Representative Node With Node The channel bandwidth and channel gain between them; Represents its transmission power; represents the background noise power; Represents the signal interference between two nodes; S12. Construct a road vehicle mobility model, using express and The duration of the interruption of the direct link between them; use and to represent the vehicles at time t With vehicle The coordinates of and To represent the speed of the vehicle at time t, the distance between vehicles is simplified as ; S13, build a communication link stability entropy model, using To represent the stability entropy value of the link, according to the current vehicle speed and , and the calculated connection time To calculate : ; (2) in, A stability entropy calculation model for the transmission link between vehicles traveling in the same direction; It indicates the situation when the vehicle is traveling in the opposite direction to the vehicle; It indicates the communication status between the vehicle and the roadside unit RSU; represents the vehicle standard speed deviation evaluation parameter; Represents connection time evaluation parameters; represents the speed consistency evaluation parameter; Represents the weight coefficient of each parameter; Indicates the maximum speed of the vehicle; S2, decompose the multi-task cooperative offloading optimization problem with dynamic resource coordination and intelligent scheduling mechanism into two sub-optimization goals: differentiated task allocation and zero-interrupt routing path planning; S21. The task offloading optimization objective is expressed as maximizing the average number of completed tasks. The total number of tasks generated in each time slot t is represented by N, and the task completion rate is: ;(16) The optimization objective of the task offloading problem is expressed as: ;(17); Indicates the number of time slots in the statistics; binary variable To indicate the task offloading decision; Represents the generated task; S22, decomposing the optimization objective of the task offloading problem into a task and service equipment matching problem and a path planning problem under the constraints of path connectivity and stability; S221. For the matching problem between tasks and service devices, the total execution time of tasks in time slot t is defined as the longest execution time among all tasks, and express: ;(18) in Represents the task allocation matrix; because The formula consists of two parts: calculation execution time and transmission time. It is impossible to obtain the accurate routing transmission time under the current time scale. To indicate the starting node With the target node Estimated transfer time between: ;(19) in is the median value of the transmission rate; represents the task size of task k; Representation Node The communication radius; Representation Node The distance between Therefore It is expressed as: ;(20) The optimization objective of the sub-problem is: ;(21); S222. For the path planning problem under the constraints of path connectivity and stability, N task allocation schemes are obtained through S221. The successful implementation of each allocation scheme will obtain a corresponding reward Rk, thereby expressing the optimization problem as a reward maximization problem: ;(22) in, represents the routing path selection of task k; S3. Based on the hierarchical reinforcement learning algorithm, the process of finding a zero-interruption route in the dynamic road topology is modeled as a partially observable Markov process, and the high-level policy network and the low-level policy network and their corresponding action space, state space and reward function are established hierarchically.
2. A method for task "0 interruption" offloading based on hierarchical reinforcement learning according to claim 1, characterized in that: S11 content is as follows: S111. Construct a time delay model for the task. Generated tasks The total time of its uninstallation execution is The total time is represented by the transmission time and calculate execution time composition: ;(3) in Indicates vehicle The generated task k is determined by j Each node device performs processing; S112, any task Execution time The calculation formula is constructed as: ;(4) in Representation Node j The computing power of the device; Indicates the computational complexity of task k; binary variable To indicate the task offloading decision; if the task is decided to be processed locally, ,but ,otherwise ; Represents a road node set; S113. Task transmission time , when the vehicle The generated k-type tasks determine the local vehicle When performing calculations, its task transfer time ; When offloading tasks to another service node j When performing task calculation, the task transmission time is: ;(5) in, represents the task size of task k; Indicates the size of data returned after task k is completed; Representation Node The downlink transmission rate between.
3. A method for task "0 interruption" offloading based on hierarchical reinforcement learning according to claim 2, characterized in that: S12 content is as follows: S121. When two vehicles are traveling in the same direction, When the two vehicles Link time between Calculated as: ;(6) in, Representation Node The communication radius; Indicates the node at time t ij The distance between when When the two vehicles Connection time between Calculated as: ;(7); S122. When two vehicles are traveling in opposite directions, Connection time between Calculated as: ;(8) S123, Vehicle Connection time with roadside unit RSU-j Calculated as: ;(9) in, Represents the horizontal and vertical coordinates of the roadside unit RSU-j; Indicates vehicle The horizontal and vertical coordinates at time t.
4. The method for unloading tasks with "0 interruption" based on hierarchical reinforcement learning according to claim 3, characterized in that: S13 content is as follows: S131, vehicle standard speed deviation evaluation parameters , which is evaluated by the difference between the vehicle speed and the road reference speed, using Represents the prescribed speed limit of the road, and the speed deviation is calculated as: ; (10) When the vehicle speed and When the speed difference from the prescribed speed limit is large, The larger the vehicle is, the more unstable the vehicle's driving will be. Perform normalization: ;(11); S132. Connection time evaluation parameters The calculation is based on the calculated connection time : ; (12) because It is an inverse proportional function, which is normalized using nonlinear function transformation: ; (13); S133: Since the speed relationship between vehicles has an impact on link stability, additional parameters are introduced Measure the relative difference between the speeds of two cars: ; (14); S134. For any multi-hop path , and use Indicates the path The total stability entropy value of is calculated using an exponential growth strategy. Multi-hop path with relay nodes The total stability entropy : ; (15) in, Representative path Middle p The stability entropy value calculated by the segment link, is the stability index, .
5. A method for task "0 interruption" offloading based on hierarchical reinforcement learning according to claim 4, characterized in that: The S3 content is as follows: S31. Model the path selection problem in a dynamic environment as a POMDP, which includes the observation O, state S, action A, and reward R of the environment; and use the deep reinforcement learning DRL algorithm to solve the optimization challenge. To solve the dynamic topology problem, a cluster-assisted hierarchical reinforcement learning CAHRL algorithm is designed; S32. At time t, the road cloud platform collects road information within the observed range, including vehicles V and roadside units RSU, where vehicle observation information The definition is as follows: ; (23) in, Represents the vehicle index list within the current road segment; Represents a list of vehicle speed information; and Represents the location information list of each vehicle; RSU observation information The definition is as follows: ;(24) in Represents the processing queue list of RSU; and Represents the location information list of RSU; Represents the index list of RSU; S33, the length is The road is gridded into several blocks to characterize the road conditions. Due to the hierarchical decision-making design framework, the state information based on high-level decisions and low-level decisions is different. is defined as follows: ;(25) in, represents the task type of task k; Indicates the current node information; Indicates the target node information; Indicates the number of nodes that can be selected in the current grid block n; Indicates the current hop count; Low level status is defined as follows: ;(26) in, Indicates node speed information; and Indicates node location information; S34. The design of the action space is divided into two layers, which are defined as follows: High-level actions Select n block options: ;(27) in, Indicates an optional block within the road; Low-level actions Select m nodes in the selected block: ;(28) in, Indicates optional nodes in the block; S35. In the reward part design, the reward formula for low-level decision making is defined as: ;(29) ;(30) in is the reward discount factor; Represents the reward for step n; Indicates reward for success; Indicates the basic step reward; The reward formula for the high-level strategy is defined as: ;(31) in, Characterization Node Whether communication is possible between them.
Citation Information
Patent Citations
Vehicle infrastructure cooperation online task scheduling method and system based on calculation result reuse
CN117255370A
V2x services for providing journey-specific QOS predictions
US20230074288A1