RSU-assisted car networking task unloading method

By using EKF and PPO algorithms in the Internet of Vehicles system, combining GPS information to predict vehicle trajectory and optimizing task offload decisions, the problem of reliability and delay in the Internet of Vehicles task offload in a dynamic environment is solved, and more efficient task processing and resource utilization are achieved.

CN120128988APending Publication Date: 2025-06-10ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510333553.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to ensure low latency and reliability of vehicle network task offloading in dynamic environments, mainly due to the inadequate consideration of vehicle mobility and GPS errors.

Method used

The extended Kalman filter (EKF) is used to combine GPS information to obtain high-precision vehicle trajectory information, and optimize task offload decisions through the deep reinforcement learning algorithm PPO, and select the optimal computation offload node to minimize task processing delay.

Benefits of technology

It significantly improves the accuracy of vehicle position prediction, enhances the reliability of task offload decisions, reduces task processing delays, and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128988A_ABST
    Figure CN120128988A_ABST
Patent Text Reader

Abstract

The invention discloses an RSU-assisted Internet of Vehicles task unloading method, and belongs to the technical field of edge computing. According to the invention, the real-time track information of the vehicle is obtained by combining the EKF with the GPS information of the vehicle, so that the effective vehicle set in the communication range of each RSU and the hop count set of the task result returned by each RSU are updated in real time; task unloading time delays of different task unloading nodes are calculated according to the time delay calculation model, and an optimization objective function is designed according to the calculated task unloading time delays; according to the objective function, converting a task decision problem into a Markov decision process problem, and selecting a calculation unloading node which enables task processing time delay to be minimum; and optimizing the unloading decision model by using a PPO algorithm, and generating an optimal task unloading decision. According to the method, the mobility influence of the vehicle is fully considered, and the accuracy of vehicle position prediction in a dynamic environment is remarkably improved, so that the reliability of task unloading decision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of edge computing, and specifically relates to an RSU-assisted vehicle networking task offloading and dynamic vehicle trajectory prediction method based on deep reinforcement learning, which is used to solve the problem of optimizing vehicle networking task offloading in a dynamic environment. Background Art

[0002] In recent years, the emergence of 5G / 6G technologies has greatly emphasized the importance of autonomous driving in the road traffic system, where RSU (Road Side Unit) plays a key role. As a fixed infrastructure, RSU can not only achieve real-time communication with vehicles, but also provide certain computing and storage capabilities.

[0003] Since the tasks generated by RSU usually have high latency requirements and the computing resources of RSU are limited, it is difficult for RSU to process all the generated computing tasks. RSU will offload the analysis tasks to other network nodes, including other RSUs and vehicles. However, in the prior art, various offloading schemes usually only use GPS to measure the mobility of vehicles, without considering the impact of GPS errors. Due to the frequent change of vehicle positions, the vehicle group within the communication range of RSU changes dynamically. The GPS positioning error further exacerbates this problem, reducing the positioning accuracy and affecting the accuracy of the association between RSU and vehicles. Therefore, how to ensure low latency and reliable task offloading in a dynamic and uncertain environment remains a major challenge. Summary of the Invention

[0004] Aiming at the problem that the reliability of the vehicle networking task offloading scheme in the prior art needs to be further improved due to the insufficient consideration of vehicle mobility, the present invention provides an RSU-assisted vehicle networking task offloading method. The present invention applies EKF to fuse the measurement trajectory based on GPS, fully considering the impact of vehicle mobility, thus significantly improving the accuracy of vehicle position prediction in a dynamic environment, and thus being beneficial to ensuring the reliability of task offloading decisions.

[0005] To achieve the above object, the technical solution provided by the present invention is as follows:

[0006] The present invention provides an RSU-assisted vehicle networking task offloading method, including:

[0007] Actively sense road information through sensors on RSU to generate sensing tasks;

[0008] Obtain the real-time trajectory information of vehicles by combining EKF with vehicle GPS information to update the effective vehicle set within the communication range of each RSU and the hop count set for each RSU to return task results in real time;

[0009] Calculate the task offloading latency of different task offloading nodes according to the latency calculation model, and design an optimization objective function based on the calculated task offloading latency to minimize the task processing latency;

[0010] According to the objective function, transform the task decision problem into a Markov decision process problem, and select the computing offloading node that minimizes the task processing latency;

[0011] Use the PPO algorithm to optimize the offloading decision model and generate the optimal task offloading decision.

[0012] Furthermore, the effective vehicle set within the communication range of each RSU and the hop count set for each RSU to return the task result are updated in real time, specifically including:

[0013] (1) The vehicle sends the historical EKF trajectory and state to the source RSU via a beacon message, denoted here as RSU r;

[0014] (2) RSU r divides the historical EKF trajectory of vehicle v into η segments, each corresponding to the movement trajectory of vehicle v within the equivalent transmission time, and these segments are denoted as Therefore, the position change of vehicle v after receiving the task is expressed as:

[0015]

[0016] where w i ∈[0,1] is the weight assigned to each historical segment;

[0017] (3) The position of vehicle v after receiving the task is determined as (x′ v (t), y′ v (t)). If vehicle v is within the communication range of RSU r, RSU r adds vehicle v to the effective vehicle set N v (t);

[0018] (4) After RSU r obtains the corresponding effective vehicle set, it calculates the time for each effective vehicle to complete the task processing according to the beacon information. RSU r predicts its position (x″ v (t), y″ v (t)) for effective vehicle v using the weighted average method based on the historical EKF trajectory of effective vehicle v;

[0019] (5) RSU r identifies the RSUs passed by effective vehicle v according to its predicted trajectory. Next, effective vehicle v transmits the task processing result to RSU via multi-hop communication through multiple RSUs, and the number of hops required to transmit the processing result is denoted as h r,v (t);

[0020] (6) Finally, RSU r calculates the number of hops required to return the task results of all valid vehicles and forms these hop arrays into a set H r (t).

[0021] Furthermore, the task offloading nodes include valid vehicles and neighboring RSUs. When an RSU generates a sensing task, the processing of the sensing task can only be offloaded to a certain task offloading node for calculation.

[0022] Furthermore, the optimization objective function of the delay calculation model is as follows:

[0023]

[0024] where T r (t) is the total processing delay for the task generated by RSU r to be offloaded to a valid vehicle for calculation and to be calculated on a neighboring RSU;

[0025] x r,β (t) is used as the offloading decision variable, N r (t) represents the set of valid vehicles, M r (t) represents the set of neighboring RSUs, x r,β (t) = 1, β ∈ N r (t) indicates that the task generated by RSU r is offloaded to a valid vehicle; x r,β (t) = 1, β ∈ M r (t) indicates that the task is offloaded to a neighboring RSU;

[0026] In constraint C3, T r,β (t) is defined as the time required to process each task. If β ∈ N r (t), then T r,β (t) is the total processing delay for the task offloaded to vehicle v If β ∈ M r (t), then T r,β (t) is the total processing delay for the task offloaded to RSU k τ r is the time delay constraint for task processing.

[0027] Furthermore, the total processing delay for the task offloaded to a neighboring RSU or a valid vehicle includes transmission delay, waiting delay, calculation delay, and result transmission delay.

[0028] Furthermore, the Markov decision process includes three key elements: state, action, and reward, where:

[0029] The system state s(t) includes the task queue information D(t), the task information f(t) of each computing node, and the number of hops H(t) during the return process of the task processing completed by the vehicle nodes that can communicate with each RSU;

[0030] The action is that after observing the state s(t) from the environment, the agent makes an offloading decision:

[0031] a t ={x 1,β (t), x 2,β (t), …, x r,β (t), …, x R,β (t)}

[0032] where β ∈ (N r (t) ∪ M r (t));

[0033] The reward function is expressed as follows:

[0034]

[0035] Furthermore, obtaining the real-time trajectory information of the vehicle by combining the vehicle GPS information through EKF specifically includes:

[0036] (1) In the prediction stage of EKF, according to the optimal state estimate of vehicle v at time slot t - 1 and the non-linear motion model of vehicle v, calculate the predicted state of vehicle v at time slot t Expressed as:

[0037]

[0038] where, represents the state vector of vehicle v at time slot t - 1, which consists of where x v (t - 1), y v (t - 1) respectively represent the horizontal and vertical coordinates of the vehicle, respectively represent the lateral and longitudinal speeds of the vehicle; where is the control vector, and respectively represent the accelerations of vehicle v in the X-axis and Y-axis directions; J A and J B are the Jacobian matrices obtained by differentiating and u v (t - 1) respectively;

[0039] (2) Evaluate the predicted state The uncertainty, based on the optimal state estimation error covariance matrix S of vehicle v at time slot t-1 v (t-1), to obtain the predicted error covariance matrix of vehicle v at time slot t, denoted as Expressed as follows:

[0040]

[0041] where Q(t-1) represents the covariance of the process noise.

[0042] (3) In the update phase, obtain the measured state of vehicle v during time period t through vehicle sensors and the measurement error covariance matrix O v (t), combine the measurement error covariance matrix O v (t) and the predicted error covariance matrix to calculate the Kalman gain K v (t), expressed as:

[0043]

[0044] (4) Use the Kalman gain to fuse the predicted state and the measured state to obtain a more accurate optimal estimated state, denoted as Expressed as:

[0045]

[0046] where J C serves as the Jacobian matrix of the measured state;

[0047] (5) Finally, use the predicted error covariance matrix to update the optimal estimation error covariance matrix of vehicle v during time period t, denoted as S v (t), expressed as:

[0048]

[0049] Furthermore, the PPO algorithm employs two neural networks: a policy network and a value network. Among them, the policy network generates action probabilities for a given state, while the value network estimates the expected value of each action.

[0050] Furthermore, the parameters of the policy network are denoted as θ π , and the parameters θ π are updated during training. The policy network is updated by maximizing the objective function in L π (θ π ) as follows:

[0051]

[0052] where represents the probability ratio between the new policy and the old policy; ∈ is a hyperparameter that controls the allowable range of this difference; Φ represents the advantage function, which is used to measure the state s t and the action a t relative to the average value of all actions.

[0053] Furthermore, the loss function of the value network represents the mean square error between the predicted value and the reward, and the specific expression is as follows:

[0054]

[0055] is the parameter of the value network, represents the transition state s t as the output of the value network; is the reward value.

[0056] Adopting the technical solution provided by the present invention, compared with the prior art, the following beneficial effects can be achieved:

[0057] (1) By using EKF in combination with vehicle GPS information, the present invention obtains high-precision vehicle trajectory information, and uses this to update the effective vehicle set within the communication range of each RSU and the hop count set of the task results returned by each RSU in real time, thereby effectively increasing the accuracy of vehicle association within the RSU communication range, reducing errors, avoiding the impact of vehicle movement on the reliability of task offloading decisions, and improving the reliability of task offloading.

[0058] (2) The present invention selects a suitable computing offloading node according to the task processing delay after task offloading. When choosing to offload the task to the vehicle for calculation, due to the mobility of the vehicle, there is a risk of communication link disconnection during the task transmission process. The present invention predicts the future trajectory of the vehicle by combining the task transmission time of the vehicle in a weighted average manner based on the vehicle's historical EKF trajectory, so as to help the RSU identify the vehicles still within the communication range and avoid communication link disconnection.

[0059] (3) Due to the complex road information, optimizing the task processing delay is a non-convex problem that is difficult to solve. The present invention can effectively explore in the policy space in deep reinforcement learning, so as to find a better solution and minimize the task processing delay. At the same time, the present invention transforms the problem into a Markov decision process and applies the proximal policy optimization (PPO) algorithm in the deep reinforcement learning algorithm, which is beneficial to further solve the non-convex problem well. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is the system scenario diagram of the embodiment of the present invention;

[0061] Figure 2It is a schematic flowchart of the task offloading method according to an embodiment of the present invention;

[0062] Figure 3 It is the pseudo code of the PPO algorithm in an embodiment of the present invention;

[0063] Figure 4 It is a comparison chart of reward values and task processing delays under different algorithms. Detailed implementation manners

[0064] Considering that the system operates in a time-slot-based manner, the present invention divides time into multiple discrete time slots, referred to as the set {1, 2, …, t, …}. As Figure 1 shown, the scenario of the present invention includes multiple RSUs and multiple vehicles. Taking RSU 1 as an example, it monitors its coverage area and generates tasks. Vehicles within the communication range apply EKF to fuse GPS trajectories and obtain accurate position information. They regularly send beacon messages to RSU 1, which contain information such as their trajectories, current positions, speeds, processing capabilities, and task queues. After receiving these messages, RSU 1 analyzes whether the vehicles are within its communication range and evaluates the task processing capabilities of the vehicles. It identifies valid vehicles, such as Vehicle 1 and Vehicle 2. Next, RSU 1 predicts the positions of Vehicle 1 and Vehicle 2 after completing the tasks based on the historical EKF trajectories. Taking Vehicle 1 as an example, RSU1 finds that Vehicle 1 will be out of its communication range and enter the communication range of RSU2 after completing the task, and multi-hop transmission is required to send the task result back to RSU1. RSU1 plans the transmission path according to the predicted trajectory and routes the task result to the adjacent RSU2. The number of hops for Vehicle 1 to transmit the task back is recorded as 1. In addition, RSU1 uses the PPO algorithm to optimize the task offloading decision and selects vehicles with low latency and low energy consumption, thereby improving the task processing efficiency and resource utilization rate.

[0065] To further understand the content of the present invention, the present invention will be described in detail below in combination with specific embodiments.

[0066] Refer to Figure 2 , the RSU-assisted vehicle-to-everything (V2X) task offloading method according to an embodiment of the present invention specifically includes the following steps:

[0067] Step 1: According to the road environment, sensors on the RSU actively sense road information and generate sensing tasks.

[0068] Step 2: The vehicle combines the vehicle GPS information through an EKF (Extended Kalman Filter) to obtain the real-time trajectory information of the vehicle.

[0069] The position of the vehicle changes frequently, resulting in a dynamic change in the set of vehicles within the communication range of the RSU. The GPS positioning error further exacerbates this problem, reducing the positioning accuracy and affecting the accuracy of the association between the RSU and the vehicle. How to ensure low latency, energy consumption, and reliable task offloading in a dynamic and uncertain environment is of great significance.

[0070] In the embodiment of the present invention, the EKF is used to fuse the predicted state and the measured state of the data, and an optimal estimated state with higher accuracy can be obtained. Specifically:

[0071] For the EKF in the prediction stage, according to the optimal state estimate of vehicle v at time slot t - 1 and the non - linear motion model of vehicle v, the predicted state of vehicle v at time slot t is calculated Expressed as:

[0072]

[0073] Where, represents the state vector of vehicle v at time slot t - 1, which consists of where x v (t - 1), y v (t - 1) represent the horizontal and vertical coordinates of the vehicle respectively, represent the lateral and longitudinal speeds of the vehicle respectively. All subsequent states represent speed and position information, where is the control vector, and represent the accelerations of vehicle v in the X - axis and Y - axis directions respectively. Here, the non - linear motion model of the vehicle is expressed as:

[0074]

[0075] Here, the model is differentiated with respect to and u v (t - 1) respectively to obtain the Jacobian matrices J A and J B , and the non - linear model is linearized through the Jacobian matrix, thus ensuring the accuracy and efficiency of the calculation.

[0076] Subsequently, the uncertainty of the predicted state v (t - 1) is evaluated using the corresponding method according to the optimal state estimation error covariance matrix S of vehicle v at time slot t - 1 , and the predicted error covariance matrix of vehicle v at time slot t is obtained, denoted as Expressed as:

[0077]

[0078] Q v (t - 1) represents the covariance of the process noise; the inverse matrix of the error covariance matrix represents the precision matrix.

[0079] In the update stage, the measured state of vehicle v at time period t is obtained through vehicle sensors and the measurement error covariance matrix O v (t), combining the measurement error covariance matrix O v (t) and the predicted error covariance matrix the Kalman gain K v (t) can be calculated as:

[0080]

[0081] Using the Kalman gain to fuse the predicted state and the measured state to obtain a more accurate optimal estimated state, denoted as expressed as:

[0082]

[0083] Here as the measured state, it is the data directly obtained through the vehicle's GPS, J C The Jacobian matrix of the measured state is used to keep the dimensions of the predicted state and the measured state the same.

[0084] Finally, use the predicted error covariance matrix to update the optimal estimated error covariance matrix of vehicle v at time period t, denoted as S v (t) is expressed as:

[0085]

[0086] The optimal estimated state and the optimal estimated error covariance matrix will be used for the next time period.

[0087] Step 3: According to the historical EKF trajectory, update the set of valid vehicles within the communication range of each RSU in real time, and at the same time update the set of hop counts for each RSU to return the task results.

[0088] In the traditional scheme, either the impact of vehicle mobility is not considered, or only the mobility challenges under basic or predefined mobility models in the vehicle network are solved. On the one hand, considering that during the process of task transmission to the vehicle, the vehicle may leave the communication range of the RSU, resulting in task transmission failure, so it is necessary to identify which vehicles are still within the communication range of the RSU during task transmission; on the other hand, after the task processing is completed, the vehicle may no longer be within the communication range of the source RSU, so it is necessary to estimate the vehicle position in order to send the task result back to the RSU through multi-hop transmission by other RSUs.

[0089] Specifically, the update process of the set of valid vehicles within the communication range of each RSU and the set of hop counts for each RSU to return task results in the embodiments of the present invention includes:

[0090] (1) After the vehicle obtains the accurate position of the vehicle by fusing the GPS trajectory using the EKF model, it sends the historical EKF trajectory and status to the source RSU (marked as RSU r here) through the beacon message;

[0091] During the time period t, the source RSU obtains a set of valid vehicles using the beacon information; when the RSU decides to offload the computing task to the vehicle, it will select the vehicle within its communication range as the offloading node. Since the vehicle is mobile, it is crucial to ensure that the vehicle is always within the communication range of the RSU during the entire task transmission process. Each vehicle periodically sends a beacon message to inform the RSU of its status, including the historical EKF trajectory, the current position and speed obtained through EKF, the processing capacity, and the task queue length.

[0092] (2) RSU r divides the historical EKF trajectory corresponding to when the vehicle v accepts the task into η segments, each segment corresponding to the movement trajectory of the vehicle v within the equivalent transmission time, including n time slot segments, and these segments are represented as represents the two-dimensional displacement vector of the vehicle v within the η-th equivalent transmission time. Therefore, the position change of the vehicle v after receiving the task is represented as:

[0093]

[0094] where w i ∈ [0, 1] is the weight assigned to each historical segment, and the weight is determined based on two key factors. The trajectory segment closest to the position where the task is accepted is assigned a higher weight to give priority to the most recent movement, which can better illustrate the future position. At the same time, the weight is also proportional to the accuracy of the EKF trajectory of each historical segment. Finally, the weight is expressed as:

[0095]

[0096] a i represents the accuracy of the EKF trajectory of the i-th segment, expressed as:

[0097]

[0098] Here represents the error covariance matrix of the optimal estimated state of the EKF, ensuring that the segments with high accuracy contribute more to the prediction.

[0099] (3) The position of the vehicle v after accepting the task is determined as (x′ v (t), y′v (t)), if vehicle v is within the communication range of RSU r, RSU r adds vehicle v to the set of valid vehicles N v (t).

[0100] (4) After RSU r obtains the corresponding set of valid vehicles, it calculates the time for each valid vehicle to complete task processing based on the beacon information. RSU r predicts the position (x″ v (t), y″ v (t)) of the valid vehicle v using the weighted average method according to the historical EKF trajectory of the valid vehicle v

[0101] Here, the method of predicting the position (x″ v (t), y″ v (t)) of the valid vehicle v using the weighted average method is the same as the method for calculating the position change of vehicle v after receiving the task in step (2).

[0102] (5) RSU r identifies the RSUs passed by the valid vehicle v according to its predicted trajectory. Next, the valid vehicle v transmits the task processing result to the RSU in a multi-hop communication manner through multiple RSUs. The number of hops required to transmit the processing result is denoted as h r,v (t).

[0103] (6) Finally, RSU r calculates the number of hops required to return the task results of all valid vehicles and forms these hop numbers into a set H r (t).

[0104] Step Four: Calculate the corresponding task offloading delay according to different task offloading nodes, and then design an optimization objective function based on the calculated task offloading delay to minimize the task processing delay.

[0105] Specifically, when the RSU generates a sensing task, the processing of the sensing task can only be offloaded to a certain computing offloading node for calculation, and the corresponding task processing delay is calculated according to different computing offloading nodes. The offloading delay model is the total processing delay of the task calculated on the valid vehicle and on the neighboring RSU.

[0106] According to Shannon's theorem, the data transmission rate between RSU r and vehicle v at time slot t is The data transmission rate between vehicle v and RSU r is The total processing delay for the source RSU (RSU r) to offload the task to the neighboring RSU includes transmission delay, waiting delay, computing delay, and result transmission delay.

[0107] Among them, the time required for the task f(t) to be transmitted from RSU r to RSU k is expressed as:

[0108]

[0109] Since the RSU positions are fixed, for convenience, the data transfer rate between adjacent RSUs is set to a constant value, denoted as δ r Denoted as the size of the task number, the above formula means that if the processing node is not the source RSU, there is a transmission delay; otherwise,

[0110] The waiting delay for processing a task at RSU k is denoted as:

[0111]

[0112] ζ q Denoted as the number of CPU cycles required to process a unit of data, that is, the processing intensity of the task; δ q is the data size of the task; f k is the processing capacity of RSU k, the task list (t, k) is the task queue of RSU k in time slot t, and the task processing delay is denoted as:

[0113]

[0114] ζ r Denoted as the number of CPU cycles required to process a unit of data, that is, the processing intensity of the task; δ r The data size of the task. The delay for returning the processing result from RSU r to RSU k is denoted as:

[0115]

[0116] δ′ r Denoted as the data size of the task result.

[0117] Finally, the total processing delay for offloading a task to RSU k The calculation formula is as follows:

[0118]

[0119] When the source RSU (RSU r) generates a task and calculates it on an active vehicle:

[0120] The transmission delay for task f(t) to be transmitted from RSU r to vehicle v is denoted as:

[0121]

[0122] The waiting delay for processing a task at vehicle v is denoted as:

[0123]

[0124] fv Let the processing capacity of vehicle v be \(P(v)\), and the task list \((t, v)\) be the task queue of vehicle v in time slot t. The task processing delay is expressed as:

[0125]

[0126] The delay in returning the processing result from RSU r to vehicle v is expressed as:

[0127]

[0128] Finally, the total processing delay of task offloading to vehicle v The calculation formula is as follows:

[0129]

[0130] According to the processing costs of different computing offloading nodes, in summary, the total processing delays for tasks to be computed on effective vehicles and on neighboring RSUs are:

[0131]

[0132] where \(V = \{1, 2, \ldots, v, \ldots, V\}\) is the set of vehicles, is the set of RSUs, \(N\) r (t) represents the set of effective vehicles, and \(M\) r (t) represents the set of neighboring RSUs.

[0133] The objective of the present invention is to continuously optimize the offloading decision to minimize the total processing delay of the tasks perceived by each RSU. The optimization objective function is as follows:

[0134]

[0135] The embodiment of the present invention defines \(x\) r,β (t) as the offloading decision variable, where \(\beta\in(N\) r (t)\(\cup M\) r (t)). Among them, when \(x\) r,β (t)=1 and \(\beta\in N\) r (t), it means that the task generated by RSU r is offloaded to an effective vehicle; when \(x\) r,β (t)=1 and \(\beta\in M\) r (t), it means that the task is offloaded to a neighboring RSU.

[0136] In constraint C3, \(T\) r,β (t) is defined as the time required to process each task. If \(\beta\in N\) r (t), then If \(\beta\in M\) r (t), then Constraints C1 and C2 ensure that each task is offloaded and processed on only one node. Constraint C3 guarantees that the task is completed within the specified time constraint. Here, C1 incorporates the previous considerations regarding vehicle mobility to ensure that the vehicle does not leave the communication range during the RSU task transmission, resulting in task failure.

[0137] Step 5: According to the objective function, transform the task decision problem into a Markov decision process problem, and select the computing offloading node that minimizes the task processing delay.

[0138] The Markov decision process consists of three key elements: state, action, and reward, which are defined as follows:

[0139] State:

[0140] The state space includes the task queue information D(t) of each computing node, the task information and the number of hops H(t) during the return process of task processing completion of the vehicle nodes that can communicate with each RSU. In each time slot t, the agent monitors the network environment and dynamically collects the system state:

[0141]

[0142] Action:

[0143] After observing the state s(t) from the environment, the agent makes an offloading decision:

[0144] a t ={x 1,β (t),x 2,β (t),…,x r,β (t),…,x R,β (t)}

[0145] where β ∈ (N r (t) ∪ M r (t)).

[0146] Reward:

[0147] The design of the reward function is the key to reinforcement learning. Generally speaking, the reward function is related to the established objective function. The optimization goal is to obtain the maximum system reward. If constraints C1 - C3 are satisfied, the reward function is the negative value of the objective function. If constraints C1 - C3 are not satisfied, the reward function is -Z. The reward function is expressed as:

[0148]

[0149] Step 6: Use the PPO algorithm to solve the co - optimization problem of task processing delay and obtain the optimal task decision.

[0150] The PPO algorithm uses two neural networks: a policy network and a value network. The policy network generates action probabilities for a given state, while the value network estimates the expected value of each action. The parameters of the policy network are denoted as θ π , and the parameters of the value network are denoted as In time slot t, the agent observes the current state s from the environment t , which serves as the input to the policy network; the agent selects an action a t in state s t with probability denoted as π(a t |s t ; θ π ). After executing action a t in state s t , the agent receives a reward and transitions to a new state s t+1 . The tuple is stored in the experience replay buffer. The schematic diagram of the PPO algorithm's pseudocode is specifically as Figure 3 shown. The pseudocode for selecting offloading decisions is presented in Algorithm 1. Consider the RSU as an agent that makes tasks as centralized training and distributed decisions. Lines 3 and 4 are used to initialize the network parameters. From lines 5 to 14, data is collected and stored in the replay buffer. During this process, the valid nodes of the source RSU are determined, and the available offloading nodes are determined through action masks. From line 16 to line 23, it is used to update the network parameters. After sampling a batch of data from the replay buffer equation. Subsequently, the parameters of the policy network are updated and the parameters of the value network are updated.

[0151] To continuously improve the policy network, the parameters θ π are updated during training. The policy network is updated by maximizing the objective function in L π (θ π ):

[0152]

[0153] where represents the probability ratio between the new policy and the old policy, ∈ is a hyperparameter that controls the allowed range of this difference, Φ represents the advantage function, which measures the value of action a t under state s t relative to the average value of all actions.

[0154] Meanwhile, the loss function of the value network is represented as the mean squared error between the predicted value and the reward, and this loss is used to evaluate the accuracy of the value function network in estimating the state value; represents the transition state s t as the output of the value network.

[0155]

[0156] Figure 4 They are respectively the comparison charts of the reward values and task processing delays under different algorithms. It can be seen from the charts that the present invention realizes greater reward values and lower task processing delays. Figure 4 As shown in the left sub-chart, the present invention is superior to the method that does not use EKF (only uses GPS for vehicle trajectory positioning) and the method of non-predictive offloading NTPO (does not obtain valid vehicles and hop counts) in terms of latency. Different from GNE, the present invention takes into account the uncertainty of GPS positioning. Through data fusion in EKF, more accurate position information can be obtained, so that vehicle nodes within the communication range of the RSU can be identified. Non-predictive offloading is worse than not using EKF, mainly because of vehicle movement, making it difficult to reliably receive the calculation results. The present invention applies trajectory prediction to improve the reliability of data transmission, thereby improving the task completion rate and overall performance.

[0157] Figure 4 The right sub-chart illustrates the reward differences between different schemes. The present invention regards the tasks that exceed the delay constraint as unsuccessful processing tasks and assigns specific penalties to them. Methods with better latency performance will generate higher reward values. Therefore, compared with non-predictive offloading and not using EKF, the method of the present invention can obtain higher reward values. In addition, for non-predictive offloading, if the vehicle does not return the task result before leaving the communication range of the source RSU, it is classified as a task failure, resulting in the worst performance of non-predictive offloading.

[0158] In summary, for the RSU-assisted single-hop offloading and dynamic vehicle prediction scheme based on deep reinforcement learning of the present invention, the key innovation is the integration of EKF and PPO. Specifically, first, EKF is applied to fuse the GPS-based measurement trajectories, thus significantly improving the accuracy of vehicle position prediction in a dynamic environment and then considering the impact of vehicle mobility. This prediction ability is crucial for reliable task offloading decisions. Then, the task offloading problem is converted into an MDP in a unique way, and the DRL algorithm PPO is used to dynamically optimize the offloading strategy under mobile conditions. The synergy between trajectory prediction and the adaptive decision-making framework enables the present invention to achieve excellent performance in terms of latency reduction and energy efficiency, solving the limitations of existing methods that either ignore mobile dynamics or rely on rigid optimization strategies.

Claims

1. A RSU-assisted vehicle networking task offloading method, characterized in that: include: Actively sense road information through sensors on the RSU and generate perception tasks; The real-time trajectory information of the vehicle is obtained by combining the vehicle GPS information with the EKF to update the valid vehicle set within the communication range of each RSU and the hop count set of each RSU returning the task result in real time; The task offloading delay of different task offloading nodes is calculated according to the delay calculation model, and the optimization objective function is designed according to the calculated task offloading delay to minimize the task processing delay; According to the objective function, the task decision problem is transformed into a Markov decision process problem, and the computation offloading node that minimizes the task processing delay is selected; The PPO algorithm is used to optimize the offloading decision model and generate the optimal task offloading decision.

2. The RSU-assisted vehicle networking task offloading method according to claim 1 is characterized in that: The valid vehicle set within the communication range of each RSU and the hop count set of each RSU returning the task result are updated in real time, including: (1) The vehicle sends the historical EKF trajectory and state to the source RSU through a beacon message, which is denoted as RSU r. (2) RSUr divides the historical EKF trajectory of vehicle v into η segments, each of which corresponds to the motion trajectory of vehicle v in the equivalent transmission time. These segments are represented as Therefore, the position change of vehicle v after receiving the task is expressed as: where w i ∈[0,1] is the weight assigned to each history segment; (3) The position of vehicle v after accepting the task is determined as (x′ v (t),y′ v (t)), if vehicle v is within the communication range of RSU r, RSU r adds vehicle v to the valid vehicle set N v (t); (4) After RSU r obtains the corresponding valid vehicle set, it calculates the time it takes for each valid vehicle to complete the task processing based on the beacon information. RSU r uses the weighted average method to predict the position (x″) of the valid vehicle v based on its historical EKF trajectory. v (t),y″ v (t)); (5) RSU r identifies the RSUs that the valid vehicle v passes through based on its predicted trajectory. Next, the valid vehicle v transmits the task processing results to the RSUs through multiple RSUs in a multi-hop communication manner. The number of hops required to transmit the processing results is represented by h r,v (t); (6) Finally, RSUr calculates the number of hops required to return the task results of all valid vehicles and groups these hops into a set H r (t).

3. The RSU-assisted vehicle networking task offloading method according to claim 2 is characterized in that: The task offloading nodes include valid vehicles and adjacent RSUs, and when the RSU generates a perception task, the processing of the perception task can only be offloaded to a certain task offloading node for calculation.

4. The RSU-assisted vehicle networking task offloading method according to claim 3 is characterized in that: The optimization objective function of the delay calculation model is: s.t. C1:x r,β (t)={0,1}, β∈(N r (t)∪M r (t)) C2: C3:x r,β (t)T r,β (t)≤τ r Among them, T r (t) The total processing delay of the tasks generated for RSU r to be offloaded to the active vehicles and calculated on the neighboring RSUs; x r,β (t) is used as the unloading decision variable, N r (t) represents the valid vehicle set, M r (t) is represented as the set of neighboring RSUs, x r,β (t)=1,β∈N r (t) indicates that the task generated by RSU r is offloaded to the valid vehicle; x r,β (t)=1,β∈M r (t) indicates that the task is offloaded to the neighboring RSU; In constraint C3, define T r,β (t) is the time required to process each task, if β∈N r (t), then T r,β (t) is the total processing delay of task offloading to vehicle v If β∈M r (t), then T r,β (t) is the total processing delay of task offloading to RSU k τ r The time delay constraint for task processing.

5. The RSU-assisted Internet of Vehicles task offloading method according to claim 4 is characterized in that: The total processing delay of task offloading to the neighboring RSU or valid vehicle includes transmission delay, waiting delay, calculation delay and result transmission delay.

6. The RSU-assisted vehicle networking task offloading method according to any one of claims 1 to 5, characterized in that: The Markov decision process includes three key elements: state, action, and reward, among which: The system state s(t) includes the task queue information D(t) of each computing node, the task information f(t) and the number of hops H(t) during the return process of the task processing completed by each vehicle node that can communicate with the RSU; The action is that after observing the state s(t) from the environment, the agent makes an unloading decision: a t ={x 1,β (t),x 2,β (t),…,x r,β (t),…,x R,β (t)} Where β∈(N r (t)∪M r (t)); The reward function is expressed as follows:

7. The RSU-assisted vehicle networking task offloading method according to any one of claims 1 to 5, characterized in that: The real-time trajectory information of the vehicle is obtained by combining the EKF with the vehicle GPS information, specifically including: (1) In the prediction stage, EKF estimates the optimal state of vehicle v at time slot t-1. and the nonlinear motion model of vehicle v, and calculate the predicted state of vehicle v at time slot t (2) Evaluate the prediction status The uncertainty of the prediction error covariance matrix of vehicle v in time slot t is obtained, which is recorded as (3) In the update phase, the vehicle sensor obtains the measured state of vehicle v in time period t and the measurement error covariance matrix O v (t), combined with the measurement error covariance matrix O v (t) and the prediction error covariance matrix Calculate the Kalman gain K v (t); (4) Using the Kalman gain to fuse the predicted state with the measured state, a more accurate optimal estimated state is obtained; (5) Finally, the prediction error covariance matrix is ​​used to update the optimal estimation error covariance matrix of vehicle v in time period t, denoted as S v (t).

8. The RSU-assisted vehicle networking task offloading method according to any one of claims 1 to 5, characterized in that: The PPO algorithm employs two neural networks: a policy network and a value network, wherein the policy network generates action probabilities for a given state, while the value network estimates the expected value of each action.

9. The RSU-assisted vehicle networking task offloading method according to claim 8 is characterized in that: The parameters of the policy network are denoted as θ π , parameter θ π During training, the policy network is updated by maximizing L π (θ π ) to update the objective function: in represents the probability ratio between the new strategy and the old strategy; ∈ is a hyperparameter that controls the allowed range of this difference; Φ represents the advantage function, which is used to measure the state s t Next action a t The value relative to the average of all actions.

10. The RSU-assisted Internet of Vehicles task offloading method according to claim 8, characterized in that: The loss function of the value network represents the mean square error between the predicted value and the reward, which is expressed as follows: are the parameters of the value network, Represents the transition state s t As the output of the value network; r t is the reward value.

Citation Information

Cited By

  • Unmanned bus stop state detection method and system based on vehicle-road cooperation

    CN120299262A