RSU auxiliary single-hop task unloading optimization method based on multi-agent reinforcement learning

Through multi-agent reinforcement learning and GRU model to predict the distribution of adjacent vehicles, combined with the MAPPO algorithm to optimize RSU task offloading, the problems of limited RSU computing power and dynamic location changes are solved, and task processing with low latency and low energy consumption are achieved.

CN120282211APending Publication Date: 2025-07-08ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510443130.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In autonomous driving environments, RSU has limited computing power and dynamic location changes, and the prior art is difficult to effectively predict environmental changes and optimize task offloading, resulting in computing delay and energy consumption problems.

Method used

The RSU-assisted single-hop task offloading method based on multi-agent reinforcement learning is adopted, and the distribution of neighboring vehicles is predicted through dynamic environment perception and GRU model, and distributed training is carried out in combination with the MAPPO algorithm to optimize task offload decisions to minimize delay and energy consumption.

Benefits of technology

The timeliness and resource utilization efficiency of task processing are significantly optimized, the system's adaptability and stability in dynamic environments are improved, and the calculation delay and energy consumption are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282211A_ABST
    Figure CN120282211A_ABST
Patent Text Reader

Abstract

The invention discloses an RSU auxiliary single-hop task unloading optimization method based on multi-agent reinforcement learning, and belongs to the technical field of Internet of Vehicles and edge computing. The method comprises the following steps: actively sensing a road environment through an RSU, and generating a sensing task; the distribution state of adjacent vehicles is predicted through a GRU model, and an adjacent node set of each RSU is updated in real time; according to task characteristics, tasks generated by the RSU are unloaded to adjacent edge computing nodes, and time delay and energy consumption unloaded to other nodes are calculated based on a time delay and energy consumption calculation model; optimizing the objective function to minimize the total time delay and the total energy consumption of calculation, and modeling a task decision problem as a Markov decision process; and performing distributed training and optimization by using an MAPPO algorithm to obtain an optimal task unloading decision. According to the method, the long-time average task processing time delay and energy consumption can be effectively reduced, and the task unloading efficiency and robustness in the dynamic Internet of Vehicles environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle networking and edge computing, and specifically relates to an RSU-assisted single-hop task offloading and dynamic vehicle prediction method based on the MAPPO algorithm. Background Art

[0002] With the development of 5G / 6G technologies, the importance of autonomous driving in modern transportation has gradually increased. However, the computing resources of vehicles themselves are limited, and problems such as latency and high energy consumption are likely to occur during execution, making it difficult to meet the requirements of complex computing tasks with high real-time requirements. To address these issues, existing research has proposed a solution to sink computing power to the network edge through mobile edge computing (MEC) to provide a low-latency and high-efficiency solution for autonomous driving tasks.

[0003] Among them, roadside units (RSUs), as nodes for mobile computing, provide computing support by sensing the environment and collaborative computing. However, due to the limited computing power of RSUs and the dynamic changes in their positions and the distribution of neighboring nodes, how to effectively predict environmental changes and optimize task offloading has become a key issue.

[0004] Currently, although centralized task offloading can consider the global environment, it cannot cope with the dynamic challenges brought by node changes; while distributed solutions are more adaptable, there is still much room for improvement in collaborative decision-making and performance optimization. Therefore, it is particularly important to design a task offloading method based on dynamic environment perception and distributed collaboration. Summary of the Invention

[0005] The present invention aims to provide an RSU-assisted single-hop task offloading optimization method based on multi-agent reinforcement learning, so as to solve the problem of task offloading of RSUs in a high-speed mobile environment in the prior art. Through dynamic environment perception and multi-agent reinforcement learning, the present invention has achieved significant optimization of task processing latency and energy consumption, and at the same time improved the adaptability and stability of the vehicle networking system in a dynamic environment.

[0006] To achieve the above object, the technical solution provided by the present invention is as follows:

[0007] The present invention provides an RSU-assisted single-hop task offloading optimization method based on multi-agent reinforcement learning, including:

[0008] Actively sensing the road environment through an RSU to generate sensing tasks;

[0009] Modeling the historical trajectories of vehicles through a gated recurrent unit (GRU) model to predict the distribution states of neighboring vehicles, and updating the neighboring node set of each RSU in real time;

[0010] According to the task characteristics, the tasks generated by the RSU are offloaded to neighboring edge computing nodes, and based on the latency and energy consumption calculation models, the latency and energy consumption of offloading tasks to other nodes are calculated;

[0011] The objective function is optimized to minimize the total computing latency and total energy consumption, and the task decision problem is modeled as a Markov decision process;

[0012] The MAPPO algorithm is used for distributed training and optimization to obtain the optimal task offloading decision.

[0013] Based on the problem of task offloading by RSU in a high-speed mobile environment, the present invention proposes an RSU-assisted single-hop task offloading optimization method based on multi-agent reinforcement learning. First, by using vehicle historical trajectory data, the GRU model can predict the vehicle positions within a future period of time, helping the RSU to identify in advance which vehicles may become suitable computing nodes, thus ensuring the efficiency and accuracy of task offloading; this prediction mechanism helps to cope with the dynamic challenges caused by vehicle position changes and improves the success rate of task offloading.

[0014] At the same time, the present invention also proposes to perform task offloading by balancing energy consumption and latency. By accurately predicting the nodes and timing of task offloading, the RSU can flexibly choose to offload tasks to different nodes (adjacent vehicles, RSU or the cloud) according to the current computing resources and network conditions, aiming to maximize the overall performance of the system. Through such a distributed offloading strategy, the present invention not only reduces the latency during task execution, but also significantly reduces the energy consumption and improves the resource utilization rate.

[0015] According to any of the technical solutions described in the present invention, the optimized objective function is:

[0016]

[0017] C3:D r (t) ≤ d r

[0018] where, D r (t) is the total latency of task calculation; represents the set of valid vehicles within the communication range of RSUr at time slot t; represents the set of adjacent RSUs of RSUr at time slot t; E r (t) is the total energy consumption of task calculation; γ is a weight factor for balancing the trade-off between calculation latency and calculation energy consumption, and its value range is [0, 1]. The determination of the value of γ depends on the relative importance of delay and energy consumption in the actual application scenario; x r,y(t) = {0, 1} is a binary variable, which indicates whether the task is offloaded to each possible offloading destination y; d r is the maximum allowable delay of the task.

[0019] Constraint C1 indicates that a task is assigned to a specific computing offloading node only when x r,y (t) is not 0; C2 indicates that each task r can be assigned to only one computing node or local device for execution, ensuring the uniqueness of task assignment; C3 means that the actual processing delay D r (t) must be less than or equal to the maximum allowable delay d r .

[0020] According to any of the technical solutions of the present invention, the edge computing nodes adjacent to each RSU include adjacent vehicle nodes, a cloud server, and the RSU; the total computing delay and total energy consumption respectively refer to the total computing delay and total energy consumption calculated on the adjacent vehicle nodes, cloud server, or RSU.

[0021] According to any of the technical solutions of the present invention, the use of the MAPPO algorithm for distributed training and optimization specifically includes:

[0022] Initializing the parameters of the policy network and value network of the agent;

[0023] At each time step t, each RSU selects an action a r (t) according to its local observation o r (t), and obtains a reward R r (t) and a new state s(t + 1) after executing this action, and at the same time records this interaction in the replay buffer for subsequent training use;

[0024] During a fixed training period, sample data from the replay buffer and use this data to calculate the advantage function A(t), thereby updating the parameters of the policy network and value network;

[0025] Repeat the above process until the training reaches a convergence state.

[0026] According to any of the technical solutions of the present invention, within each training period of the MAPPO algorithm, based on the sampled data, the new edge computing and caching policy network π θ is optimized, and the objective function is as follows:

[0027]

[0028] Among them, represents the probability ratio of the new and old policy networks, which is used to control the amplitude of policy update, a t represents the perception or information of the RSU about the environment at the current time slot t, At Denote the advantage function, which is used to measure the current action a t The advantage relative to the average policy; ∈ represents the clipping parameter, which is used to limit the policy update amplitude to ensure training stability.

[0029] According to any one of the technical solutions described in the present invention, predicting the distribution state of neighboring vehicles through a gated recurrent unit model to update the neighboring node set of each RSU in real time specifically includes:

[0030] Collect the historical position data of vehicle v in the k time slots before time slot t:

[0031] L v (t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)}

[0032] Among them, L v (t) is the set of historical position sequences input to the GRU model, k is the window length of the GRU input sequence, and the number of consecutive position data input to the GRU each time; l v (t - k) is the position data in the input sequence, representing the position of vehicle v at time slot t - k;

[0033] Input the data into the GRU model to generate a position sequence:

[0034]

[0035] The predicted value output by the model is:

[0036]

[0037] Among them, Y′ is the set of predicted positions output by the GRU model, is the position predicted by the GRU for vehicle v at time slot t - k + l.

[0038] Its true position is:

[0039]

[0040] According to any one of the technical solutions described in the present invention, the loss function of the GRU model is: Among them: n is the time slot index of iteration, and the range is l′ v (n) is the position of vehicle v predicted by the GRU model at time slot n; l v (n) is the true position of vehicle v at time slot n; is the normalization factor of the loss function, representing the number of time slots for calculating the error. After the model training is completed, the following can be input:

[0041] L v (t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)}

[0042] The model will output:

[0043] l′ v (t)

[0044] That is, the predicted position of the vehicle at time t.

[0045] According to any one of the technical solutions of the present invention, the Markov decision process includes three key elements: state, action, and reward, where: the state space describes the environmental and task distribution of the system at a certain moment, and is used to guide the decision-making of the offloading strategy. The state of the system is expressed as: o r (t) = {task information, neighboring node information, vehicle distribution information}; the action space represents the task offloading decision of the RSU in each time slot, and is defined as offloading the task to different nodes; the reward function is used to minimize the task delay and energy consumption on the basis of meeting the system constraint conditions.

[0046] According to any one of the technical solutions of the present invention, the specific action space is:

[0047] a r (t) = x r,y (t)

[0048]

[0049] Among them, a r (t) represents the action of RSUr in time slot t; x r,y (t) represents a binary decision variable, indicating whether the task is offloaded to node y; the specific action set includes:

[0050] Neighboring vehicles: There are a total of actions, indicating offloading the task to the currently communicable neighboring vehicles;

[0051] Neighboring RSU: There are a total of actions, indicating offloading the task to other neighboring RSU;

[0052] Cloud server: 1 action, indicating offloading the task to the remote cloud for processing.

[0053] According to any one of the technical solutions of the present invention, the reward when the constraint conditions are met is:

[0054] R r (t) = -[γD r (t) + (1 - γ)E r (t)]

[0055] The penalty when the constraint conditions are not met is:

[0056] R r (t) = -Z

[0057] where D r (t) represents the total delay of the task; E r (t) represents the total energy consumption of the task; γ represents a parameter for balancing delay and energy consumption, and its value range is [0, 1]; Z represents a fixed large positive value, which is generally tuned through experiments. Usually, a penalty value that can effectively balance the system optimization objectives (such as delay and energy consumption) and constraint conditions is selected, so as to punish the actions that violate the constraint conditions, and it is further preferably between 180 and 200.

[0058] Adopting the technical solution provided by the present invention, compared with the prior art, it has the following beneficial effects:

[0059] (1) By combining the dynamic trajectory prediction of the GRU model, the distributed offloading strategy for balancing energy consumption and delay, and the collaborative work of RSU, the present invention optimizes the task offloading process between vehicles and computing nodes, significantly improving the timeliness, accuracy and resource utilization efficiency of task processing.

[0060] (2) By optimizing the objective function of the delay and energy consumption calculation model and selecting appropriate computing offloading and energy consumption decisions, the present invention can achieve a balance between minimizing computing delay and minimizing energy consumption.

[0061] (3) By considering the collaborative effects of multiple edge nodes (such as RSU, vehicles, cloud servers), the MAPPO algorithm can dynamically generate optimal offloading decisions in a complex network environment, thus solving the common optimization problem of task delay and energy consumption. At the same time, the MAPPO algorithm performs online learning and interaction through the policy network, and uses the clipped objective function to limit the policy update amplitude, ensuring the stability and efficiency of training. In addition, MAPPO uses the advantage function to evaluate decision-making actions, guiding the policy network to perform precise optimization, so as to improve the overall performance of the system in multi-node and multi-task scenarios. Compared with traditional single-node or static scheduling methods, this algorithm has stronger dynamic adaptability and robustness, and can effectively cope with the changes and uncertainties of network states. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is the flowchart of the present invention;

[0063] Figure 2 It is a schematic diagram of the MAPPO parameterized target policy in the embodiments of the present invention;

[0064] Figure 3 It is a system scenario diagram in the embodiments of the present invention;

[0065] Figure 4 and 5 Figures 6 are comparison diagrams of the reward value, task processing delay, and energy consumption under different algorithms (centralized task offloading scheme, greedy algorithm, and the method of the present application). The method of the present application in the figure corresponds to the algorithm effect diagram of the embodiments of the present invention. Detailed implementation manners

[0066] To further understand the content of the present invention, the present invention will be described in detail below in conjunction with specific implementation manners.

[0067] The embodiments of the present invention provide an RSU-assisted single-hop task offloading optimization method based on multi-agent reinforcement learning. In a road environment, the RSU actively senses road environment information and generates sensing tasks. During the process of task processing, it can choose to offload the tasks to nearby edge computing nodes for calculation. The edge computing nodes on the road can choose vehicle nodes, cloud servers, and the RSU itself as computing nodes. Combining Figure 1 and Figure 2 shown, the task offloading optimization method specifically includes:

[0068] Actively sense the road environment through the RSU to generate sensing tasks;

[0069] Model the historical trajectories of vehicles through a gated recurrent unit (GRU) model, predict the distribution states of neighboring vehicles, and update the neighboring node sets of each RSU in real time;

[0070] According to the task characteristics, offload the tasks generated by the RSU to neighboring edge computing nodes, and calculate the delay and energy consumption of offloading the tasks to other nodes based on the delay and energy consumption calculation models;

[0071] Optimize the objective function to minimize the total calculation delay and total energy consumption, and model the task decision problem as a Markov decision process;

[0072] Use the MAPPO algorithm for distributed training and optimization to obtain the optimal task offloading decision.

[0073] Specifically, when the RSU generates sensing tasks, the processing of the tasks can only be offloaded to a certain computing offloading node for calculation, and there is no direct communication between vehicles.

[0074] Among them, the gated recurrent unit (GRU) model is used to model the historical trajectory of the vehicle to predict the distribution state of neighboring vehicles, so as to update the neighboring node set of each RSU in real time. This process specifically includes:

[0075] (1) Collect the historical position data of vehicle v in the k time slots before time slot t:

[0076] L v (t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)}

[0077] Among them, L v (t) is the set of historical position sequences input to the GRU model, k is the window length of the GRU input sequence, that is, the number of consecutive position data input to the GRU each time; l v (t - k is the position data in the input sequence, indicating the position of vehicle v at time slot t - k;

[0078] (2) Input the above historical data into the GRU model to generate a position sequence:

[0079]

[0080] (3) The predicted value output by the GRU model is:

[0081]

[0082] Among them, Y′ is the set of predicted positions output by the GRU model, is the position predicted by the GRU for vehicle v at time slot t - .

[0083] Its true position is:

[0084]

[0085] As a further preferred method, the loss function of the GRU model in the present invention is:

[0086]

[0087] Among them: n is the time slot index of iteration, and the range is l′ v (n) is the position of vehicle v predicted by the GRU model at time slot n; l v (n) is the true position of vehicle v at time slot n; is the normalization factor of the loss function, indicating the number of time slots for calculating the error. After the model training is completed, it can be input:

[0088] Lv l(t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)}

[0089] The model will output:

[0090] l′ v (t)

[0091] That is, the predicted position of the vehicle at time t.

[0092] Through the above steps, the RSU can predict the distribution status of neighboring vehicles, determine whether the vehicle position is within the RSU communication range, and thus update the neighboring node set in real time to provide support for subsequent task offloading.

[0093] As a preferred implementation of the present invention, the neighboring edge computing nodes include neighboring vehicles, other RSUs or cloud servers. To further optimize task offloading, the present invention designs a total delay and total energy consumption model for tasks.

[0094] 1. Delay model:

[0095] The total delay includes transmission delay, waiting delay, and computing delay, which are specifically as follows.

[0096] (1) When the task is offloaded to the neighboring RSU m, the total time delay is denoted as including three parts: transmission delay, waiting delay, and computing delay, which are respectively denoted as That is:

[0097]

[0098] Among them, is the transmission delay of data upload, is the waiting delay of the task on the target RSU, is the computing delay of the task on the target RSU.

[0099] The formula for data upload transmission delay is as follows:

[0100]

[0101] where s r is the size of the input data of the task, is the transmission rate between RSU r and the target RSU m (r and m are different RSU numbers).

[0102] The formula for waiting delay is:

[0103]

[0104] Among them, q is the task index in the task queue, and s q is the input data size of task q in the queue, and c q is the computational complexity coefficient of task q in the queue (computational requirement per unit data of the task), and f m is the computing power of the RSUm (amount of data that can be processed per unit time).

[0105] The calculation delay formula is:

[0106]

[0107] Among them, c r is the computational complexity coefficient of task r, and s r is the input data size of the task, and f m is the computing power of the RSU m (amount of data that can be processed per unit time).

[0108] (2) When the task is offloaded to vehicle v, similar to the calculation when the task is offloaded to the RSU, the total time delay is denoted as including the transmission delay the waiting delay and the calculation delay The total time delay for calculating the task on the vehicle is:

[0109]

[0110] Among them, is the transmission rate between the RSUr and the target vehicle v, and f v is the computing power of vehicle v (amount of data that can be processed per unit time), q is the task index in the task queue, and s q is the input data size of task q in the queue, and c q is the computational complexity coefficient of task q in the queue (computational requirement per unit data of the task), and c r is the computational complexity coefficient of task r, and s r is the input data size of the task.

[0111] (3) Similarly, when the task is offloaded to the cloud server, the total time delay is denoted as:

[0112]

[0113] Among them is the transmission rate between the RSUr and the cloud server C, and f C is the computing power of the cloud server C, q is the task index in the task queue, and s q is the input data size of task q in the queue, and cq is the computational complexity coefficient of task q in the queue (computing requirement per unit data of the task), c r is the computational complexity coefficient of task r, s r is the input data size of the task.

[0114] Therefore, the total latency for the task to be computed on the RSU, on the vehicle, and on the cloud server is:

[0115]

[0116] where R2V, R2R, and R2C are the sets of vehicles, RSUs, and cloud servers respectively, and x r,y (t) = {0, 1} is a binary variable that indicates whether the task is offloaded to that location for each possible offloading destination y, where

[0117] 2. Energy consumption model:

[0118] The total energy consumption includes transmission energy consumption and computing energy consumption, where:

[0119] (1) Energy consumption of the RSU computing node includes transmission energy consumption and computing energy consumption

[0120] Transmission energy consumption:

[0121] Computing energy consumption:

[0122]

[0123] where P R is the transmission power of the RSU, ξ m is the unit computing energy consumption coefficient of the RSU, c r is the computational complexity coefficient of task r, s r is the input data size of the task, is the transmission latency for data uploading on the RSU.

[0124] (2) Energy consumption of the vehicle computing node includes transmission energy consumption and computing energy consumption

[0125] Transmission energy consumption:

[0126] Computing energy consumption:

[0127]

[0128] Among them, P R is the transmission power of the RSU, and ξ v is the unit computing energy consumption coefficient of the vehicle, and c r is the computing complexity coefficient of task r, and s r is the size of the input data of the task, is the transmission delay of data upload on the vehicle.

[0129] (3) Energy consumption of the cloud server includes transmission energy consumption and computing energy consumption

[0130] Transmission energy consumption:

[0131] Computing energy consumption:

[0132]

[0133] Among them, ξ C is the unit computing energy consumption coefficient of the cloud server.

[0134] Therefore, the total energy consumption of the task when computing on the RSU, on the vehicle, and on the cloud server is:

[0135]

[0136] Among them, R2V, R2R, and R2C are the sets of vehicles, RSUs, and cloud servers respectively, and x r,y (t) = {0, 1} is a binary variable, which indicates whether the task is offloaded to that location for each possible offloading destination y. Among them,

[0137] The optimization objective of the present invention is to minimize the computing delay and the energy consumption. By further optimizing the objective function and selecting appropriate computing offloading and energy consumption decisions, a balance can be achieved between minimizing the computing delay and the minimum energy consumption. Specifically, the objective function of the embodiment of the present invention is:

[0138]

[0139] C3:D r (t) ≤ d r

[0140] Among them, D r (t) is the total delay of task computing; represents the set of valid vehicles within the communication range of RSUr at time slot t; Denote the set of adjacent RSU of RSUr at time slot t; E r (t) is the total energy consumption for task calculation; γ is a weight factor that balances the trade-off between calculation delay and calculation energy consumption, and its value range is [0, 1]. The determination of the value of γ depends on the relative importance of delay and energy consumption in the actual application scenario; x r,y (t) = {0, 1} is a binary variable, which indicates whether the task is offloaded to this location for each possible offloading destination y; d r is the maximum allowable delay of the task.

[0141] Constraint C1 indicates that a task is assigned to a specific calculation offloading node only when x r,y (t) is not 0; C2 indicates that each task r can only be assigned to one calculation node or local device for execution, ensuring the uniqueness of task assignment; C3 represents the actual processing delay D r (t) must be less than or equal to the maximum allowable delay d r of the task.

[0142] Furthermore, the Markov decision process includes three key elements: state, action, and reward. Among them: the state space describes the environment and task distribution of the system at a certain moment, and is used to guide the decision-making of the offloading strategy. In the embodiment of the present invention, the state of the system is represented as: o r (t) = {task information, neighboring node information, vehicle distribution information}.

[0143] The action space represents the task offloading decision of the RSU in each time slot, and is defined as offloading the task to different nodes;

[0144] a r (t) = x r,y (t)

[0145]

[0146] where, a r (t) represents the action of RSUr at time slot t; x r,y (t) represents a binary decision variable, indicating whether the task is offloaded to node y; specifically, the action set includes:

[0147] Neighboring vehicles: a total of actions, indicating offloading the task to the currently communicable neighboring vehicle;

[0148] Neighboring RSU: a total of actions, indicating offloading the task to other neighboring RSU;

[0149] Cloud server: 1 action, indicating offloading the task to the remote cloud for processing.

[0150] The design of the reward function is the key to reinforcement learning. Generally speaking, the reward function is related to the established objective function. In the embodiments of the present invention, on the basis of satisfying the system constraint conditions, the latency and energy consumption of the task are minimized. If the constraint conditions C1 - C3 are satisfied, the reward function is expressed as:

[0151] R r (t)= - [γD r (t)+(1 - γ)E r (t)]

[0152] If the constraint conditions C1 - C3 are not satisfied, the reward function is expressed as:

[0153] R r (t)= - Z

[0154] Where, D r (t) represents the total latency of the task; E r (t) represents the total energy consumption of the task; γ represents a parameter for weighing latency and energy consumption, and its value range is [0, 1]; Z represents a fixed large positive value used to punish actions that violate the constraint conditions.

[0155] As a further preferred implementation, the use of the MAPPO algorithm for distributed training and optimization specifically includes:

[0156] First, by initializing the parameters of the policy network and value network of the agent, a foundation is laid for the training of reinforcement learning;

[0157] Next, at each time step t, each RSU will select an action a r (t) according to its local observation o r (t), and after executing this action, obtain a reward R r (t) and a new state s(t + 1), and at the same time record this interaction in the replay buffer for subsequent training use;

[0158] Then, within a fixed training period, sample data from the replay buffer, and use this data to calculate the advantage function A(t), thereby updating the parameters of the policy network and value network;

[0159] Finally, repeat the above process until the training reaches a convergence state.

[0160] The schematic diagram of the MAPPO parameterized target policy is as Figure 2 shown. Through such an iterative optimization process, the MAPPO algorithm can continuously improve the decision-making ability of task offloading to achieve effective optimization of latency and energy consumption. In each training period of the MAPPO algorithm, based on the sampled data, the new edge computing and caching policy network π θ is optimized, and the objective function is as follows:

[0161]

[0162] Among them, represents the probability ratio of the new and old policy networks, controls the amplitude of policy update, and A t represents the advantage function, which measures the advantage of the current action a t relative to the average policy. ∈ represents the clipping parameter, which is used to limit the amplitude of policy update to ensure training stability.

[0163] For example, Figure 4 、 Figure 5 and Figure 6 show the comparison charts of the reward value, task processing delay, and task processing cost under different algorithms (centralized task offloading scheme, greedy algorithm, and the method of the present invention). It can be seen from the figure that the algorithm of the present invention achieves a greater reward value, lower task processing delay, and task processing cost.

[0164] In summary, the present invention introduces the deep reinforcement learning method into the RSU-assisted edge computing and the GRU model to predict the distribution state of neighboring vehicles. First, the RSU actively senses the road environment to generate a sensing task, then uses the GRU model to predict neighboring vehicles, and then selects a suitable task computing offloading according to the computing delay and energy consumption after task offloading. In a dynamic vehicle networking environment, each RSU (Road Side Unit) acts as an independent agent, and can only make decisions based on its own local observation information (such as the queue state and computing capabilities of surrounding vehicles and RSU). The present invention transforms the problem into a Markov decision process and applies the MAPPO algorithm to solve this non-convex problem, and can make effective decisions on edge computing and energy consumption optimization. The simulation experiment results show that the method of the present invention can effectively reduce the long-term average task processing delay and energy consumption, and improve the efficiency and robustness of task offloading in a dynamic vehicle networking environment.

Claims

1. An RSU-assisted single-hop task offloading optimization method based on multi-agent reinforcement learning, characterized in that Including: Actively sense the road environment through the RSU to generate sensing tasks; Model the vehicle's historical trajectory through the Gated Recurrent Unit (GRU) model, predict the distribution state of neighboring vehicles, and update the neighboring node set of each RSU in real time; According to the task characteristics, offload the tasks generated by the RSU to neighboring edge computing nodes, and calculate the latency and energy consumption of offloading tasks to other nodes based on the latency and energy consumption calculation model; Optimize the objective function to minimize the total computing latency and total energy consumption, and model the task decision problem as a Markov decision process; Use the MAPPO algorithm for distributed training and optimization to obtain the optimal task offloading decision.

2. The RSU-assisted single-hop task offloading optimization method according to claim 1, wherein The optimized objective function is: Among them, D r (t) is the total delay of task calculation; represents the set of valid vehicles within the communication range of RSUr at time slot t; represents the set of adjacent RSUs of RSUr at time slot t; E r (t) is the total energy consumption of task calculation; γ is a weight factor that balances the trade-off between calculation delay and calculation energy consumption, and its value range is [0, 1]; x r,y (t) = {0, 1} is a binary variable, which indicates whether the task is offloaded to this location for each possible offloading destination y; d r is the maximum allowable delay of the task.

3. The RSU-assisted single-hop task offloading optimization method according to claim 2, wherein The edge computing nodes neighboring each RSU include neighboring vehicle nodes, cloud servers, and RSUs; the total computing latency and total energy consumption respectively refer to the total computing latency and total energy consumption calculated on neighboring vehicle nodes, cloud servers, or RSUs.

4. The RSU-assisted single-hop task offloading optimization method according to any one of claims 1-3, characterized in that, The distributed training and optimization using the MAPPO algorithm specifically includes: Initialize the parameters of the policy network and value network of the agent; At each time step t, each RSU selects an action a r (t) based on its local observation o r (t), and obtains a reward R r (t) and a new state s(t + 1) after executing this action, while recording this interaction in the replay buffer for subsequent training; During a fixed training period, sample data from the replay buffer and use this data to calculate the advantage function A(t), thereby updating the parameters of the policy network and value network; Repeat the above process until the training reaches a convergence state.

5. The RSU-assisted single-hop task offloading optimization method according to claim 4, wherein During each training epoch of the MAPPO algorithm, the new edge computing and caching policy network π is optimized based on the sampled data θ The objective function is as follows: Among them, represents the probability ratio of the new and old policy networks, a t represents the perception or information of the RSU about the environment in the current time slot t, which is used to control the amplitude of policy update, A t represents the advantage function, which is used to measure the current action a t 's advantage relative to the average policy; ∈ represents the clipping parameter, which is used to limit the amplitude of policy update to ensure training stability.

6. The RSU-assisted single-hop task offloading optimization method according to any one of claims 1-3, characterized in that, The modeling of the vehicle's historical trajectory through the Gated Recurrent Unit (GRU) model, predicting the distribution state of neighboring vehicles, and updating the neighboring node set of each RSU in real time specifically includes: Collect the historical position data of vehicle v in the k time slots before time slot t: L v l(t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)} Among them, L v (t) is the set of historical position sequences input into the GRU model, k is the window length of the GRU input sequence, and the number of consecutive position data input into the GRU each time; l v (t - k) is the position data in the input sequence, representing the position of vehicle v at time slot t - k; Input the data into the GRU model to generate a position sequence: L = {[l v (t - k), …, l v (t - k + l - 1)], …, [l v (t - l - 1), l v (t - l), …, l v (t - 2)]}; The predicted value output by the model is: Y′ = {l′ v (t - k + l), …, l′ v (t - 1)} Among them, Y′ is the set of predicted positions output by the GRU model, the time series length is l, l′ v (t - k + l) is the position of vehicle v predicted by GRU at time slot t - k + l; Its true position is: Y = {l v (t - k + l), l v (t - k + l + 1), …, l v (t - 1)}.

7. The RSU-assisted single-hop task offloading optimization method according to claim 6, wherein The loss function of the GRU model is as follows: where: n is the time slot index of iteration, ranging from [t - k + l, t - 1]; l′ v (n) is the position of vehicle v predicted by the GRU model at time slot n; l v (n) is the actual position of vehicle v at time slot n; k - l is the normalization factor of the loss function, indicating the number of time slots for calculating the error; after the model training is completed, input: L v l(t) = {l v (t - k), l v (t - k + 1), …, l v (t - 1)} The model will output: l′ v (t) That is, the predicted position of the vehicle at time t.

8. The RSU-assisted single-hop task offloading optimization method according to any one of claims 1-3, characterized in that The Markov decision process includes three key elements: state s t , action a t , and reward r t , where: The state space describes the environmental and task distribution of the system at a certain moment and is used to guide the decision-making of the offloading strategy. The state of the system is represented as: o r (t) = {task information, neighboring node information, vehicle distribution information}; The action space represents the task offloading decision of the RSU in each time slot and is defined as offloading tasks to different nodes; The reward function is used to minimize the task delay and energy consumption on the basis of meeting the system constraints.

9. The RSU-assisted single-hop task offloading optimization method according to claim 8, wherein The specific action space is: a r x(t) = x r,y x(t) where a r (t) represents the action of RSU r at time slot t; x r,y (t) represents a binary decision variable indicating whether a task is offloaded to node y; specifically, the set of actions includes: Proximity vehicle: total actions, indicating offloading the task to the currently communicable neighboring vehicle; Proximity RSU: total actions, indicating offloading tasks to other proximity RSUs; Cloud server: 1 action, indicating offloading the task to the remote cloud for processing.

10. The RSU-assisted single-hop task offloading optimization method according to claim 8, wherein, The reward when the constraint conditions are met is: R r (t) = -[γD r (t) + (1 - γ)E r (t)] The penalty when the constraint conditions are not met is: R r (t) = -Z Among them, D r (t) represents the total delay of the task; E r (t) represents the total energy consumption of the task; γ represents a parameter for balancing delay and energy consumption, and its value range is [0, 1]; Z represents a fixed large positive value used to penalize actions that violate the constraints.

Citation Information

Cited By

  • Industrial task unloading optimization method integrating edge computing and time-sensitive network

    CN120892105A