Edge-end collaborative resource allocation algorithm oriented to load balancing between RSUs
By introducing Kalman filtering and DDPG reinforcement learning algorithms in the Internet of Vehicles, we evaluate link quality and optimize resource allocation, and solve the delay problem caused by high RSU load, and realize the reduction of load balancing and video acquisition delay.
Patent Information
- Application Number
- CN202510323276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the Internet of Vehicles, the RSU may not be able to complete the transcoding task in a timely manner under high load state, resulting in an increase in delay, and the computing resources of adjacent RSUs may be idle, increasing the complexity of resource allocation.
A side-end collaborative resource allocation algorithm for load balancing among RSUs is proposed. Through Kalman filtering, historical data is fused, link quality is evaluated, and a system model is established to minimize video acquisition delay through collaborative transcoding and calculation of offload decisions, and the solution is done using DDPG reinforcement learning algorithm.
It effectively reduces the video acquisition delay, realizes balanced load distribution between RSUs, and improves the overall performance of the Internet of Vehicles system.
Smart Images

Figure CN120186679A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile communications, and more specifically, to a resource allocation algorithm for video stream services in the Internet of Vehicles (IoV). Background Art
[0002] With the continuous development of IoV technology, in-vehicle entertainment video services have become an important part of enhancing the travel experience of passengers. This growth is mainly due to the rapid development of emerging technologies such as video streaming, online gaming, and virtual reality and augmented reality. The popularization of 5G technology has further promoted this trend, greatly improving the quality and efficiency of video services by providing higher data transmission rates and lower latency. Therefore, the application of in-vehicle video services in IoV not only enhances the comfort of travel but also promotes the rapid development of the intelligent transportation ecosystem. However, based on edge-cloud collaboration, due to differences in vehicle distribution, video task requirements, and link status, some Road Side Units (RSUs) may not be able to complete transcoding tasks in a timely manner under high load conditions, resulting in significant delays even when their link quality is good, while the computing resources of adjacent RSUs may be idle. This imbalance exacerbates the complexity of resource allocation and has become a key challenge that urgently needs to be solved.
[0003] Currently, there are already quite a few literatures focusing on the collaborative problem of RSUs in mobile edge computing (MEC), and certain achievements have been made. In response to the challenges brought about by the high bandwidth and low latency requirements in video stream transmission, the authors of the literature [YANG T, TAN Z, XU Y, et al. Collaborative edge caching and transcoding for 360° video streaming based on deep reinforcement learning [J]. IEEE Internet of Things Journal, 2022, 9(24): 25551 - 25564.] proposed a strategy to collaboratively optimize caching and transcoding resources in the edge cluster to reduce quality mismatch, latency, and transmission cost. By modeling the problem as a Markov decision process (MDP) and adopting the deep deterministic policy gradient (DDPG) algorithm to optimize cache replacement and computing resource allocation. In the case of limited resources, regarding the problem of how to meet the strict quality of service (QoS) requirements of terminals and reduce caching costs, the authors of the literature [LIU Q, ZHANG H, ZHANG X, et al. Joint service caching, communication and computing resource allocation in collaborative MEC systems: A DRL-based two-timescale approach [J]. IEEE Transactions on Wireless Communications, 2024, 23(10): 15493 - 15506.] proposed a collaborative MEC framework to maximize long-term QoS and reduce caching costs by optimizing service caching, collaborative offloading, and computing and communication resource allocation. In the scenario of adaptive video streaming, to solve the problem of network congestion caused by large video data volume and reduce the quality of experience (QoE), the authors of the literature [LIU W, ZHANG H, DING H, et al. QoE-aware collaborative edge caching and computing for adaptive video streaming [J]. IEEE Transactions on Wireless Communications, 2024, 23(6): 6453 - 6466.] proposed a method of collaborative edge caching and computing, aiming to maximize QoE by optimizing cache placement, computing resource allocation, and user bitrate adaptation.Aiming at the problems of limited communication and computing resources, unknown global information, and difficult multi-agent collaboration in multimedia edge networks, the authors of [DU Y, HUANG Z, YANG S, et al. Collaborative Task Offloading Based on Deep Reinforcement Learning in Heterogeneous Edge Networks[C]. 2024 International Wireless Communications and Mobile Computing (IWCMC), Ayia Napa, Cyprus, 2024: 375-380.] proposed a distributed computing task scheduling algorithm based on deep reinforcement learning (DRL). By optimizing the collaboration and decision-making process among agents, this algorithm aims to minimize latency and energy consumption while improving convergence. In various vehicular applications and time-varying network states, the authors of [NING Z, ZHANG K, WANG X, et al. Joint computing and caching in 5G-envisioned Internet of vehicles: A deep reinforcement learning-based traffic control system[J]. IEEE Transactions on Intelligent Transportation Systems, 2020, 22(8): 5201-5212.] proposed an intention-based traffic control system. This system improves the profits of mobile network operators and effectively allocates network resources by dynamically coordinating edge computing and content caching.
[0004] Although the existing literature has considered the cooperation problem among RSU in the MEC scenario, and to a certain extent ensured the load balance of RSU and optimized the system objectives. However, these methods still have deficiencies in the resource allocation problem and fail to optimize from a global perspective. First, video transmission is closely related to the link quality, while existing research usually allocates communication and computing resources independently and fails to effectively associate the two. Good link quality is a prerequisite for ensuring high-quality video transmission. Second, existing research does not make full use of the resources of the entire system. In the IoV scenario, not only can RSU cooperate through computing offloading, but RSU and the vehicle side can also jointly transcode video segments to achieve more efficient cooperation. Based on the above analysis, the present invention proposes an edge-side collaborative resource allocation algorithm for load balancing among RSU. On the basis of edge-side collaboration, it further combines computing offloading among RSU for load balancing, and optimizes the resource allocation of the entire system from a global perspective to minimize the video acquisition delay. Summary of the Invention
[0005] The present invention aims to solve the above problems of the prior art and proposes an edge-side collaborative resource allocation algorithm for load balancing among RSU. The technical solution of the present invention is as follows:
[0006] An edge-side collaborative resource allocation algorithm for load balancing among RSU, which includes the following steps:
[0007] First, use Kalman filtering to fuse historical data, and evaluate the link quality through interval type-2 fuzzy logic to provide a more accurate decision-making basis for collaborative transcoding. Second, the collaborative transcoding tasks of RSU can be offloaded to adjacent RSU, and the resource allocation is optimized through the computing offloading decision. Third, based on the collaborative transcoding and computing offloading decisions, a system model for minimizing the video acquisition delay is established, and the RSU load is used as a constraint condition. Subsequently, the DDPG reinforcement learning algorithm is used for solving. Finally, RSU coordinates with the vehicle and adjacent RSU according to the information of each scheme to jointly complete the entire video service process.
[0008] Further, the construction of the link quality model for fusing historical data includes the following steps: First, the vehicle sends a video request to the RSU, and the RSU collects the bandwidth, signal-to-noise ratio, and jitter information of the vehicle. Subsequently, the RSU uses Kalman filtering technology to fuse historical data and real-time measurement data to obtain the fused link information parameters. Then, a link quality model is constructed based on type-2 fuzzy logic, and the link quality is scored by calculating the link information parameters. Finally, a threshold is set. When the score is higher than the threshold, the link quality is determined to be good, otherwise it is average.
[0009] Further, the vehicle request information includes the local computing power p at this time u,loc and the quality video requested It indicates that vehicle u requests video f with quality format l within the coverage of RSU r, and vice versa indicates that the vehicle does not request the video. At the same time, it is assumed that the vehicle can only request video content of one quality format within the coverage of the RSU, so it satisfies:
[0010]
[0011] Furthermore, the three key parameters collected by the RSU and the vehicle are:
[0012] Transmission bandwidth refers to the amount of data that can pass through the link per unit time and represents the transmission capacity of the wireless channel. High-definition video streams require high bandwidth support to improve the user's viewing experience.
[0013] Signal-to-noise ratio is the ratio of the received useful signal strength to the received interference signal strength, which reflects the quality of the wireless link to a certain extent. High signal-to-noise ratio is the basis for ensuring accurate transmission.
[0014] Jitter is the short-term deviation of a signal at a specific moment from its ideal time position, representing the volatility and discontinuity of data transmission. Reducing the impact of jitter on the transmission process can improve the coherence and smoothness of the user's video viewing.
[0015] Furthermore, the link parameters are vulnerable to environmental noise and short-term interference, resulting in the deviation of single-sampling values from the true state or the triggering of incorrect rules, affecting the evaluation accuracy. Therefore, a method of fusing historical data with real-time measurement is proposed to alleviate the problems of short-term fluctuations and random errors and improve the accuracy and robustness of the evaluation. Historical data provides long-term trends and stable characteristics, can smooth abnormal real-time data, and enhance the reliability of link quality evaluation, thus providing support for collaborative transcoding.
[0016] Furthermore, common data fusion methods include weighted average method, least squares method, particle filter, and Kalman filter, etc. Among them, the Kalman filter combines real-time measurement with historical data, estimates the system state by recursively optimizing the weights, can adapt to dynamic changes in real time, and effectively reduces the influence of errors and noise. Therefore, the link quality model for fusing historical data used in the present invention mainly includes the following five steps:
[0017] 1) Kalman filtering process. The filtering processes of signal-to-noise ratio and jitter are the same as that of bandwidth. The present invention takes bandwidth as an example to illustrate the filtering process. Its overall modeling process is as follows
[0018] b t = Ab t-1 + Bμ t + ω t-1 (2)
[0019] χ t = Hbt +υ t (3)
[0020] Wherein, b t and χ t are the estimated value and the actual measured value at time slot t, respectively. A is the state transition matrix. It is considered that within a very short time, the value of the bandwidth remains unchanged, so A = 1. B is an optional correlation matrix between the control input and the bandwidth, and μ t is the control input. Since there is no control input, μ t = 0. ω t-1 represents the process noise, which is usually expressed as a Gaussian random variable with a variance of , that is H is a correlation matrix between the state and the measured value, which is usually considered to be unchanged. In the experiment, the measured value is the state value, so H = 1. υ t represents the measurement noise, similar to ω t-1 , and is usually expressed as a Gaussian random variable with a variance of , that is Therefore, the above formula can be rewritten as
[0021] b t = b t-1 + ω t-1 (4)
[0022] χ t = b t + υ t (5)
[0023] The time update equation of the Kalman filter is
[0024]
[0025] Wherein, represents the predicted value of the bandwidth at time slot t, represents the estimated value of the bandwidth at time slot t, which is the final estimated value corrected by combining the actual measured value. represents the error covariance of the predicted state, which is used to measure the uncertainty of the prediction, and O t represents the error covariance after updating the state, which is used to measure the uncertainty of the final estimate. is the process noise covariance, which represents the uncertainty of the dynamic changes in the model.
[0026] The measurement update equation is
[0027]
[0028] Wherein, t t is the Kalman gain, is the measurement noise covariance.
[0029] To ensure the normal operation of the Kalman filter, it is necessary to know and the value of O0. The accuracy of the value only affects the convergence speed of the evaluation and does not affect the accuracy. Before the system starts, the variance of the first set of historical bandwidth data is calculated to estimate Since may change slowly over a long period of time, it is set to a fixed value. The variance of a set of actually measured bandwidth values is used to estimate and is set to the value of the bandwidth measured for the first time. The value of O0 is usually obtained by calculation. Through the above Kalman filtering process, the bandwidth, signal-to-noise ratio, and jitter information that fuse historical data can be obtained as the input to the subsequent link quality evaluation model.
[0030] 2) Fuzzification. Through the above Kalman filtering process, the bandwidth, signal-to-noise ratio, and jitter information that fuse historical data can be obtained and fuzzified. In an interval type-2 fuzzy set, the membership degree of each parameter is converted into an interval rather than a definite value. To intuitively represent the membership degree of the fuzzy set and facilitate calculation, the present invention selects the commonly used Gaussian membership function. Generally, the upper limit function and the lower limit function of the Gaussian membership function are respectively expressed as
[0031]
[0032] where represents the range of variation of the mean. When , the type-2 membership function can be reduced to a type-1 membership function, and the membership degree at this time is an exact value.
[0033] 3) Fuzzy inference. Based on the established fuzzy rule set, fuzzy logic operations are performed on the fuzzified input, and finally a comprehensive fuzzy output can be obtained. The relationship between the fuzzy input and the fuzzy output is usually expressed in the form of "if - then". At the same time, to ensure accurate evaluation of the link quality, in the design of the fuzzy rules, the importance of the bandwidth is the highest, the signal-to-noise ratio is the second, and the jitter is the lowest. If the bandwidth is good, the signal-to-noise ratio is good, and the jitter is good, the link quality is the best; while if the bandwidth is poor, the signal-to-noise ratio is poor, and the jitter is poor, the link quality is the worst, and at this time, high-quality video transmission tasks cannot be supported.
[0034] 4) Defuzzification. To reduce the computational complexity, the present invention uses the average value method to obtain the defuzzification result, that is
[0035]
[0036] 5) Defuzzification. The present invention uses the centroid method for defuzzification, converting the fuzzy output into a specific score value and normalizing the score value between 0 and 1. The threshold θ = 0.55. When the score is greater than θ, it indicates that the link quality is good, and then the link quality z u,r = 1. Otherwise, the link quality is average, and z u,r = 0.
[0037] Furthermore, it is determined whether to perform collaborative transcoding according to the link quality. When z u,r = 1, the RSU further splits the video stream into two parts, and the splitting ratio is h u,r ∈[0,1], that is, the collaborative transcoding ratio of the RSU. Only when the link quality is good, the RSU collaborates with the vehicle for transcoding, so it satisfies
[0038]
[0039] In addition, Η = {h u,r |u∈U; r∈R} is used to represent the collaborative transcoding strategy of RSU r.
[0040] Furthermore, the RSU can offload the video transcoding task to the adjacent RSU through a wired link for mutual cooperation, and use to represent the computing offloading decision variable. When , it means that the video requested by vehicle u is finally transmitted by RSU r, but is transcoded to the target quality by RSU j. The transcoding task of vehicle u is at most processed by one RSU. When the link quality is not good, only local transcoding is performed, so it satisfies
[0041]
[0042] And use to represent the computing offloading strategy.
[0043] Furthermore, the resource allocation process includes the allocation of communication resources and computing resources.
[0044] Furthermore, the allocation of communication resources is mainly the process of bandwidth allocation. In the long-term evolution vehicle network, the entire radio resource is divided into resource blocks (RBs) in the time domain and frequency domain. Considering the dynamic characteristics of fast vehicle speed and fast network topology change in the vehicle network, use B r (t) to represent the number of RBs contained in RSU r at the t-th time slot. Considering that B r (t) is discrete and finite, it is modeled as a finite state Markov chain (FSMC). At the same time, the available spectrum resources of the RSU are divided into M states, that is Use b u,rDenote the number of RBs allocated by RSU r to vehicle u, and the number of allocated RBs cannot be greater than the number of RBs it currently has. Thus, it satisfies
[0045]
[0046] Use Β = {b u,r | u ∈ U; r ∈ R} to represent the allocation strategy of the bandwidth resources of RSU r.
[0047] Furthermore, the allocation of the computing resources is mainly the allocation of the computing power of the RSU. When it means that the video requested by vehicle u is finally transmitted by RSU r, but transcoded to the target quality by RSU j. At this time, RSU j allocates computing resources to vehicle u. When a vehicle requests a video, the RSU may have other computing tasks, and the vehicle randomly occupies and releases the computing resources. Although the computing resources of the RSU are continuous and limited, they can be discretized. Thus, similar to the bandwidth resources, use P j (t) to represent the computing resources of RSU j at time slot t. And the computing resources are evenly discretized into N states, that is Let p u,j be the computing resources allocated by RSU j to vehicle u. The allocated computing resources cannot be greater than the computing resources it currently has. Thus, it satisfies
[0048]
[0049] And when it means that the video is not transcoded by this RSU j. In addition, use Ρ = {p u,r | u ∈ U; r ∈ R} to represent the computing resource allocation strategy of RSU r.
[0050] Furthermore, from the above collaborative transcoding strategy Η and the computing offloading strategy the total number of cpu cycles consumed by the transcoding task at each RSU at time slot t can be obtained, specifically expressed as
[0051]
[0052] In the formula, h u,r represents the magnitude of the collaboration ratio, d f,l represents the number of bits contained in the video f with the quality version l, and Y(l, l') represents the number of cpu cycles consumed to transcode the video with the quality version l into the content with the version l'. At the same time, according to the computing power P j (t) of RSU j at this time, the load of RSU j can be obtained as
[0053]
[0054] Further, when vehicle u requests video f with quality version l' from RSU r, RSU r first obtains z u,r through the link quality model. Then, together with the relevant resource information, it is sent to the back-end data center. The back-end determines the collaborative transcoding decision based on z u,r , that is, h u,r . Then, based on h u,r , the computing offloading decision is obtained, that is . Finally, h u,r , and the resource allocation decision are sent to RSU r. RSU r collaborates with the vehicle and adjacent RSU according to various decisions to complete this video task. Therefore, the total video acquisition delay D u,r is divided into the vehicle-side processing delay T u and the RSU-side processing delay T r .
[0055] Further, the vehicle-side processing delay includes the transmission delay of the low-quality l = 1 video stream and the transcoding delay of transcoding it into quality l'. Therefore, the processing delay on the vehicle side is:
[0056]
[0057] Further, the is of size
[0058]
[0059] where G u,r is the transmission rate of the V2I link. According to the Shannon formula, the V2I link communication rate from RSU r to vehicle u is
[0060]
[0061] where B0 is the bandwidth amount contained in a single RB, and Υ u,r is the signal-to-noise ratio between vehicle u and RSU r link. Since SNR is a continuous random variable that is difficult to analyze, for ease of calculation, SNR is modeled as an FSMC. In this model, Υ u,r is discretized and quantized into K states: if then Υ u,r = σ0; if then Υ u,r = σ1; …; if then Υ u,r = σ K-1 . Each level represents a state of the FSMC, and the corresponding state space is written as λ = {σ0, σ1,..., σ K-1}.
[0062] Furthermore, the size is
[0063]
[0064] where p u,loc represents the local computing power of vehicle u.
[0065] Furthermore, the processing delay at the RSU r side includes the offloading delay of offloading the low-quality video task to the RSU j the video transcoding delay at the RSU j and the transmission delay of transmitting the video stream of quality l' to the vehicle Therefore
[0066]
[0067] where includes the transmission delay of transmitting the low-quality video with l = 1 from the RSU r to the RSU j and the transmission delay of the RSUj sending the high-quality video back to the RSU r Therefore
[0068]
[0069] Specifically, they are respectively expressed as
[0070]
[0071] where g represents the transmission rate of the wired link between adjacent RSUs, which is a constant. When j = r, the RSU r itself executes the transcoding task, that is Expressed as
[0072]
[0073] Expressed as
[0074]
[0075] Considering that the RSU r can perform transcoding of the remaining segments while transmitting the low-quality video stream. Therefore, the total delay for vehicle u to obtain the video f with quality version l' is the maximum of the RSU-side processing delay and the vehicle-side processing delay, that is D u,r = Max(T r , T u ).
[0076] Furthermore, in order to implement an efficient resource allocation algorithm and provide low-latency video streaming services to users while ensuring the load balance of RSU, a joint optimization problem that simultaneously considers edge-edge collaboration and edge-cloud collaboration needs to be designed. This problem comprehensively considers bandwidth resource allocation, computing resource allocation, collaborative transcoding strategy, and computing offloading strategy, aiming to minimize the video acquisition delay of the entire system. Therefore, the entire problem is described as
[0077]
[0078] In the formula, c1 indicates that vehicle u can only request one format of video content under the coverage of RSU r; c2 ensures that the RSU will only perform collaborative transcoding when the link quality is good; c3 ensures that when performing collaborative transcoding, the computing task is offloaded to at most one RSU for execution, and when the link is average, the computing offloading decision is 0; c4 ensures that the bandwidth allocated by the RSU to the vehicle does not exceed the total bandwidth of the RSU; c5 ensures that the computing power allocated by the RSU to the vehicle does not exceed the total computing power of the RSU; c6 ensures that the difference in the load between two RSUs is less than to ensure load balance; c7, c8, and c9 respectively ensure that the variables z u,r and are both binary variables; c10 ensures that h u,r is a continuous variable between 0 and 1.
[0079] Furthermore, since the V2I channel state and available resources are dynamically changing, Equation (30) is a dynamic optimization problem. At the same time, both computing resource allocation and collaborative transcoding ratio belong to continuous variables. Traditional optimization methods are often inapplicable when facing such complex problems with dynamically changing resources and continuous variables. The DDPG method, combining the advantages of deep learning and deterministic policy gradient, can handle optimization problems in dynamic environments and continuous action spaces and make optimal decisions.
[0080] Furthermore, MDP can be used for problem modeling in reinforcement learning. Therefore, Equation (30) is modeled as an MDP, and then the resource allocation algorithm based on DDPG is described in detail. The elements of MDP, namely the state space S, action space A, and reward R t are defined as follows.
[0081] The state space S is jointly determined by RSU r, vehicle u, and their environment. It includes the basic resource information of the RSU and the vehicle, the request information of the vehicle, and the current V2I link quality information, that is
[0082] 1)
[0083] 2)
[0084] 3)
[0085] 4)
[0086] 5)
[0087] 6)
[0088] Therefore, the state s at the t-th time slot t ∈ S is represented as
[0089]
[0090] The action space A includes bandwidth allocation strategy, computing power allocation strategy, collaborative transcoding strategy, and computing offloading strategy, that is
[0091] 1) Β = {b u,r | u ∈ U; r ∈ R};
[0092] 2) Ρ = {p u,r | u ∈ U; r ∈ R};
[0093] 3) Η = {h u,r | u ∈ U; r ∈ R};
[0094] 4)
[0095] Therefore, the action a at the t-th time slot t ∈ A is represented as
[0096]
[0097] The reward R t is the reward feedback from the environment after the agent RSU interacts with the environment, combined with the optimization goal. In the case of the state being s t , taking the action a t to maximize the reward R t . Taking the negative of the delay directly as the reward can intuitively achieve the goal of maximizing the reward. This design not only simplifies the optimization process but also shows significant performance improvement in experiments and theoretical analysis. Therefore, the reward R at the t-th time slot t is defined as
[0098]
[0099] Furthermore, the basic idea of the DDPG algorithm is to integrate a Deep Neural Network (DNN) into the deterministic policy gradient algorithm to improve the learning efficiency. In DDPG, the core concept is to use two different DNN networks to approximate the policy function and the Q-value function respectively, namely the actor network and the critic network.
[0100] The actor network is a neural network used to represent and learn the policy function. Its main function is to map the state of the environment to a specific action and optimize this policy function during training. The expression of the actor network is usually π(s|θ π ), where s is the state of the environment, π is the actor network (policy function), and θ π are the weight parameters of the actor network. The critic network is a neural network used to estimate the value function of an action. Different from the actor network, the critic network does not directly output an action but evaluates the value (or Q-value) of the state-action pair. The expression of the critic network is usually Q(s,a|θ Q ), where a is the action output by the actor network and θ Q are the weight parameters of the critic network.
[0101] For the actor network, the output of the critic network is used to calculate the policy gradient. The policy gradient is usually expressed as
[0102]
[0103] In the formula, J(θ π ) is the expected return of the policy function, and π(s|θ π ) is the output of the actor network. Using the calculated policy gradient information, the parameters θ π of the actor network are updated by gradient ascent to increase the expected return of the policy function
[0104]
[0105] Here, α π is the learning rate, which controls the step size of parameter update.
[0106] For the critic network, the target Q-value
[0107] ρ t = R t + γQ′(s t+1 ,π'(s t+1 |θ π' )|θ Q′ ). (36)
[0108] In the formula, R t is the current return, and st+1 is the next state, π' and Q' are the outputs of the target actor network and the target critic network, and γ is the discount factor. A regression loss function (such as mean squared error) is used to update the parameters θ of the critic network Q
[0109]
[0110] where W is the size of the batch data sampled from the replay buffer, and α Q is the learning rate of the critic network
[0111] Furthermore, the RSU allocates corresponding resources to the vehicles according to the solved resource allocation scheme, and then collaborates with adjacent RSUs and vehicles according to the collaborative transcoding strategy and the computing offloading strategy to jointly complete the video transmission process
[0112] The advantages and beneficial effects of the present invention are as follows
[0113] 1. Introduce Kalman filtering to fuse historical data and optimize the link quality evaluation model, so as to provide a more accurate basis for edge-side collaborative transcoding
[0114] 2. Design a resource allocation algorithm that comprehensively considers edge-side collaboration and edge-edge collaboration, and focuses on solving the load balancing problem among RSUs. By computing offloading among RSUs, balance the transcoding delay of each RSU, realize the optimal configuration of system resources, and thus effectively reduce the video acquisition delay BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 is the system scenario diagram
[0116] Figure 2 is the link quality model diagram for fusing historical data
[0117] Figure 3 is the update model diagram of Kalman filtering
[0118] Figure 4 is the membership function diagram of bandwidth
[0119] Figure 5 is the relationship diagram between the average collaborative transcoding ratio and the number of vehicles
[0120] Figure 6 is the relationship diagram between delay and the number of vehicles
[0121] Figure 7 is the load situation diagram of each RSU under different algorithms
[0122] Figure 8 is the relationship diagram between the average delay and the number of vehicles under different thresholds
[0123] Figure 9 It is a schematic diagram of the edge - end collaborative resource allocation algorithm process for load balancing among RSUs. Specific implementation manners
[0124] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0125] The technical solution of the present invention is as follows:
[0126] In the vehicle - to - everything network integrating mobile edge computing, from the perspectives of edge - end collaboration and RSU load balancing, a link quality model integrating historical data is designed, and based on this, an edge - end collaborative resource allocation algorithm for load balancing among RSUs is proposed. While ensuring load balancing, it can further reduce the video acquisition delay and improve the overall performance of the vehicle - to - everything system.
[0127] The edge - end collaborative resource allocation algorithm for load balancing among RSUs proposed by the present invention includes the following steps:
[0128] Step 1: When a vehicle requests a video, the RSU collects real - time link parameter information and uses Kalman filtering to fuse historical link parameters to obtain the fused link parameter information. And through the interval type - 2 fuzzy system, the current link quality is obtained and used as the decision - making information for collaborative transcoding.
[0129] Step 2: The RSU sends the request information, resource information, and link information to the back - end data center, and the back - end generates a collaborative transcoding, computing offloading, and resource allocation scheme according to the global system information. During the computing offloading process, the transcoding tasks offloaded to adjacent RSUs need to send the results back to the original RSU after completion, and the original RSU is responsible for transmitting the final video content to the vehicle. The RSU collaborates with the vehicle and adjacent RSUs according to each scheme information to jointly complete the entire video service process.
[0130] To evaluate the performance of the algorithm proposed by the present invention, the system scenario is simulated through the PyCharm simulation platform. Assume that on a straight two - lane road, 3 RSUs are evenly distributed on one side of the lane. They are all connected by wired links, and the coverage range of each is 500×15 m². Vehicles are randomly distributed among the 3 RSUs.
[0131] In the simulation experiment, to verify the effectiveness of the proposed resource allocation algorithm, three typical algorithms were selected for comparison. (1) The resource allocation algorithm with only edge-cloud collaboration: According to the results of link quality, edge-cloud collaborative transcoding is performed, denoted as "ATT". (2) The edge-cloud collaborative resource allocation algorithm based on improved link quality: Using the link quality model of the present invention, and then edge-cloud collaboration, denoted as "ATT_M". (3) The resource allocation algorithm that only considers edge-cloud cooperation: Video transcoding is performed collaboratively among RSU, that is, the algorithm proposed in the literature [YANG T, TAN Z, XU Y, et al. Collaborative edge caching and transcoding for 360° video streaming based on deep reinforcement learning [J]. IEEE Internet of Things Journal, 2022, 9(24): 25551-25564.], denoted as Col. At the same time, the resource allocation algorithm proposed by the present invention is denoted as "ATT_CM".
[0132] Figure 5 It shows the relationship between the average collaborative transcoding ratio at the RSU side and the number of vehicles. In the figure, the Col algorithm only considers the collaboration among RSU and does not consider edge-cloud collaboration, so the ratio is always 1 and remains unchanged. As the number of vehicles increases, the collaborative ratios of the other three algorithms all decrease. Because as the number of vehicles increases, the resources allocated to each vehicle become less, so the ratio of the video transcoded collaboratively at the RSU side also decreases. From the comparison of the curves of the ATT and ATT_M algorithms, it can be seen that the collaborative transcoding ratio of the ATT_M algorithm is slightly higher than that of the ATT algorithm, indicating that in the same environment, more vehicle link qualities are judged to be good, which verifies the effectiveness of the improved link quality model proposed by the present invention. When the number of vehicles is small, the system resources are sufficient and the load under each RSU is small. Even if the task is offloaded to an RSU with stronger computing power, it will not significantly increase the transcoding ratio. Therefore, the transcoding ratios of each algorithm are less different. When the number of vehicles is between 60 and 120, the collaborative ratio of the ATT_CM algorithm is significantly better than the other algorithms, and the gap is large. This is because the improved link quality model can more accurately evaluate the link quality, and the computing offloading among RSU helps to increase the offloading ratio. When the number of vehicles further increases, the ATT_CM algorithm gradually approaches the ATT_M algorithm. Because at this time, the load under each RSU is heavy and has tended to be saturated, so only a few computing offloading tasks are generated at this time, and the transcoding ratio also approaches the ATT_M algorithm. And it should be noted that the algorithm ATT_CM proposed by the present invention is better than all algorithms.
[0133] Figure 6Shows the relationship between the video acquisition delay and the number of vehicles. Among them, subfigure (a) shows the total delay of all vehicles to acquire the video, and subfigure (b) shows the average delay of vehicles to acquire the video. As the number of vehicles increases, the delay of all algorithms increases. Because the system resources are limited, the increase in the number of vehicles leads to a decrease in the resources allocated to each vehicle, so the delay increases. Among them, the delay of the Col algorithm increases the fastest with the increase in the number of vehicles, higher than other algorithms. Because it only considers the cooperation between RSUs and fails to utilize the computing power of the vehicle side. And when the number of vehicles is more than 120, the delay increases significantly. Because at this time, transmitting a large number of high-quality videos will consume a large amount of bandwidth, resulting in a large transmission delay. The total delay curves of the ATT and ATT_M algorithms are close, but the delay of the ATT_M algorithm is slightly lower than that of the ATT algorithm, verifying the superiority of the proposed improved link quality model. The total delay of the ATT_CM algorithm is the lowest, especially when the number of vehicles is between 80 and 130, and the gap with other algorithms is significant. And as can be seen from subfigure (b), when the number of vehicles is less than 60, the delay of the ATT_CM algorithm has little difference from that of the other algorithms, while when the number of vehicles is about 105, the gap is the largest. At this time, the ATT_CM algorithm can make full use of the resources of the entire system to cooperate in video transcoding, so the delay is significantly lower than that of the other algorithms. And when the number of vehicles is greater than 140, the delay is slightly lower than that of the ATT_M algorithm, indicating that at this time the load of each RSU is heavy, and continuing to offload to adjacent RSUs may not be worth the loss. This figure shows that the ATT_CM algorithm proposed by the present invention can significantly reduce the delay in the medium load scenario, and also maintains a low delay in the low load and high load scenarios, verifying its superiority.
[0134] Figure 7 Shows the load conditions of each RSU under different algorithms when the number of vehicles in each RSU is 35, and the computing power of RSU1 is the lowest and that of RSU3 is the highest. Among them, neither the ATT nor the ATT_M algorithm considers the computing offloading between RSUs, so the computing tasks under each RSU tend to be 35. Since only vehicles with good link quality will generate cooperative transcoding tasks at the RSU, the number of computing tasks at some RSUs is less than 35. And the total number of computing tasks under the ATT_M algorithm is more than that under the ATT algorithm, which also verifies the effectiveness of the improved link quality model proposed by the present invention. The task distribution of each RSU under the ATT_CM and Col algorithms is positively correlated with the computing power of the RSU, that is, the number of tasks on RSU1 is the least and the number of tasks on RSU3 is the most, fully reflecting the effectiveness of cooperation between RSUs to achieve load balancing. And all tasks of the Col algorithm are carried out at the edge end, so the total number of tasks is the total number of vehicles, slightly more than that of the ATT_CM algorithm.
[0135] Figure 8 Shows different load thresholds Under this condition, the relationship between the average video acquisition delay of the ATT_CM algorithm and the number of vehicles. As the number of vehicles increases, each The delay under all conditions shows an upward trend. This is because the increase in the number of vehicles leads to a reduction in the resources that can be allocated to each vehicle, thus causing an increase in the delay. When the delay gap between the ATT_CM algorithm and the ATT_M algorithm is not significant. This indicates that under larger conditions, the load difference between RSU is large, and only a small number of computing tasks are offloaded to adjacent RSU. Therefore, the delay performance of the ATT_CM algorithm is similar to that of the non-cooperative ATT_M algorithm. As is further reduced, the load difference between RSU shrinks, and more computing tasks are offloaded to adjacent RSU, thus significantly reducing the average delay, and at the same time indicating the effectiveness of the proposed RSU-based computing offloading for load balancing algorithm. When the strict load balancing strategy can slightly reduce the delay in the case of a small number of vehicles. However, as the number of vehicles increases, the delay increases. This shows that in order to pursue strict load balancing, the delay performance is sacrificed to a certain extent. And it is not difficult to find that compared with other thresholds, has the optimal average delay.
[0136] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0137] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in the present invention, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0138] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0139] The above embodiments should be understood as being only for the purpose of illustrating the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. An edge collaborative resource allocation algorithm for load balancing between RSUs, characterized in that: The following steps are involved:
101. The RSU collects vehicle requests and link parameter information and evaluates the link quality with the vehicle in combination with historical data; 102. Determine the collaborative transcoding strategy based on the link quality obtained in step 101. When the link quality is average, the RSU directly transmits the low-quality video stream to the vehicle, and the vehicle uses local computing power to transcode it into high-quality video. When the link quality is good, the video stream is divided into two parts: one part is transmitted after the RSU completes transcoding, and the other part is directly transmitted to the vehicle and transcoded locally; 103. According to the collaborative transcoding decision obtained in step 102, the transcoding task at the RSU can be offloaded to the adjacent RSU for load balancing, thereby obtaining a computation offloading strategy; 104. According to the collaborative transcoding strategy of step 102 and the computation offloading strategy of step 103, the DDPG algorithm is used to globally allocate system resources to minimize the video acquisition delay; 105. The RSU allocates system resources through the resource allocation scheme of step 104, and cooperates with the vehicle and the adjacent RSU to complete the video transmission task according to the collaborative transcoding strategy of step 102 and the calculation offloading strategy of step 103.
2. According to the edge collaborative resource allocation algorithm for load balancing between RSUs described in claim 1, it is characterized by: The link quality assessment model integrating historical data specifically includes:
201. The RSU collects bandwidth, signal-to-noise ratio and jitter information of the vehicle, and fuses historical data using Kalman filtering to obtain comprehensive link parameter information; 202. According to the link information collected in step 201, dynamically evaluate the link quality based on interval type-2 fuzzy logic.
3. According to the edge collaborative resource allocation algorithm for load balancing between RSUs described in claim 2, it is characterized by: The link quality assessment model based on Kalman filtering and historical data fusion includes: The collected link parameters are easily affected by environmental noise and short-term interference, causing the single sampling value to deviate from the actual state or trigger the wrong rule, affecting the evaluation accuracy. To this end, a method of fusing historical data with real-time measurements is proposed to alleviate the problems of short-term fluctuations and random errors and improve the accuracy and robustness of the evaluation. Historical data provides long-term trends and stable characteristics, which can smooth abnormal real-time data and enhance the reliability of link quality assessment, thereby providing support for collaborative transcoding. Common data fusion methods include weighted averaging, least squares method, particle filtering and Kalman filtering. Among them, Kalman filtering combines real-time measurements with historical data, estimates the system state by recursively optimizing weights, can adapt to dynamic changes in real time, and effectively reduce the impact of errors and noise. The filtering process of signal-to-noise ratio and jitter is the same as that of bandwidth. The filtering process is explained using bandwidth as an example. The overall modeling process is as follows: b t =Ab t-1 +Bμ t +oh t-1 (1) x t =Hb t +u t (2) Where b t and χ t are the estimated value and the actual measured value at time slot t, respectively. A is the state transition matrix. It is assumed that the bandwidth value is constant in a very short time, so A = 1. B is an optional control input and bandwidth association matrix, μ t is the control input. Since there is no control input, μ t =0.ω t-1 represents the process noise, usually expressed as a variance of A Gaussian random variable, that is H is a correlation matrix between states and measurements, and is usually considered invariant. In experiments, the measurements are the state values, so H = 1. t represents the measurement noise, and ω t-1 Similarly, the variance is usually expressed as A Gaussian random variable, that is Therefore, the above formula can be rewritten as b t =b t-1 +oh t-1 (3) x t =b t +u t (4) The time update equation of the Kalman filter is: In the formula, represents the bandwidth prediction value at time slot t, It represents the bandwidth estimation value at time slot t, which is the final estimation value obtained after correction in combination with the actual measurement value. Represents the error covariance of the predicted state, which is used to measure the uncertainty of the prediction, O t Represents the error covariance after updating the state, which is used to measure the uncertainty of the final estimate. is the process noise covariance, which represents the uncertainty of dynamic changes in the model. The measurement update equation is Where, t t is the Kalman gain, is the measurement noise covariance. In order to ensure the normal operation of Kalman filtering, it is necessary to know and the value of O0. The accuracy of the value only affects the convergence speed of the evaluation, not the accuracy. Before the system starts, the variance of the first set of historical bandwidth data is calculated to estimate because It may change slowly over a long period of time, so it is set to a constant value. The variance of a set of bandwidth values measured in practice is estimated and Set to the value of the bandwidth measured for the first time. The value of O0 is usually determined by Through the above Kalman filtering process, the bandwidth, signal-to-noise ratio and jitter information of the fused historical data can be obtained as the input of the subsequent interval type-2 fuzzy logic. The interval type-2 fuzzy logic system is used to evaluate the link quality. It mainly includes the following four steps: 1) Fuzzification. The fused link parameters are selected as the input parameters of the type II fuzzy logic system, and are fuzzified into three quality levels: low, medium, and high. The present invention selects the commonly used Gaussian membership function. Generally, the upper limit function and the lower limit function of the Gaussian membership function are respectively expressed as In the formula, Indicates the range of variation of the mean. When , the type II membership function can be reduced to a type I membership function, and the membership is now an exact value. 2) Fuzzy reasoning. Based on the established fuzzy rule set, fuzzy logic operations are performed on the fuzzified input, and a comprehensive fuzzy output is finally obtained. The relationship between fuzzy input and fuzzy output is usually expressed in the form of "if-then". At the same time, in order to ensure accurate evaluation of link quality, in the design of fuzzy rules, bandwidth is of the highest importance, followed by signal-to-noise ratio, and jitter is the lowest. If the bandwidth is good, the signal-to-noise ratio is good, and the jitter is good, the link quality is the best; if the bandwidth is poor, the signal-to-noise ratio is poor, and the jitter is poor, the link quality is the worst, and it cannot support high-quality video transmission tasks. 3) Fuzzy reduction. In order to reduce the computational complexity, the present invention adopts the average value method to obtain the reduction result, that is, 4) Defuzzification. The present invention uses the centroid method to defuzzify, converting the fuzzy output into a specific score value, and normalizing the score value between 0 and 1. Threshold θ = 0.55, when the score is greater than , it means that the link quality is good, then the link quality z u,r =1. On the contrary, the link quality is average, z u,r =0.
4. According to the edge collaborative resource allocation algorithm for load balancing between RSUs described in claim 1, it is characterized by: The collaborative transcoding strategy specifically includes: When z u,r =1, RSU further splits the video stream into two parts, with a split ratio of h u,r ∈[0,1], which is the ratio of RSU cooperative transcoding. Only when the link quality is good, RSU will cooperate with the vehicle to transcode, so it satisfies In addition, using H={h u,r |u∈U; r∈R} represents the collaborative transcoding strategy of RSU r.
5. The edge collaborative resource allocation algorithm for load balancing between RSUs according to claim 1, characterized in that: Computation offloading strategies include: RSU can offload video transcoding tasks to adjacent RSUs through wired links to collaborate with each other. represents the computational offloading decision variable. When , it means that the video requested by vehicle u is finally transmitted by RSU r, but it is transcoded to the target quality by RSU j. The transcoding task of vehicle u is handled by at most one RSU. When the link quality is poor, it is only transcoded locally, so it satisfies And use Indicates the calculation offloading strategy.
6. The edge collaborative resource allocation algorithm for load balancing between RSUs according to claim 1, characterized in that: The resource allocation process includes the allocation of communication resources and computing resources. Specifically, it includes: The allocation of communication resources is mainly the process of allocating bandwidth. In the long-term evolving Internet of Vehicles, the entire radio resource is divided into resource blocks (RBs) in the time domain and frequency domain. Considering the dynamic characteristics of the Internet of Vehicles, such as the fast speed of vehicles and the fast change of network topology, B r (t) represents the number of RBs contained in RSU r in the tth time slot. Considering B r (t) is discrete and finite, and is modeled as a finite state Markov chain (FSMC). At the same time, the available spectrum resources of RSU are divided into M states, namely Use b u,r It indicates the number of RBs allocated by RSU r to vehicle u, and the number of allocated RBs cannot be greater than the number of RBs contained at this time, so it satisfies Using Β={b u,r |u∈U; r∈R} represents the allocation strategy of RSU r’s bandwidth resources. The allocation of computing resources is mainly the allocation of RSU computing power. When , it means that the video requested by vehicle u is finally transmitted by RSU r, but it is transcoded to the target quality by RSU j. At this time, RSU j allocates computing resources to vehicle u. When the vehicle requests a video, the RSU may have other computing tasks, and the vehicle randomly occupies and releases computing resources. Although the computing resources of RSU are continuous and limited, they can be discretized. Therefore, similar to bandwidth resources, P is used j (t) represents the computing resources of RSU j at time slot t. And the computing resources are discretized into N states on average, that is, Let p u,j is the computing resource allocated by RSU j to vehicle u. The allocated computing resource cannot be greater than the computing resource currently available, so When When , the RSUj does not perform video transcoding. u,r |u∈U; r∈R} represents the computing resource allocation strategy of RSUr.
7. The edge collaborative resource allocation algorithm for load balancing between RSUs according to claim 6, characterized in that: According to the collaborative transcoding strategy Η and the computation offloading strategy €, the total number of CPU cycles consumed by the transcoding task at each RSU at time slot t can be obtained, which is specifically expressed as In the formula, h u,r Indicates the size of the synergy ratio, d f,l represents the number of bits contained in the video f with quality version l, and Y(l,l') represents the number of CPU cycles consumed to transcode the video with quality version l into the content with version l'. At the same time, according to the computing power size P of RSU j at this time j (t), the load of RSU j is 8. The edge collaborative resource allocation algorithm for load balancing between RSUs according to claim 1, characterized in that: When vehicle u requests video f with quality version l' from RSU r, RSU r first obtains z through the link quality model. u,r The value of . Then, it is sent to the backend data center together with relevant resource information. u,r The collaborative transcoding decision is obtained, that is, h u,r . Then according to h u,r The computation offloading decision is obtained, namely Finally, h u,r , The resource allocation decision is sent to RSU r, and RSU r cooperates with vehicles and adjacent RSUs to complete the video task according to various decisions. Therefore, the total video acquisition delay D u,r Divided into vehicle-side processing delay T u and RSU processing delay T r . Vehicle-side processing delay includes the transmission delay of low-quality l=1 video stream and the transcoding delay to transcode it to quality l' Therefore, the processing delay on the vehicle side is: The size is Where G u,r is the transmission rate of the V2I link. According to Shannon’s formula, the V2I link communication rate from RSU r to vehicle u is Where B0 is the bandwidth contained in a single RBs, u,r is the signal-to-noise ratio between the vehicle u and RSU r links. Since SNR is a continuous random variable that is difficult to analyze, SNR is modeled as FSMC for ease of calculation. In this model, Υ u,r is discretized and quantized into K states: if Then u,r =σ0; if Then u,r =σ1; ...; if Then u,r =σ K-1 Each level represents a state of FSMC, and the corresponding state space is written as λ = {σ0,σ1,...,σ K-1 }. The size is In the formula, p u,loc Indicates the local computing power of vehicle u. The processing delay at RSU r includes the offloading delay of the low-quality video task to RSU j. Video transcoding latency at RSU j and the transmission delay of the quality l' video stream to the vehicle Therefore In the formula, Includes the transmission delay of low-quality l = 1 video from RSU r to RSU j and the transmission delay of RSU j transmitting high-quality video back to RSU r Therefore Specifically, they are In the formula, g represents the transmission rate of the wired link between adjacent RSUs, which is a constant. When j = r, RSU r performs the transcoding task itself, that is, Expressed as Expressed as Considering that RSU r can transcode the remaining segments while transmitting low-quality video streams, the total delay for vehicle u to obtain video f with quality version l' is the maximum of the RSU processing delay and the vehicle processing delay, that is, D u,r =Max(T r ,T u ). In order to implement an efficient resource allocation algorithm and provide users with low-latency video streaming services while ensuring RSU load balancing, it is necessary to design a joint optimization problem that considers both edge-to-edge collaboration and edge-to-edge collaboration. This problem comprehensively considers bandwidth resource allocation, computing resource allocation, collaborative transcoding strategy, and computing offloading strategy, aiming to minimize the video acquisition latency of the entire system. Therefore, the entire problem is described as In the formula, c1 indicates that vehicle u can only request video content of one format under the coverage of RSU r; c2 ensures that RSU will cooperate in transcoding only when the link quality is good; c3 ensures that when cooperating in transcoding, the computing task is offloaded to at most one RSU for execution, and when the link is normal, the computing offloading decision is 0; c4 ensures that the bandwidth allocated by RSU to the vehicle does not exceed the total bandwidth of the RSU; c5 ensures that the computing power allocated by RSU to the vehicle does not exceed the total computing power of the RSU; c6 Ensure that the load difference between the two RSUs is less than To ensure load balancing; c7, c8 and c9 respectively ensure that variables z u,r and are binary variables; c10 ensures that h u,r is a continuous variable ranging from 0 to 1.
9. The edge collaborative resource allocation algorithm for load balancing between RSUs according to claim 8, characterized in that: Since the V2I channel status and available resources are dynamically changing, Equation (29) is a dynamic optimization problem. At the same time, computing resource allocation and collaborative transcoding ratio are both continuous variables. Traditional optimization methods are often difficult to apply when faced with complex problems such as dynamic resource changes and continuous variables. The DDPG method, which combines the advantages of deep learning and deterministic policy gradients, can handle optimization problems in dynamic environments and continuous action spaces and make optimal decisions. MDP can be used to model reinforcement learning problems. Therefore, equation (29) is modeled as an MDP, and the resource allocation algorithm based on DDPG is described in detail. The elements of MDP, namely, state space S, action space A, reward R t The definitions are as follows. The state space S is determined by the RSU r and the vehicle u and their environment. It includes the basic resource information of the RSU and the vehicle, the request information of the vehicle, and the current V2I link quality information, which is 1) 2) 3) 4) 5) 6) Therefore, the state s in the tth time slot is t ∈S is represented by Action space A includes bandwidth allocation strategy, computing power allocation strategy, collaborative transcoding strategy and computing offloading strategy, which is 1) B = {b u,r |u∈U;r∈R}; 2)P={p u,r |u∈U;r∈R}; 3)H={h u,r |u∈U;r∈R}; 4) Therefore, action a in the tth time slot t ∈A is represented by Return R t It is the reward of the environment feedback after the intelligent agent RSU interacts with the environment, combined with the optimization goal. t In the case of t , so that the return R t Maximize. Taking the opposite of the delay as the reward directly can intuitively achieve the goal of maximizing the reward. This design not only simplifies the optimization process, but also shows significant performance improvement in experiments and theoretical analysis. Therefore, the reward R in the tth time slot is t Defined as The DDPG algorithm is then used to solve the MDP problem. Finally, the RSU allocates corresponding resources to the vehicle according to the solved resource allocation scheme, and then cooperates with adjacent RSUs and vehicles to complete the video transmission process according to the collaborative transcoding strategy and computation offloading strategy.