Internet of vehicles resource allocation method applied to 5G remote driving
By building a distribution model network architecture of the vehicle layer, edge layer and center layer, and utilizing the computing power of the vehicle layer and edge layer for collaborative processing, the delay and network congestion problems caused by the limitation of vehicle computing power in 5G remote driving are solved, and stable and low-latency communication support is achieved.
Patent Information
- Application Number
- CN202510877686.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
AI Technical Summary
In 5G remote driving, the latency requirements caused by the vehicle's own computing power and storage space limitations cannot be guaranteed, and as the number of connected vehicles increases, the problems of excessive transmission latency and network congestion are difficult to solve.
Construct a distribution model network architecture of the vehicle layer, edge layer and center layer, pre-process task information through the vehicle layer, and use the computing power of the vehicle layer and edge layer for collaborative processing to avoid task transmission to the center layer and ensure that tasks are completed within the standard delay.
It achieves stable and low-latency communication support in 5G remote driving, avoids waste of computing power and network congestion, and meets real-time requirements.
Smart Images

Figure CN120751441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 5G remote driving technology, and in particular to the technical field of a vehicle network resource allocation method applied to 5G remote driving. Background Art
[0002] In recent years, with the accelerated deployment of fifth-generation mobile communication technology, automobiles are constantly developing towards intelligent driving. The advantages of 5G communication technology such as high bandwidth, low latency and low power consumption can basically meet the needs of modern intelligent transportation systems and provide good technical support for autonomous driving. However, due to the limitations of other technologies, manufacturing costs and other factors, the current level of autonomous driving technology is at the L3 level, that is, conditional autonomous driving. The development of vehicle network technology provides new solutions to the above problems.
[0003] At this stage, when autonomous driving technology is not yet fully mature, remote driving technology, as one of the many applications of the Internet of Vehicles, can serve as a backup function in emergency situations such as when the autonomous driving system fails. Due to the limitations of the vehicle's own computing power, storage space, and transmission power, when the autonomous driving system fails or faces a massive amount of data to process, relying solely on the vehicle itself to handle tasks will reduce the frequency of the central processing unit, resulting in the delay requirements of the connected vehicle not being guaranteed and making it difficult to meet real-time requirements. Moreover, as the number of connected vehicles increases, it will also lead to problems such as excessive transmission delay and network congestion.
[0004] Therefore, how to achieve stable and low-latency communication support during 5G remote driving has become a difficult problem that needs to be solved urgently in this field. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a vehicle network resource allocation method applied to 5G remote driving. Obtaining stable and low-latency communication support during 5G remote driving has become a difficult problem that needs to be solved urgently in this field.
[0006] The present invention provides a method for allocating Internet of Vehicles resources for 5G remote driving, comprising: constructing an allocation model network architecture;
[0007] The distribution model network architecture includes a vehicle layer, an edge layer, and a center layer;
[0008] There is 1 vehicle in the vehicle layer;
[0009] The edge layer includes M edge base stations;
[0010] The central layer includes the central controller of the 5G remote control system;
[0011] The vehicle layer is communicatively connected to the edge layer via a wireless network;
[0012] The central layer is connected to the edge layer through optical fibers;
[0013] The vehicle layer obtains task information; the vehicle layer preprocesses the task information to obtain task preprocessing output information; the task preprocessing output information includes task category, quantitative indicator information corresponding to the task category, task offloading ratio Vehicle layer delay
[0014] The vehicle layer calculates the task offloading ratio by using formula 1 for the task category and the quantitative indicator information corresponding to the task category. when When , the vehicle layer completes the task calculation and processing within the standard delay time and outputs the task processing plan information;
[0015] when When the task is completed, the vehicle layer transmits the task pre-processing output information to the edge layer, which processes it or processes it together with the edge layer and the vehicle layer. The task calculation is completed within the standard delay time and the task processing solution information is output.
[0016] The task processing plan information is output from the edge layer to the central layer; the central layer transmits the received task processing plan information to the vehicle layer for 5G remote driving.
[0017] Compared with the existing technology, the present invention has the following advantages: obtaining task information through the vehicle layer; preprocessing the task information by the vehicle layer to obtain task preprocessing output information; the task preprocessing output information includes task category, quantitative index information corresponding to the task category, task offloading ratio, and task offloading information. Vehicle layer delay It is possible to confirm whether the delay in obtaining the processing solution information by the computing power of the vehicle layer for different tasks meets the standard delay requirements. If the standard delay requirements are not met, the task is assigned to the edge layer, and the computing power of the edge layer and the computing power of the vehicle layer are used to process the task simultaneously, so as to complete the task within the standard delay. At the same time, the computing power of the vehicle layer is utilized to avoid wasting computing power, while the problem of long transmission time when sending tasks to the central layer for processing is avoided, and the problem of increased latency or network congestion caused by the central layer accepting too many tasks is avoided.
[0018] Obtaining stable and low-latency communication support when achieving 5G remote driving has become a difficult problem that needs to be solved urgently in this field.
[0019] Furthermore, Formula 1 is: The task j generated by vehicle i at time t is where i∈{1, 2, …, I}, j∈{1, 2, …, J}, t∈{1.2, …, T};
[0020]
[0021] is the task offloading ratio, The available computing power on the vehicle side, T is the number of CPU clock cycles required to process each bit of data in the task category; max Maximum local processing delay on the vehicle side.
[0022] The beneficial effect of adopting the above step is that, through the above steps, the vehicle layer can pre-process different tasks to obtain different unloading ratios.
[0023] Furthermore, the vehicle layer obtains the vehicle layer delay through formula 2 according to the task category and the quantitative indicator information corresponding to the task category
[0024] The formula 2 is
[0025] is the task offloading ratio, is the task data size of vehicle layer i, is the CPU frequency of vehicle layer i, The number of CPU clock cycles required to process each bit of data for different task categories.
[0026] The beneficial effect of adopting the previous step is that when the vehicle layer processes different tasks, the vehicle layer and the edge layer collaborate to complete the task under the premise of the unloading ratio, which is the delay data of the vehicle layer. This is beneficial for the subsequent edge layer to process the received tasks. It can obtain the delay required for the overall task processing to complete, and then analyze how to allocate resources to this task to the edge layer processor based on this ratio.
[0027] Furthermore, the task categories include emergency obstacle avoidance, extreme weather pre-processing, remote parking, and network fluctuation processing;
[0028] The task information includes one or more of emergency obstacle avoidance task information, extreme weather pre-processing task information, remote parking task information, and network fluctuation processing task information;
[0029] Emergency obstacle avoidance mission information includes: obstacles in front of the vehicle lane, videos and pictures of the vehicle in front, and videos and pictures of obstacles in adjacent lanes, vehicles in front, and vehicles behind;
[0030] Extreme weather pre-processing task information includes: road photos, videos, and weather information;
[0031] Remote parking task information includes: the horizontal distance between the vehicle's current position and the center point of the target parking space; the yaw angle between the vehicle's heading and the parking space entry direction;
[0032] The network fluctuation processing task information includes: communication delay, network packet loss rate, and bandwidth utilization.
[0033] Furthermore, the process of processing the task information at the vehicle layer includes the following steps:
[0034] Obtain quantifiable indicators from task information, compare the quantitative indicators with the preset task trigger threshold to obtain the task category, so that the vehicle layer can obtain the task category and the quantitative indicator information corresponding to the task category;
[0035] The quantifiable indicators corresponding to the emergency obstacle avoidance task include: the distance between the front obstacle in the ego vehicle's lane and the ego vehicle, the distance between the front vehicle in the ego vehicle's lane and the ego vehicle, the distance between the obstacle in the adjacent lane and the ego vehicle, the distance between the front vehicle in the adjacent lane and the ego vehicle, the distance between the rear vehicle in the adjacent lane and the ego vehicle, the speed and acceleration of the front vehicle in the ego vehicle's lane, the speed and acceleration of the front vehicle in the adjacent lane, and the speed and acceleration of the rear vehicle in the adjacent lane;
[0036] The quantifiable indicators corresponding to the extreme weather pre-processing task include: light intensity, rainfall rate, and road friction coefficient;
[0037] The quantifiable indicators corresponding to the remote parking task include: distance and yaw angle of the parking target;
[0038] The quantifiable indicators corresponding to the network fluctuation processing task include: network delay time and network packet loss rate;
[0039] The preset task trigger thresholds include emergency obstacle avoidance trigger thresholds, extreme weather pre-processing trigger thresholds, remote parking trigger thresholds, and network fluctuation processing trigger thresholds;
[0040] The emergency obstacle avoidance triggering threshold includes triggering emergency obstacle avoidance when the distance between the obstacle and the preceding vehicle is less than 5m, or when the relative speed between the vehicle and the preceding vehicle is greater than 2m / s;
[0041] The extreme weather pre-processing trigger thresholds include starting 5G remote driving when the illumination is less than 50 lux and the rainfall rate is greater than 10 mm / hr, or when the friction coefficient is less than 0.5;
[0042] The network fluctuation processing trigger threshold includes enabling 5G remote driving when the network delay is greater than 50ms or the packet loss rate is greater than 5%.
[0043] The beneficial effect of adopting the previous step is that the vehicle layer preprocesses the task information to obtain task preprocessing output information; the task preprocessing output information includes the task category and the quantitative indicator information corresponding to the task category.
[0044] Furthermore, the edge layer processing or the edge layer and vehicle layer processing process includes:
[0045] The edge layer is provided with a task preprocessing output information set, an edge layer state information set, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer state information;
[0046] The edge layer status information includes the number of tasks processed by the edge layer and the degree of edge layer congestion when the edge layer receives the vehicle layer task preprocessing output information;
[0047] The edge layer task processing parameters include the resource allocation ratio and task processing delay when the edge layer processes the task preprocessing output information;
[0048] Get the number of tasks processed by the edge server of edge layer m at time t;
[0049] The edge layer processes the task preprocessing output information through formula 3, and obtains the blocking degree of the edge server m at time t: When the degree of obstruction When the blocking criteria are met, the edge server of the edge layer m processes the task preprocessing output information; when the blocking degree When the blocking criteria are not met, edge layer m transmits the task preprocessing output information to the adjacent edge layer m+1 or edge layer m-1 through the wireless network, and judges the blocking degree. or Whether the blocking criteria are met, whether the task pre-processing output information can be processed at the edge server of edge layer m+1 or edge layer m-1, and finally determining the edge layer that processes the task pre-processing output information;
[0050] The edge server of the edge layer m processes the task preprocessing output information including:
[0051] The optimal edge layer task processing parameters of the edge layer m at time t corresponding to the task preprocessing output information of the vehicle i received by the edge layer m at time t are obtained by comparing the task preprocessing output information of the vehicle i received by the edge layer m at time t with the edge layer state information of the edge layer m at time t and the edge layer task processing optimal parameter set corresponding to the edge layer state information.
[0052] The edge server of edge layer m processes the task preprocessing output information according to the optimal parameters of edge layer task processing at time t to obtain task processing solution information.
[0053] Furthermore, the blocking degree of the edge server of edge layer m at time t is is calculated by formula 3, which is:
[0054] is the maximum amount of data that can be stored in the edge server; i∈I means that vehicle i belongs to set I, and I is the set of all vehicles in the vehicle layer.
[0055] The beneficial effect of the previous step is that the edge layer m pre-processes the output information according to the task of vehicle i, and the blocking degree of the edge server of the edge layer m at time t is obtained as Determine whether the task of vehicle i can be processed at edge layer m. If not, send the task preprocessing output information of vehicle i to the adjacent edge layer for processing.
[0056] When the task of vehicle i is processed at the edge layer m, a task preprocessing output information set, an edge layer status information set, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer status information are set according to the edge layer; the task preprocessing output information of vehicle i received by the edge layer m at time t, and the edge layer status information of the edge layer m at time t, are compared with the edge layer task processing optimal parameter set corresponding to the edge layer status information, and the edge layer task processing optimal parameters of the edge layer m at time t corresponding to the task preprocessing output information of vehicle i received by the edge layer m at time t are obtained; thereby, the task processing solution information can be obtained quickly, and fewer edge server resources are required.
[0057] Furthermore, a SAC-E reinforcement learning model is constructed in the central layer; the central layer performs offline training and compression on the SAC-E reinforcement learning model to obtain a set of task preprocessing output information, a set of edge layer state information, and an optimal solution set of edge layer task processing parameters corresponding to both the task preprocessing output information and the edge layer state information;
[0058] Then, during the non-working hours of the edge layer, the optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer status information is transmitted to the edge layer, and the edge layer is regularly optimized and updated.
[0059] The beneficial effect of the previous step is that the SAC-E reinforcement learning model is built in the central layer; the central layer performs offline training and compression through the SAC-E reinforcement learning model, avoiding processing at the edge layer and reducing the processing time at the edge layer.
[0060] Furthermore, the SAC-E reinforcement learning model includes the following:
[0061] The task is calculated by formula 4 Latency offloaded to edge servers
[0062] The task of edge layer m to vehicle i is calculated by formula 5 The delay generated is
[0063] The task is calculated by formula 6 Delay
[0064] The task of the edge server of edge layer m to vehicle i is calculated by formula 7 Resource consumption costs
[0065] Based on MEC The vehicle i task is calculated by formula 8 when multiple vehicle output tasks are processed in edge layer m. Resource allocation ratio at edge layer m
[0066] Furthermore, formula 4 is:
[0067] in
[0068] B is the bandwidth between the vehicle layer and the edge layer, P is the upload power of vehicle i, N0 is the Gaussian white noise power, is the channel gain between the edge layer server and the vehicle layer device, d m,i is the distance between the edge service layer m and vehicle i, and s is the path loss index;
[0069] Formula 5 is:
[0070] in
[0071] When multiple vehicle output tasks are processed at the edge layer m, vehicle i task The resource allocation ratio at the edge layer m; F mec is the computing power of the edge server at edge layer m;
[0072] Formula 6 is:
[0073] Formula 7 is:
[0074] Among them, the price per unit of computing power in the edge server is K;
[0075] Formula 8 is:
[0076]
[0077] Among them, μ1 and μ2 are the weight coefficients of delay and resource consumption cost respectively, usually μ1+μ2=1; constraint C1 specifies the offloading ratio Constraint C2 stipulates that the total amount of computing resources allocated to each edge server shall not exceed its maximum computing capacity; constraint C3 stipulates that the total amount of offloaded data received by the edge server shall not exceed its maximum storage capacity.
[0078] The beneficial effect of the above step is that the vehicle i task is processed by the edge layer m when the multiple vehicle output tasks obtained by the above step are processed. Resource allocation ratio at edge layer m And according to this resource allocation ratio Process the vehicle i task to satisfy the task The delay generated meets the standard delay time, and the task of vehicle i The resource consumption cost is low.
[0079] Preferably, the optimal solution of the edge layer task processing parameters obtained by the optimization formula eight under the Markov decision process includes the edge layer m to vehicle layer i task When processing the preprocessing output information Resource allocation ratio at edge layer m Task Delay
[0080] Markov's state space, action space, and reward function are defined as:
[0081] The state space at time t is defined as:
[0082]
[0083] The action space is defined as: in For processing The edge layer m;
[0084] The reward function is:
[0085] P(x) represents the probability of event x occurring, and the entropy is defined as follows:
[0086]
[0087] The Q function value is as follows:
[0088]
[0089] Among them, represents the entropy of the strategy π in the state space s, and the value function of the V function with the entropy term added is:
[0090]
[0091] The global optimal strategy of the SAC algorithm is expressed as:
[0092]
[0093] The neural network with parameter θ is used to approximate Q(s t ,a t ) value, the probability P(x) of realizing the action tends to be decentralized;
[0094] Objective function J Q The gradient of (θ) can be expressed as:
[0095]
[0096] And the parameters θ of the Q critic network q Update it with the following formula:
[0097]
[0098] Approximate V with a neural network with parameter ψ ψ (s t ), its objective function J V The gradient of (ψ) is expressed as:
[0099]
[0100] The parameters are updated as follows:
[0101]
[0102] In order to obtain the decision action of the strategy, a noisy neural network is set up for reparameterization. The loss function of the strategy network is defined as:
[0103]
[0104] fφ ( t ;s t ) is a reparameterized representation of the action space, is the input noise vector, sampled from a normal distribution with an expected value of 0 and a variance of 1, and φ represents the parameters of the policy network, which is updated by the following formula:
[0105]
[0106] The gradient of the policy can be expressed as:
[0107] The priority experience replay technology is introduced into the SAC algorithm. The improved algorithm can use the TD error to calculate the priority of the sample. The larger the TD error, the more worthy the agent is to learn from the experience. The average of the absolute values of the TD errors of the Critic network and the Actor network is defined as the absolute error of the experience sample in our algorithm, as shown in the following formula:
[0108]
[0109] Get the corresponding experience j The priority indicators are:
[0110] p j =|δ j |+ε;
[0111] Among them, ε is a very small number, in order to ensure that all p j >0, and this patent introduces a priority adjustment factor ζ to indicate the degree of priority playback. If ζ = 0, it indicates a random uniform sampling method; if 0 < ζ < 1, it indicates partial priority sampling; if ζ = 1, it indicates full priority sampling. Therefore, the probability of experience being sampled is expressed as:
[0112]
[0113] Further preferably, the intelligent agent part includes an Actor network and a Critic network, and the Critic network includes four Q networks, two of which are used to reduce over-estimation of actions.
[0114] The beneficial effect of adopting the previous step is to obtain the optimal solution for the calculation of Formula 8. DETAILED DESCRIPTION
[0115] In order to better understand the technical solution of the present invention, the present invention is further described below in conjunction with specific embodiments.
[0116] Example 1:
[0117] According to this embodiment, a method for allocating Internet of Vehicles resources for 5G remote driving is provided, including: constructing an allocation model network architecture;
[0118] The distribution model network architecture includes a vehicle layer, an edge layer, and a center layer;
[0119] There is 1 vehicle in the vehicle layer;
[0120] The edge layer includes M edge base stations;
[0121] The central layer includes the central controller of the 5G remote control system;
[0122] The vehicle layer is communicatively connected to the edge layer via a wireless network;
[0123] The central layer is connected to the edge layer through optical fibers;
[0124] The vehicle layer obtains task information; the vehicle layer preprocesses the task information to obtain task preprocessing output information; the task preprocessing output information includes task category, quantitative indicator information corresponding to the task category, task offloading ratio Vehicle layer delay
[0125] The task categories include emergency obstacle avoidance, extreme weather pre-processing, remote parking, and network fluctuation processing;
[0126] The task information includes one or more of emergency obstacle avoidance task information, extreme weather pre-processing task information, remote parking task information, and network fluctuation processing task information;
[0127] Emergency obstacle avoidance mission information includes: obstacles in front of the vehicle lane, videos and pictures of the vehicle in front, and videos and pictures of obstacles in adjacent lanes, vehicles in front, and vehicles behind;
[0128] Extreme weather pre-processing task information includes: road photos, videos, and weather information;
[0129] Remote parking task information includes: the horizontal distance between the vehicle's current position and the center point of the target parking space; the yaw angle between the vehicle's heading and the parking space entry direction;
[0130] The network fluctuation processing task information includes: communication delay, network packet loss rate, and bandwidth utilization.
[0131] The process of processing task information at the vehicle layer includes the following steps:
[0132] Obtain quantifiable indicators from task information, compare the quantitative indicators with the preset task trigger threshold to obtain the task category, so that the vehicle layer can obtain the task category and the quantitative indicator information corresponding to the task category;
[0133] The quantifiable indicators corresponding to the emergency obstacle avoidance task include: the distance between the front obstacle in the ego vehicle's lane and the ego vehicle, the distance between the front vehicle in the ego vehicle's lane and the ego vehicle, the distance between the obstacle in the adjacent lane and the ego vehicle, the distance between the front vehicle in the adjacent lane and the ego vehicle, the distance between the rear vehicle in the adjacent lane and the ego vehicle, the speed and acceleration of the front vehicle in the ego vehicle's lane, the speed and acceleration of the front vehicle in the adjacent lane, and the speed and acceleration of the rear vehicle in the adjacent lane;
[0134] The quantifiable indicators corresponding to the extreme weather pre-processing task include: light intensity, rainfall rate, and road friction coefficient;
[0135] The quantifiable indicators corresponding to the remote parking task include: distance and yaw angle of the parking target;
[0136] The quantifiable indicators corresponding to the network fluctuation processing task include: network delay time and network packet loss rate;
[0137] The preset task trigger thresholds include emergency obstacle avoidance trigger thresholds, extreme weather pre-processing trigger thresholds, remote parking trigger thresholds, and network fluctuation processing trigger thresholds;
[0138] The emergency obstacle avoidance triggering threshold includes triggering emergency obstacle avoidance when the distance between the obstacle and the preceding vehicle is less than 5m, or when the relative speed between the vehicle and the preceding vehicle is greater than 2m / s;
[0139] The extreme weather pre-processing trigger thresholds include starting 5G remote driving when the illumination is less than 50 lux and the rainfall rate is greater than 10 mm / hr, or when the friction coefficient is less than 0.5;
[0140] The network fluctuation processing trigger threshold includes enabling 5G remote driving when the network delay is greater than 50ms or the packet loss rate is greater than 5%.
[0141] The vehicle layer calculates the task offloading ratio by using formula 1 for the task category and the quantitative indicator information corresponding to the task category. when When t, the vehicle layer completes the task calculation and processing within the standard delay time and outputs the task processing plan information; Formula 1 is: The task j generated by vehicle i at time t is where i∈{1, 2, …, I}, j∈{1, 2, …, J}, t∈{1.2, …, T};
[0142]
[0143] is the task offloading ratio, The available computing power on the vehicle side, T is the number of CPU clock cycles required to process each bit of data in the task category;max Maximum local processing delay on the vehicle side.
[0144] The vehicle layer delay is obtained by formula 2 based on the task category and the quantitative indicator information corresponding to the task category.
[0145] The formula 2 is
[0146] is the task offloading ratio, is the task data size of vehicle layer i, is the CPU frequency of vehicle layer i, The number of CPU clock cycles required to process each bit of data for different task categories.
[0147] when When the task is completed, the vehicle layer transmits the task pre-processing output information to the edge layer, which processes it or processes it together with the edge layer and the vehicle layer. The task calculation is completed within the standard delay time and the task processing solution information is output.
[0148] The processes processed by the edge layer or the edge layer and vehicle layer together include:
[0149] The edge layer is provided with a task preprocessing output information set, an edge layer state information set, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer state information;
[0150] The edge layer status information includes the number of tasks processed by the edge layer and the degree of edge layer congestion when the edge layer receives the vehicle layer task preprocessing output information;
[0151] The edge layer task processing parameters include the resource allocation ratio and task processing delay when the edge layer processes the task preprocessing output information;
[0152] Get the number of tasks processed by the edge server of edge layer m at time t;
[0153] The edge layer processes the task preprocessing output information through formula 3, and obtains the blocking degree of the edge server m at time t: When the degree of obstruction When the blocking criteria are met, the edge server of the edge layer m processes the task preprocessing output information; when the blocking degree When the blocking criteria are not met, edge layer m transmits the task preprocessing output information to the adjacent edge layer m+1 or edge layer m-1 through the wireless network, and judges the blocking degree. or Whether it meets the blocking criteria, it is determined whether the task pre-processing output information can be processed at the edge server of edge layer m+1 or edge layer m-1, and finally the edge layer that processes the task pre-processing output information is determined; the blocking degree of the edge server of edge layer m at time t is is calculated by formula 3, which is:
[0154] is the maximum amount of data that can be stored in the edge server; i∈I means that vehicle i belongs to set I, and I is the set of all vehicles in the vehicle layer;
[0155] The edge server of the edge layer m processes the task preprocessing output information including:
[0156] The optimal edge layer task processing parameters of the edge layer m at time t corresponding to the task preprocessing output information of the vehicle i received by the edge layer m at time t are obtained by comparing the task preprocessing output information of the vehicle i received by the edge layer m at time t with the edge layer state information of the edge layer m at time t and the edge layer task processing optimal parameter set corresponding to the edge layer state information.
[0157] The edge server of edge layer m processes the task preprocessing output information according to the optimal parameters of edge layer task processing at time t to obtain task processing solution information.
[0158] A SAC-E reinforcement learning model is constructed in the central layer; the central layer obtains a set of task preprocessing output information, a set of edge layer state information, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer state information through offline training and compression of the SAC-E reinforcement learning model;
[0159] Then, during the non-working hours of the edge layer, the optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer status information is transmitted to the edge layer, and the edge layer is regularly optimized and updated.
[0160] The SAC-E reinforcement learning model includes the following:
[0161] The task is calculated by formula 4 Latency offloaded to edge servers
[0162] The task of edge layer m to vehicle i is calculated by formula 5 The delay generated is
[0163] The task is calculated by formula 6 Delay
[0164] The task of the edge server of edge layer m to vehicle i is calculated by formula 7 Resource consumption costs
[0165] Based on MEC The vehicle i task is calculated by formula 8 when multiple vehicle output tasks are processed in edge layer m. Resource allocation ratio at edge layer m
[0166] Furthermore, formula 4 is:
[0167] in
[0168] B is the bandwidth between the vehicle layer and the edge layer, P is the upload power of vehicle i, N0 is the Gaussian white noise power, is the channel gain between the edge layer server and the vehicle layer device, d m,i is the distance between the edge service layer m and vehicle i, and s is the path loss index;
[0169] Formula 5 is:
[0170] in
[0171] When multiple vehicle output tasks are processed at the edge layer m, vehicle i task The resource allocation ratio at the edge layer m; F mec is the computing power of the edge server at edge layer m;
[0172] Formula 6 is:
[0173] Formula 7 is:
[0174] Among them, the price per unit of computing power in the edge server is K;
[0175] Formula 8 is:
[0176]
[0177] Among them, μ1 and μ2 are the weight coefficients of delay and resource consumption cost respectively, usually μ1+μ2=1; constraint C1 specifies the offloading ratio Constraint C2 stipulates that the total amount of computing resources allocated to each edge server shall not exceed its maximum computing capacity; constraint C3 stipulates that the total amount of offloaded data received by the edge server shall not exceed its maximum storage capacity.
[0178] The task processing plan information is output from the edge layer to the central layer; the central layer transmits the received task processing plan information to the vehicle layer for 5G remote driving.
[0179] Example 2:
[0180] The contents of this embodiment that are the same as those of Example 1 will not be repeated here. The differences between this embodiment and Example 1 are as follows: This embodiment provides a method for allocating Internet of Vehicles resources for 5G remote driving, further comprising:
[0181] Through the optimization formula eight problem under the Markov decision process, the optimal solution of the edge layer task processing parameters includes the edge layer m to vehicle layer i task When processing the preprocessing output information Resource allocation ratio at edge layer m Task Delay
[0182] Markov's state space, action space, and reward function are defined as:
[0183] The state space at time t is defined as:
[0184]
[0185] The action space is defined as: in For processing The edge layer m;
[0186] The reward function is:
[0187] P(x) represents the probability of event x occurring, and the entropy is defined as follows:
[0188]
[0189] The Q function value is as follows:
[0190]
[0191] Among them, represents the entropy of the strategy π in the state space s, and the value function of the V function with the entropy term added is:
[0192]
[0193] The global optimal strategy of the SAC algorithm is expressed as:
[0194]
[0195] The neural network with parameter θ is used to approximate Q(s t ,a t ) value, the probability P(x) of realizing the action tends to be decentralized;
[0196] Objective function J Q The gradient of (θ) can be expressed as:
[0197]
[0198] And the parameters θ of the Q critic network q Update it with the following formula:
[0199]
[0200] Approximate V with a neural network with parameter ψ ψ (s t ), its objective function J V The gradient of (ψ) is expressed as:
[0201]
[0202] The parameters are updated as follows:
[0203]
[0204] In order to obtain the decision action of the strategy, a noisy neural network is set up for reparameterization. The loss function of the strategy network is defined as:
[0205]
[0206] f φ ( t ;s t ) is a reparameterized representation of the action space, is the input noise vector, sampled from a normal distribution with an expected value of 0 and a variance of 1, and φ represents the parameters of the policy network, which is updated by the following formula:
[0207]
[0208] The gradient of the policy can be expressed as:
[0209] The priority experience replay technology is introduced into the SAC algorithm. The improved algorithm can use the TD error to calculate the priority of the sample. The larger the TD error, the more worthy the agent is to learn from the experience. The average of the absolute values of the TD errors of the Critic network and the Actor network is defined as the absolute error of the experience sample in our algorithm, as shown in the following formula:
[0210]
[0211] Get the corresponding experience j The priority indicators are:
[0212] p j =|δ j |+ε;
[0213] Among them, ε is a very small number, in order to ensure that all p j >0, and this patent introduces a priority adjustment factor ζ to indicate the degree of priority playback. If ζ = 0, it indicates a random uniform sampling method; if 0 < ζ < 1, it indicates partial priority sampling; if ζ = 1, it indicates full priority sampling. Therefore, the probability of experience being sampled is expressed as:
[0214]
[0215] The intelligent agent consists of an Actor network and a Critic network. The Critic network consists of four Q networks, two of which are used to reduce overestimation of actions.
[0216] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention herein is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, the above-mentioned features may have similar functions to (but not limited to) those disclosed in this application.
Claims
1. A method for allocating Internet of Vehicles resources for 5G remote driving, characterized in that: include: Build the allocation model network architecture; The distribution model network architecture includes a vehicle layer, an edge layer, and a center layer; There is 1 vehicle in the vehicle layer; The edge layer includes M edge base stations; The central layer includes the central controller of the 5G remote control system; The vehicle layer is communicatively connected to the edge layer via a wireless network; The central layer is connected to the edge layer through optical fibers; The vehicle layer obtains task information; The vehicle layer preprocesses the task information to obtain task preprocessing output information; the task preprocessing output information includes task category, quantitative indicator information corresponding to the task category, task offloading ratio Vehicle layer delay The vehicle layer calculates the task offloading ratio by using formula 1 for the task category and the quantitative indicator information corresponding to the task category. when When , the vehicle layer completes the task calculation and processing within the standard delay time and outputs the task processing plan information; when When the task is completed, the vehicle layer transmits the task pre-processing output information to the edge layer, which processes it or processes it together with the edge layer and the vehicle layer. The task calculation is completed within the standard delay time and the task processing solution information is output. The task processing plan information is output from the edge layer to the central layer; the central layer transmits the received task processing plan information to the vehicle layer for 5G remote driving.
2. The test scenario construction method for an early warning system according to claim 1, characterized in that: Formula 1 is: The task j generated by vehicle i at time t is where i∈{1, 2, …, I}, j∈{1, 2, …, J}, t∈{1.2, …, T}; is the task offloading ratio, The available computing power on the vehicle side, T is the number of CPU clock cycles required to process each bit of data in the task category; max Maximum local processing delay on the vehicle side.
3. The test scenario construction method for an early warning system according to claim 2, It is characterized in that The vehicle layer delay is obtained by formula 2 based on the task category and the quantitative indicator information corresponding to the task category. The formula 2 is is the task offloading ratio, is the task data size of vehicle layer i, is the CPU frequency of vehicle layer i, The number of CPU clock cycles required to process each bit of data for different task categories.
4. The test scenario construction method for an early warning system according to claim 1, characterized in that: The task categories include emergency obstacle avoidance, extreme weather pre-processing, remote parking, and network fluctuation processing; The task information includes one or more of emergency obstacle avoidance task information, extreme weather pre-processing task information, remote parking task information, and network fluctuation processing task information; Emergency obstacle avoidance mission information includes: obstacles in front of the vehicle lane, videos and pictures of the vehicle in front, and videos and pictures of obstacles in adjacent lanes, vehicles in front, and vehicles behind; Extreme weather pre-processing task information includes: road photos, videos, and weather information; Remote parking task information includes: the horizontal distance between the vehicle's current position and the center point of the target parking space; the yaw angle between the vehicle's heading and the parking space entry direction; The network fluctuation processing task information includes: communication delay, network packet loss rate, and bandwidth utilization.
5. The test scenario construction method for an early warning system according to claim 1, characterized in that: The process of processing task information at the vehicle layer includes the following steps: Obtain quantifiable indicators from task information, compare the quantitative indicators with the preset task trigger threshold to obtain the task category, so that the vehicle layer can obtain the task category and the quantitative indicator information corresponding to the task category; The quantifiable indicators corresponding to the emergency obstacle avoidance task include: the distance between the front obstacle in the ego vehicle's lane and the ego vehicle, the distance between the front vehicle in the ego vehicle's lane and the ego vehicle, the distance between the obstacle in the adjacent lane and the ego vehicle, the distance between the front vehicle in the adjacent lane and the ego vehicle, the distance between the rear vehicle in the adjacent lane and the ego vehicle, the speed and acceleration of the front vehicle in the ego vehicle's lane, the speed and acceleration of the front vehicle in the adjacent lane, and the speed and acceleration of the rear vehicle in the adjacent lane; The quantifiable indicators corresponding to the extreme weather pre-processing task include: light intensity, rainfall rate, and road friction coefficient; The quantifiable indicators corresponding to the remote parking task include: distance and yaw angle of the parking target; The quantifiable indicators corresponding to the network fluctuation processing task include: network delay time and network packet loss rate; The preset task trigger thresholds include emergency obstacle avoidance trigger threshold, extreme weather pre-processing trigger threshold, remote parking trigger threshold, and network fluctuation processing trigger threshold; The emergency obstacle avoidance triggering threshold includes when the distance between the obstacle or the preceding vehicle is less than 5m, or when the relative speed between the vehicle and the preceding vehicle is greater than 2m / s; The extreme weather pre-processing trigger thresholds include starting 5G remote driving when the illumination is less than 50 lux and the rainfall rate is greater than 10 mm / hr, or when the friction coefficient is less than 0.5; The network fluctuation processing trigger threshold includes enabling 5G remote driving when the network delay is greater than 50ms or the packet loss rate is greater than 5%.
6. The test scenario construction method for an early warning system according to claim 2, characterized in that: The processes processed by the edge layer or the edge layer and vehicle layer together include: The edge layer is provided with a task preprocessing output information set, an edge layer state information set, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer state information; The edge layer status information includes the number of tasks processed by the edge layer and the degree of edge layer congestion when the edge layer receives the vehicle layer task preprocessing output information; The edge layer task processing parameters include the resource allocation ratio and task processing delay when the edge layer processes the task preprocessing output information; Get the number of tasks processed by the edge server of edge layer m at time t; The edge layer processes the task preprocessing output information through formula 3, and obtains the blocking degree of the edge server m at time t: When the degree of obstruction When the blocking criteria are met, the edge server of the edge layer m processes the task preprocessing output information; when the blocking degree When the blocking criteria are not met, edge layer m transmits the task preprocessing output information to the adjacent edge layer m+1 or edge layer m-1 through the wireless network, and judges the blocking degree. or Whether the blocking criteria are met, whether the task pre-processing output information can be processed at the edge server of edge layer m+1 or edge layer m-1, and finally determining the edge layer that processes the task pre-processing output information; The edge server of the edge layer m processes the task preprocessing output information including: The optimal edge layer task processing parameters of the edge layer m at time t corresponding to the task preprocessing output information of the vehicle i received by the edge layer m at time t are obtained by comparing the task preprocessing output information of the vehicle i received by the edge layer m at time t with the edge layer state information of the edge layer m at time t and the edge layer task processing optimal parameter set corresponding to the edge layer state information. The edge server of edge layer m processes the task preprocessing output information according to the optimal parameters of edge layer task processing at time t to obtain task processing solution information.
7. The method for constructing a test scenario for an early warning system according to claim 6, characterized in that: The blocking degree of the edge server of edge layer m at time t is is calculated by formula 3, which is: is the maximum amount of data that can be stored in the edge server; i∈I means that vehicle i belongs to set I, and I is the set of all vehicles in the vehicle layer.
8. The test scenario construction method for an early warning system according to claim 1, characterized in that: A SAC-E reinforcement learning model is constructed in the central layer; the central layer obtains a set of task preprocessing output information, a set of edge layer state information, and an optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer state information through offline training and compression of the SAC-E reinforcement learning model; Then, during the non-working hours of the edge layer, the optimal solution set of edge layer task processing parameters corresponding to the task preprocessing output information and the edge layer status information is transmitted to the edge layer, and the edge layer is regularly optimized and updated.
9. The test scenario construction method for an early warning system according to claim 1, characterized in that: The SAC-E reinforcement learning model includes the following: The task is calculated by formula 4 Latency offloaded to edge servers The task of edge layer m to vehicle i is calculated by formula 5 The delay generated is The task is calculated by formula 6 Delay The task of the edge server of edge layer m to vehicle i is calculated by formula 7 Resource consumption costs Based on MEC The vehicle i task is calculated by formula 8 when multiple vehicle output tasks are processed in edge layer m. Resource allocation ratio at edge layer m 10. The test scenario construction method for an early warning system according to claim 9, characterized in that: Formula 4 is: in B is the bandwidth between the vehicle layer and the edge layer, P is the upload power of vehicle i, N0 is the Gaussian white noise power, is the channel gain between the edge layer server and the vehicle layer device, is the distance between the edge service layer m and vehicle i, and s is the path loss index; Formula 5 is: in When multiple vehicle output tasks are processed at the edge layer m, vehicle i task The resource allocation ratio at the edge layer m; F mec is the computing power of the edge server at edge layer m; Formula 6 is: Formula 7 is: Among them, the price per unit of computing power in the edge server is K; Formula 8 is: Among them, μ1 and μ2 are the weight coefficients of delay and resource consumption cost respectively, usually μ1+μ2=1; constraint C1 specifies the offloading ratio Constraint C2 stipulates that the total amount of computing resources allocated to each edge server shall not exceed its maximum computing capacity; constraint C3 stipulates that the total amount of offloaded data received by the edge server shall not exceed its maximum storage capacity.