Cloud-side cooperative task unloading multi-strategy SAC joint optimization method based on electric vehicle networking
By employing the SAC joint optimization method in electric vehicle networking, the problems of task offloading and resource allocation in dynamic network environments of traditional deep reinforcement learning are solved, achieving more efficient task processing and resource utilization, and adapting to random traffic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEILONGJIANG UNIV
- Filing Date
- 2026-03-13
- Publication Date
- 2026-05-01
AI Technical Summary
Existing traditional deep reinforcement learning methods are difficult to adapt to random traffic and dynamic network environments in vehicle-to-everything (V2X) cloud-edge collaborative optimization, resulting in poor task offloading and resource allocation performance.
A cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking is adopted. By constructing a cloud-edge collaborative scenario model of vehicle networking, the SAC deep reinforcement learning algorithm is used for joint online updates to optimize the task offloading ratio, forwarding ratio and bandwidth allocation. Combined with multi-step temporal differential learning and exponential decay noise strategy, it can adapt to dynamic network environment.
It improves the efficiency of task offloading and resource allocation between edge nodes and the cloud, adapts to random traffic and dynamic networks, reduces task processing costs, and improves system operating efficiency.
Smart Images

Figure CN121967483A_ABST
Abstract
Description
A Cloud-Edge Collaborative Task Offloading Multi-Strategy SAC Joint Optimization Method Based on Electric Vehicle Networking Technical Field
[0001] This invention relates to the field of electric vehicle networking technology, and specifically to a cloud-edge collaborative task offloading optimization method based on electric vehicle networking. Background Technology
[0002] With the rapid development of intelligent and connected technologies in electric vehicles, electric vehicles have become important mobile intelligent terminals, generating a large number of computing tasks during operation, such as environmental perception, path planning, and in-vehicle application services. Limited by the computing power, battery capacity, and communication bandwidth of in-vehicle terminals, local execution is insufficient to meet the application requirements of low latency, high reliability, and low energy consumption. Cloud-edge collaborative task offloading has become a key supporting technology for electric vehicle networking systems. Cloud-edge collaborative task offloading refers to dynamically selecting whether to execute in-vehicle computing tasks locally, at edge nodes, or in the cloud based on real-time network status, node load, energy consumption, and cost constraints, and collaboratively allocating communication, computing, and storage resources to improve the overall system operating efficiency.
[0003] When traditional deep reinforcement learning is applied to cloud-edge collaborative optimization in vehicle-to-everything (V2X) networks, it is difficult to adapt to random traffic and dynamic network environments, resulting in poor performance in task offloading, resource allocation, and system optimization between edge nodes and the cloud, thus limiting its practical application. Summary of the Invention
[0004] To overcome the technical limitations of existing cloud-edge collaborative optimization methods for vehicle-to-everything (V2X) networks in random traffic and dynamic network environments, this invention provides a cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking.
[0005] This invention is achieved through the following technical solution:
[0006] A cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking includes the following steps:
[0007] Step S1: Modeling the scenario of offloading tasks in the Internet of Vehicles (IoV) cloud-edge collaborative process, including system architecture design, communication model construction, definition of decision variables, setting of constraints, construction of objective function, and modeling of total task processing delay;
[0008] Step S2: Dynamically generate computing tasks at the vehicle layer, including task type, data size, generation timestamp, remaining battery power, local CPU frequency, and current task load;
[0009] Step S3: The cloud layer collects the task data set, edge server information, and channel status information, and generates a state space vector based on the collected information;
[0010] Step S4: The deep reinforcement learning agent based on the actor critic network solves the state space vector, outputs the initial decision action, and performs a value evaluation on the initial decision action to obtain a value estimate.
[0011] Step S5: Based on the initial decision action, perform tasks to be offloaded from the vehicle layer to the edge layer, tasks to be forwarded from the edge layer to the cloud layer, and link bandwidth allocation operations; the vehicle layer, edge layer, and cloud layer respectively execute the corresponding computing tasks to obtain the task execution results, including total computing power, task transmission latency, energy consumption cost, and task success rate;
[0012] Step S6: Calculate the current time slot based on the task execution result. System processing costs Receive instant rewards Using system state space vectors Initial decision-making actions, immediate rewards Next time slot System state space vector Construct experience samples, store them in the priority experience replay pool, and calculate the initial priority based on the value estimate described in S4;
[0013] When the number of experience samples in the priority experience replay pool reaches the batch size, a small batch of experience samples is sampled from the experience replay buffer, the multi-step temporal difference target value is calculated, the parameters of the decision network and value network are updated by gradient descent, and the target network is softly updated, while the priority of the experience samples is updated.
[0014] The updated network will be used for decision reasoning in the next time slot.
[0015] Furthermore, the decision variables include the proportion of tasks offloaded from the vehicle to the edge layer. The proportion of tasks forwarded from the edge layer to the cloud layer and the bandwidth allocation ratio for each link;
[0016] The proportion of tasks offloaded from the vehicle to the edge layer Represented as:
[0017] ,
[0018] when When this occurs, it means that all computational tasks generated by the vehicle are executed locally at the vehicle layer and are not offloaded to the edge layer; when When this occurs, it indicates that all computational tasks generated by the vehicle are offloaded to the edge layer; when Values When this occurs, it indicates that the computational tasks generated by the vehicle are partially unloaded to the edge layer;
[0019] The task forwarding ratio from the edge layer to the cloud layer Represented as:
[0020] ,
[0021] when When this occurs, it means that all computing tasks at the edge layer are executed locally at the edge layer and are not offloaded to the cloud layer; when When this occurs, it indicates that all computing tasks at the edge layer are offloaded to the cloud layer; when Values When this occurs, it indicates that some computing tasks from the edge layer are being offloaded to the cloud layer.
[0022] Furthermore, taking minimizing task processing cost as the objective function, the formula is:
[0023] ,
[0024] In the formula, The decision variable represents the task offloading from the vehicle layer to the edge layer; Decision variables for offloading tasks from the edge layer to the cloud layer; This indicates the bandwidth allocated to the roadside units in the edge layer; This represents the total number of time slots, i.e., the length of the time window considered in the optimization problem; This represents the time slot index, which is the number of the discrete time unit during system operation; Indicates time slot The system processing cost.
[0025] Furthermore, the system processing cost Including computing costs and communication costs :
[0026] ,
[0027] Calculation cost:
[0028] ,
[0029] In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic computing resource pricing cost of the edge server; This indicates the dynamic pricing cost of computing resources for cloud servers; Indicates roadside unit Vehicle collection within the covered two-way road area In the time slot Total computational load generated within the system; This represents energy consumption costs, including processing energy costs and communication energy costs; Indicates energy trade-off parameters;
[0030] Communication cost:
[0031] ,
[0032] In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic rental cost based on real-time bandwidth contention. This indicates the total available bandwidth for communication between the current edge layer and the cloud layer; Indicates roadside unit bandwidth; Indicates the delay penalty factor; This indicates the task transmission delay.
[0033] Furthermore, the constraints are as follows:
[0034] ,
[0035] ,
[0036] ,
[0037] ,
[0038] ,
[0039] In the formula, Indicates the index of the roadside unit in the edge layer; Indicates the number of roadside units; Indicates in time slot Internal allocation to the first Bandwidth ratio coefficient of each roadside unit; Indicates in time slot No. Each roadside unit and the vehicle collection within the covered two-way road area. Average transmission rate between; This indicates the minimum transmission rate required for the communication link between the vehicle and the roadside unit. Indicates in time slot No. Average transmission rate between each roadside unit and the base station in the cloud layer; This indicates the minimum transmission rate required for the communication link between the roadside unit and the base station; Indicates energy consumption cost; This represents the upper limit threshold for energy consumption costs.
[0040] Furthermore, the edge server information includes: currently available computing resources, CPU load status, task queue length, and dynamic pricing of edge computing resources; the channel status information includes: channel status information between the vehicle unit and the roadside unit, channel information between the roadside unit and the base station, real-time interference level of each channel link, and available bandwidth resources.
[0041] The channel state information between the vehicle-mounted unit and the roadside unit includes channel gain, Rice factor, line-of-sight and non-line-of-sight states; the channel information between the roadside unit and the base station includes channel gain, obstacle probability, shape parameters of the Nakagami-m distribution, and average power.
[0042] Furthermore, the process of outputting the initial decision action in step S4 includes:
[0043] The system state space vector is input into the decision network, which outputs the mean of the action distribution using a Gaussian strategy. The standard deviation of the action exploration noise is adaptively adjusted through an exponential decay mechanism. A Gaussian distribution is constructed based on the mean and the decayed standard deviation, and the initial decision action is sampled from this distribution.
[0044] Furthermore, the formula for calculating the attenuated standard deviation is as follows: In the formula, Indicates the initial noise. Indicates the attenuation coefficient. Indicates minimum noise. Indicates the decay time interval.
[0045] Furthermore, in step S6, a small batch of experience samples is sampled from the experience playback buffer using an adaptive sampling mechanism based on TD error. extraction probability for:
[0046] ,
[0047] ,
[0048] In the formula, Indicates the index number of the experience sample in the priority experience replay buffer; Indicates the first Temporal difference error of an empirical sample; Indicates the value network to the first The spatial state vector of an empirical sample and action vectors Value estimation; It is a very small positive number; Indicates the first Each empirical sample corresponds to Step returns the target value.
[0049] Furthermore, the formula for calculating the multi-step temporal difference target value is as follows:
[0050] ,
[0051] In the formula, Indicates time slot Calculated The step returns the target value, i.e., from the current time slot. Start from the beginning Step-by-step cumulative discount rewards plus Value estimation of the state after the step; Indicates the summation index; Indicates the discount factor; Indicates in time slot The instant reward received; The identifier representing the value network; Indicates will The weighting coefficients for converting post-step value estimates to the current time slot; Represents a value network; This represents the system state n steps ahead from the current time slot; The target policy network is state-based. The generated target action; Represents the output of the value network Post-step state - estimated action value corresponding to the action; This indicates the step size for multi-step temporal difference learning.
[0052] The beneficial effects of this invention are:
[0053] This invention constructs a vehicle-to-everything (V2X) cloud-edge collaborative scenario model, using task offloading ratio, forwarding ratio, and bandwidth allocation as decision variables, and minimizing task processing cost as the objective function. It utilizes a SAC-based deep reinforcement learning algorithm for joint online update optimization, adapting to random traffic and dynamic network environments, and improving task offloading, resource allocation, and system optimization between edge nodes and the cloud. This invention employs an exponentially decaying noise strategy, ensuring global exploration capability in the early training phase and rapid convergence to the optimal strategy in the later training phase, avoiding training oscillations or slow convergence caused by traditional fixed-noise exploration. This invention introduces multi-step temporal difference learning, calculating multi-step cumulative rewards and future state value estimation, making it more suitable for task offloading scenarios in vehicular networks with long time dependencies. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0055] Figure 1 is a schematic diagram of the structure of a cloud-edge collaborative task offloading multi-strategy SAC joint optimization system based on electric vehicle networking in one embodiment of the method of the present invention.
[0056] Figure 2 is a schematic diagram of the convergence performance of MP-SAC in one embodiment of the method of the present invention;
[0057] Figure 3 is a schematic diagram comparing the system overhead of each algorithm under different vehicle number probability distribution scenarios in one embodiment of the method of the present invention.
[0058] Figure 4 is a schematic diagram comparing the system overhead of various algorithms based on the Markov chain task generation model under different numbers of vehicles in one embodiment of the method of the present invention.
[0059] Figure 5 is a schematic diagram comparing the average latency of various algorithms under different numbers of vehicles based on the probability distribution task generation model in one embodiment of the method of the present invention.
[0060] Figure 6 is a schematic diagram comparing the average time delay of each algorithm based on the probability distribution task generation model under different numbers of vehicles at 5kW power conditions in one embodiment of the method of the present invention.
[0061] Figure 7 is a schematic diagram comparing the average latency of various algorithms under different numbers of vehicles based on the Markov chain task generation model in one embodiment of the method of the present invention.
[0062] Figure 8 is a schematic diagram comparing the average time delay of various algorithms based on the Markov chain task generation model under different numbers of vehicles at 5kW power conditions in one embodiment of the method of the present invention.
[0063] Figure 9 is a schematic diagram comparing the task success rates of various algorithms under different numbers of vehicles in one embodiment of the method of the present invention.
[0064] Figure 10 is a comparison of the resource utilization of various algorithms under different numbers of vehicles in one embodiment of the method of the present invention;
[0065] Figure 11 is a comparison of the system overhead of various algorithms based on the Rice and Nakagami-m hybrid fading channel model under different numbers of vehicles in one embodiment of the method of the present invention.
[0066] Figure 12 is a comparison of the average latency of various algorithms based on the Rice and Nakagami-m hybrid fading channel model under different numbers of vehicles in one embodiment of the method of the present invention.
[0067] Figure 13 is a comparison of the average time delay of each algorithm based on the Rice and Nakagami-m hybrid fading channel model in one embodiment of the method of the present invention, under the condition of 5kW power and different number of vehicles.
[0068] Figure 14 is a comparison of the system costs of SAC and MP-SAC under different numbers of vehicles in one embodiment of the method of the present invention;
[0069] Figure 15 is a comparison of the average processing latency of SAC and MP-SAC under different numbers of vehicles in one embodiment of the method of the present invention.
[0070] Figure 16 shows a comparison of the average processing delay of SAC and MP-SAC under different numbers of vehicles in one embodiment of the method of the present invention under 5kW power conditions. Detailed Implementation
[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0072] Specific implementation method one:
[0073] Step S1: Modeling the scenario of vehicle-to-everything (V2X) cloud-edge collaborative task offloading, including V2X system architecture design, communication model construction, decision variable definition, constraint setting, objective function construction, and task total processing delay modeling.
[0074] (1) System architecture design: A three-layer edge-cloud collaborative vehicle network architecture is adopted, including the vehicle layer. Edge layer With cloud layer The vehicle layer includes a two-way road scenario with several electric vehicles traveling on the two-way road; the edge layer includes N roadside units. The system consists of N roadside units and N edge servers, with each group of roadside units and edge servers forming an edge node. The cloud layer includes a central controller, base stations, and intelligent agents. The central controller establishes communication links with the edge nodes through the base stations. Please refer to Figure 1, which is a schematic diagram of the electric vehicle network structure in this embodiment. The bidirectional road in the vehicle layer is evenly divided into N regions, and the edge layer deploys N roadside units (RSUs) and N edge servers. Each of the N regions corresponds one-to-one with the N roadside units. Each roadside unit is only responsible for information interaction with electric vehicles within its designated region, and cross-regional communication is prohibited. Each roadside unit uniquely connects to one edge server, and each edge server uniquely connects to one roadside unit. The roadside units and edge servers are connected by wires. The set of roadside units (RSUs) is defined as follows: , Indicates the first One roadside unit; define the edge server set as... , Indicates the first An edge server.
[0075] The agent is an MP-SAC deep reinforcement learning agent based on SAC improvement, with an actor-critic network as the base network, including a twin value network. , And actors network The value network adopts a three-layer fully connected structure:
[0076] (1)
[0077] In the formula, These represent fully connected networks at layers 3, 2, and 1, respectively. Indicates the activation function; Represents the system state vector; Represents the action vector; , This represents the bias vector.
[0078] (2) Communication model construction, including the communication model between the vehicle layer and the edge layer, the communication model between the edge layer and the cloud layer, the link interference noise between the vehicle layer and the edge layer, and the link interference noise between the edge layer and the cloud layer.
[0079] (2.1) Communication between the vehicle layer and the edge layer is modeled using Ricean distribution, and the channel gain is... for:
[0080] (2)
[0081] in, , where is the Rice factor, representing the power ratio of the line-of-sight component to the multipath component; Let be a Gaussian random variable. When the distance between the vehicle layer and the edge layer is less than or equal to a preset line-of-sight transmission distance threshold, the channel is determined to be in line-of-sight mode. The value should be between 10dB and 20dB; otherwise, the channel is considered to be in a non-line-of-sight state. The value ranges from 0dB to 3dB.
[0082] (2.2) Communication between the edge layer and the cloud layer adopts the Nakagami-m distribution combined with the obstacle probability model, and the channel gain... It follows a Nakagami-m distribution:
[0083] (3)
[0084] in, The shape parameter determines the severity and depth of channel fading; Average power represents the statistical expectation of the average power or energy of the received signal.
[0085] The obstacle probability is determined by an exponential model. The probability that the communication link between the roadside unit in the edge layer and the base station in the cloud layer is blocked by an obstacle can be expressed as:
[0086] (4)
[0087] In the formula, Indicates roadside unit Transmission distance to base station The larger the value, the greater the probability of obstacles appearing. This represents the obstacle density coefficient or shading intensity coefficient, reflecting the density of buildings or the concentration of obstacles in an urban environment. The larger the value, the higher the probability of encountering an obstacle per unit distance. To determine whether the current channel is in line-of-sight or non-line-of-sight mode, specifically, when a random number... If the channel enters a non-line-of-sight state and follows a Nakagami-m distribution, it is determined that the channel remains in a line-of-sight state and follows a Rice distribution with weaker fading.
[0088] (2.3) Link interference noise between vehicles in the vehicle layer and roadside units in the edge layer The formula is:
[0089] (5)
[0090] Indicates time slot Roadside Unit A collection of vehicles within the covered two-way road area; This indicates the vehicle's transmission power during communication, i.e., the signal transmission strength.
[0091] (2.4) The link between the roadside unit in the edge layer and the base station in the cloud layer is in a non-line-of-sight state, and the interference noise is high. The formula is:
[0092] (6)
[0093] in, Indicates roadside unit The signal transmission power; Indicates from roadside unit The channel gain to the base station is the attenuation coefficient in a non-line-of-sight environment, reflecting the impact of obstacles on the signal; It is the attenuation coefficient, which characterizes the exponential nature of path loss; Indicates Gaussian white noise; Indicates the additional attenuation factor of the non-line-of-sight propagation model. Represents the natural constant. (3) Definition of decision variables: including the proportion of vehicle task unloading to the edge layer. The proportion of tasks forwarded from the edge layer to the cloud layer And the bandwidth allocation ratio for each link.
[0094] The proportion of tasks offloaded from the vehicle to the edge layer Represented as:
[0095] (7)
[0096] when When this occurs, it means that all computational tasks generated by the vehicle are executed locally at the vehicle layer and are not offloaded to the edge layer; when When this occurs, it indicates that all computational tasks generated by the vehicle are offloaded to the edge layer; when Values When this occurs, it indicates that the computational tasks generated by the vehicle are being offloaded to the edge layer.
[0097] The task forwarding ratio from the edge layer to the cloud layer Represented as:
[0098] (8)
[0099] When decision variables When this occurs, it means that all computational tasks at the edge layer are executed locally at the edge layer and are not offloaded to the cloud layer; when the decision variable When this occurs, it indicates that all computing tasks at the edge layer are offloaded to the cloud layer; when the decision variable... Values When this occurs, it indicates that some computing tasks from the edge layer are being offloaded to the cloud layer.
[0100] (4) Objective function and constraints: The objective function is to minimize the task processing cost. The formula for the objective function is:
[0101] (9)
[0102] In the formula, The decision variable represents the task offloading from the vehicle layer to the edge layer; Decision variables for offloading tasks from the edge layer to the cloud layer; This indicates the bandwidth allocated to the roadside units in the edge layer; This represents the total number of time slots, i.e., the length of the time window considered in the optimization problem; This represents the time slot index, which is the number of the discrete time unit during system operation; Indicates time slot The system processing cost, including computing cost. and communication costs The calculation formula is:
[0103] (10)
[0104] Calculation cost:
[0105] (11)
[0106] In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic computing resource pricing cost of the edge server; This indicates the dynamic pricing cost of computing resources for cloud servers; Indicates roadside unit Vehicle collection within the covered two-way road area In the time slot Total computational load generated within the system; This represents energy consumption costs, including processing energy costs and communication energy costs; This represents the energy consumption trade-off parameter.
[0107] Communication cost:
[0108] (12)
[0109] In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic rental cost based on real-time bandwidth contention. This indicates the total available bandwidth for communication between the current edge layer and the cloud layer; Indicates roadside unit bandwidth; Indicates the delay penalty factor; This indicates the task transmission delay.
[0110] The constraints include: decision variables for task offloading from the vehicle layer to the edge layer. The range of values for the value and the decision variables for task offloading from the edge layer to the cloud layer. The constraints include the range of values for the value, the range of bandwidth allocated to the roadside units in the edge layer, the average transmission rate between the vehicle layer and the edge layer, the average transmission rate between the edge layer and the cloud layer, and the energy consumption cost constraint.
[0111] The formula for calculating the average transmission rate between the vehicle layer and the edge layer is as follows:
[0112] (13)
[0113] In the formula, Indicates in time slot No. Each roadside unit and the vehicle collection within the covered two-way road area. Average transmission rate between; Indicates the first The bandwidth of each roadside unit; Indicates from vehicle to number The channel gain of each roadside unit reflects the combined effect of path loss and antenna gain during signal transmission.
[0114] The formula for calculating the average transmission rate constraint between the edge layer and the cloud layer is as follows:
[0115] (14)
[0116] In the formula, Indicates in time slot No. Average transmission rate between each roadside unit and the base station in the cloud layer; Indicates in time slot Internal allocation to the first The bandwidth ratio coefficient of each roadside unit, ranging from 0 to 1, is used to control the bandwidth share occupied by that roadside unit in the backhaul link. This represents the total available bandwidth of the backhaul link between the roadside unit and the base station in time slot t; This indicates the transmission power of the roadside unit.
[0117] The constraint condition is expressed as:
[0118] (15)
[0119] (16)
[0120] (17)
[0121] (18)
[0122] (19)
[0123] In the formula, Indicates energy consumption cost; This represents the upper limit threshold for energy consumption costs; , These represent the minimum transmission rate required for the communication link between the vehicle and the roadside unit, and the minimum transmission rate required for the communication link between the roadside unit and the base station, respectively.
[0124] (5) Modeling of total task processing delay:
[0125] (20)
[0126] This indicates the total processing latency of the vehicle's task load; This represents the average local latency of electric vehicles at the vehicle layer; Indicates the total edge computation latency; Represents a set of vehicles With roadside units The average transmission delay; Indicates roadside unit Average transmission delay with base station.
[0127] The average local latency of the electric vehicles in the vehicle layer The calculation formula is:
[0128] (twenty one)
[0129] In the formula, This represents the number of CPU cycles required per unit of computational task. Indicates roadside unit Vehicle collection within the covered two-way road area In the time slot Total computational load generated within the system; Indicates the vehicle is in a time slot Average dynamic CPU frequency within the range, This reflects the change in the vehicle's actual available computing power over time.
[0130] The total edge calculation latency The calculation formula is:
[0131] (twenty two)
[0132] In the formula, This indicates the currently available computing resources of the edge server; This indicates the queuing delay.
[0133] The vehicle collection With roadside units Average transmission delay The calculation formula is:
[0134] (twenty three)
[0135] The roadside unit Average transmission delay with base station The calculation formula is:
[0136] (twenty four)
[0137] The vehicle in the time slot Average dynamic CPU frequency within , This is the vehicle's initial maximum CPU frequency. It is an adjustment factor based on the battery state. This is the maximum central processing unit frequency supported by the vehicle hardware. The battery state adjustment factor. It can be represented as:
[0138] (25)
[0139] In the formula, Indicates that the vehicle is within the time slot The remaining battery power, when the battery power is sufficient ( When the CPU frequency is unrestricted, it remains at its maximum value. When the battery power is low ( When the battery is fully charged, the CPU frequency gradually decreases as power decreases to save energy. This segmented adjustment strategy ensures both computing performance when the battery is fully charged and energy optimization when the battery is low, effectively extending the vehicle's driving range while ensuring the execution of critical tasks.
[0140] The roadside unit Vehicle collection within the covered two-way road area In the time slot Total computational load generated The calculation formula is:
[0141] (26)
[0142] In the formula, This represents the index of computational task types, which include low-latency real-time control tasks, medium-latency environmental awareness tasks, and high-latency data upload tasks. Indicates the first The average computational cost of a class of computational tasks; Indicates the first The generation probability of the class computation task, and satisfies ; Indicates time slot Inner road side unit The number of electric vehicles in the corresponding area.
[0143] The task generation frequency is simulated using a Poisson distribution, where the average arrival rate of the k-th type of task is... :
[0144] (27)
[0145] In consecutive time slots, the task types generated by each electric vehicle follow a Markov switching mechanism, switching according to a preset probability, i.e., with a first preset probability. Keeping the current task type unchanged, with the second preset probability The system randomly switches to other task types. Based on the current task type after the switch, three specific computational tasks are generated according to the corresponding probability distribution: low-latency real-time control tasks, medium-latency environmental perception tasks, and high-latency data upload tasks. The task size of each type follows a normal distribution, reflecting the randomness of task size.
[0146] The task type state switching process can be represented as:
[0147] (28)
[0148] In the formula, Indicates time slot The task type status generated at that time, Indicates electric vehicles in time slots The generated task type status, Indicates task type With time slot Generated task types The same probability, Indicates task type With task type Different probabilities. In this embodiment, , .
[0149] In the time slot The generation probability of the k-th type of task Dynamically adjust based on the current task type status:
[0150] (29)
[0151] The average computational cost It follows a normal distribution and is always positive. Its expression is:
[0152] (30)
[0153] In the formula, These represent the mean and standard deviation, respectively.
[0154] Average vehicle speed in each two-way road zone , The maximum speed limit allowed on the road; This indicates the average number of vehicles in the current area; and These represent the upper and lower limits of the number of vehicles that the area can accommodate in the traffic design. The adjustment factor, ranging from 0 to 1, is used to characterize the effect of low-density traffic flow on overall vehicle speed. When the number of vehicles is small, this factor can alleviate the potential overestimation of speed due to the reduced number of vehicles, and better reflect the characteristics of actual traffic where vehicles are still limited by factors such as road structure and driver behavior even when traffic is flowing smoothly.
[0155] Step S2: Vehicle dynamic generation calculation task.
[0156] Each electric vehicle in the vehicle layer is equipped with an onboard unit for generating task data sets in real time. These task data sets include computational tasks (task type, data size, generation timestamp) and vehicle status information (remaining battery power, local CPU frequency, current task load). After generation, the task data sets are transmitted to the roadside units in the edge layer.
[0157] The rules for classifying task types include:
[0158] Set the data size of the computation task to The required response time is ,
[0159] when and At that time, the computing task type is divided into low-latency real-time control tasks;
[0160] when and At that time, the computational task is divided into medium-latency environmental perception tasks;
[0161] when and At that time, the computing task is divided into high-latency data upload tasks.
[0162] , These represent the first preset threshold. Second preset threshold ; , These represent the first preset response time. Second preset response time In this embodiment, , , , The values are 100kb, 1MB, 10ms, and 100ms.
[0163] Step S3: The cloud layer collects the task data set, edge server information, and channel status information, and generates a state space vector based on the collected information.
[0164] The central controller in the cloud layer establishes communication links with each edge node in the edge layer through base stations, and collects task datasets, edge server information and channel status information in real time.
[0165] Edge server information includes: currently available computing resources, CPU load status, task queue length, and dynamic pricing of edge computing resources.
[0166] Channel state information includes: channel state information between vehicle-mounted units and roadside units (channel gain, Rice factor, line-of-sight and non-line-of-sight status), channel information between roadside units and base stations (channel gain, obstacle probability, shape parameters of Nakagami-m distribution and average power), real-time interference levels and available bandwidth resources for each channel link.
[0167] The central controller uses the collected data to construct the current time slot. The system state space vector. The high-dimensional continuous system state space vector contains key state information of the vehicle layer, edge layer, and cloud layer in the current time slot.
[0168] Step S4: The deep reinforcement learning agent based on the actor critic network solves the state space vector, outputs the initial decision action, and performs a value evaluation on the initial decision action to obtain a value estimate.
[0169] Step S41: Input the system state space vector into the policy network. The policy network uses a Gaussian policy to output the mean of the action distribution. Adaptively adjust the standard deviation of the action exploration noise through an exponential decay mechanism. Construct a Gaussian distribution based on the mean and the decayed standard deviation, and sample the initial decision actions from this distribution. Initial decision-making action It includes three continuous variables: the vehicle layer task offloading ratio decision variable, the edge layer task forwarding ratio decision variable, and the bandwidth allocation ratio variable for each link.
[0170] The formula for calculating the attenuated standard deviation is:
[0171] (31)
[0172] In the formula, Indicates the initial noise. Indicates the attenuation coefficient. Indicates minimum noise. This indicates the attenuation time interval. In this embodiment, , , , .
[0173] Step S42: Connect the state space vector with the initial decision action. The initial decision action is obtained by inputting the value network in the actor-critic algorithm structure. The corresponding action value estimate.
[0174] Step S5: Based on the initial decision action, perform tasks to be offloaded from the vehicle layer to the edge layer, tasks to be forwarded from the edge layer to the cloud layer, and link bandwidth allocation operations; the vehicle layer, edge layer, and cloud layer respectively execute the corresponding computing tasks to obtain the task execution results, including total computing power, task transmission latency, energy consumption cost, and task success rate.
[0175] Within each area of the vehicle layer, vehicles execute a unified local unloading operation based on the vehicle layer task unloading decision variable in the decision instruction: when the vehicle layer task unloading decision variable is 0, all tasks are processed by the electric vehicle's local system; when the vehicle layer task unloading decision variable is 1, all tasks are uploaded to the roadside unit; when the vehicle layer task unloading decision variable is between 0 and 1, the part of the computing task data that cannot be handled by the local computing capacity is uploaded to the corresponding roadside unit through the V2R link.
[0176] Each edge server in the edge layer processes the received tasks according to the edge layer task offloading decision variable in the decision instruction: when the server has sufficient available computing resources, all tasks are processed at the edge layer; when the server load is too high, a corresponding proportion of task data is forwarded to the cloud server via the R2B link. When the edge layer task offloading decision variable is 0, all computing tasks in the edge layer are executed locally at the edge layer and are not offloaded to the cloud layer; when the edge layer task offloading decision variable is 1, all computing tasks in the edge layer are offloaded to the cloud layer; when the edge layer task offloading decision variable is between 0 and 1, the portion of computing tasks that cannot be handled by local computing capacity are offloaded to the cloud layer.
[0177] Each node adjusts its communication resource usage according to the bandwidth allocation ratio variables for each link in the decision instruction. These nodes include server nodes capable of performing computing tasks in the vehicle layer, edge layer, and cloud layer.
[0178] The vehicle-mounted terminal, edge server, and cloud central server each execute their respective computing tasks and return the task execution results to the corresponding vehicles along the original path. The task execution results include the total computing load, task transmission latency, energy consumption cost, and task success rate.
[0179] Step S6: Calculate the current time slot based on the task execution result. System processing costs Receive instant rewards Using system state space vectors Initial decision-making actions, immediate rewards Next time slot System state space vector Construct experience samples, store them in the priority experience replay pool, and calculate the initial priority based on the value estimate described in S4;
[0180] When the number of experience samples in the priority experience replay pool reaches the batch size, a small batch of experience samples is sampled from the experience replay buffer, the multi-step temporal difference target value is calculated, the parameters of the decision network and value network are updated by gradient descent, and the target network is softly updated, while the priority of the experience samples is updated.
[0181] The updated network will be used for decision reasoning in the next time slot.
[0182] This invention addresses the problem of low utilization of key samples in the priority experience replay pool by employing an adaptive sampling mechanism based on TD error, which utilizes experience samples... The probability of being drawn is:
[0183] (32)
[0184] (33)
[0185] In the formula, Indicates the index number of the experience sample in the priority experience replay buffer; Indicates the first The temporal difference error of each empirical sample is used to measure the difference between the current value network estimate and the multi-step target value. The larger the absolute value of the temporal difference error, the higher the priority. Value network for the first The spatial state vector of an empirical sample and action vectors Value estimation; It is a very small positive number; Indicates the first Each empirical sample corresponds to Step returns the target value.
[0186] The formula for calculating the multi-step time-difference target value is as follows:
[0187] (34)
[0188] In the formula, Indicates time slot Calculated The step returns the target value, i.e., from the current time slot. Start from the beginning Step-by-step cumulative discount rewards plus Value estimation of the state after the step; Indicates the summation index; Indicates the discount factor; Indicates in time slot The instant reward received; An index representing a twin value network; Indicates will The weighting coefficients for converting post-step value estimates to the current time slot; Represents a value network; This represents the system state n steps ahead from the current time slot; Represents the target policy network; This represents the number of bootstrapping steps, i.e., the step size of multi-step temporal difference learning, which controls the mixing ratio of actual rewards and estimated values. express.
[0189] Please refer to Table 1 for the specific values of each parameter in this embodiment and experiment.
[0190] To verify the beneficial effects of the present invention, the following simulation experiments were conducted:
[0191] The MP-SAC algorithm proposed in this invention is implemented through the collaboration of vehicles, edge servers, and cloud platforms. Algorithms compared to MPSAC include: DDPG (Deep Deterministic Policy Gradient): a gradient-based policy optimization method that uses deep neural networks to represent deterministic policies; TD3 (Double Delay DDPG): an improved version of DDPG that reduces overestimation and improves stability by using a double Q-network, delayed policy updates, and target policy smoothing; and PPO-RNN (RNN-based Nearest Neighbor Policy Optimization): a neighbor policy optimization method combining recurrent neural networks suitable for partially observed environments, and improves policy learning by using policy pruning and modeling time dependencies.
[0192] The ablation experiment algorithm compared to this invention is the SAC (Soft Actor-Critic) algorithm, a reinforcement learning algorithm based on the maximum entropy framework, which combines stochastic policy optimization and a dual-Q network.
[0193] The simulation scenario is a 400-meter-long two-way road with roadside units deployed on one side. The road is divided into four equally spaced zones, each covered by one RSU. The number of vehicles in each zone ranges from 5 to 40. Each electric vehicle has a maximum battery capacity of 50 kWh, and the vehicle speed ranges from 15 to 28 m / s. Tasks are divided into three types, with different tasks generated according to specific probabilities. The algorithm runs in a Python 3.11 and PyTorch 2.0.1 environment. The algorithm training rounds are set to 400 rounds, with each round containing 300 time steps (slots). The random seed for Python, NumPy, and PyTorch is fixed at 42; the actor network, twin critic network, and temperature parameter all use a learning rate of 0.05. The Adam optimizer is used; the mini-batch sample size is 512; the target network is updated after each gradient step via a soft update mechanism; the actor network and each critic network are implemented using a multilayer perceptron with two 256-unit fully connected hidden layers and ReLU activation functions. More specific vehicle-to-everything (V2X) environment parameters are detailed in Table 1.
[0194] Figure 2 shows the convergence performance of SAC, DDPG, PPO-RNN, and TD3. It can be seen that in the early stages of training, due to the random initialization of all neural network parameters, the task offloading decision and resource allocation strategies perform poorly, resulting in a relatively high average cost for the computational tasks. As the number of training rounds increases, these parameters begin to update in the direction of minimizing the system's average cost, thus the convergence curve shows a downward trend. After a certain number of training rounds, the actor network and the critic network tend to converge, and the average weighted cost reaches its minimum. It can be seen that the algorithm of this invention outperforms other comparative algorithms in terms of convergence speed.
[0195] Figure 3 shows the processing cost of the task generation model under different vehicle quantity probability distribution scenarios. It can be seen that the MP-SAC algorithm of this invention achieves the lowest system cost under different vehicle quantities, which verifies that the method of this invention is superior to other traditional task offloading and resource allocation methods.
[0196] Figure 4 shows the system costs of different algorithms in the Markov chain-based task generation model as the number of vehicles increases from 10 to 40. The MP-SAC of this invention maintains the lowest system cost across all vehicle scales, demonstrating excellent adaptability and optimization capabilities, especially under heavy load conditions.
[0197] Figure 5 illustrates the average processing latency performance of different algorithms in the probabilistic task generation model under varying vehicle numbers. As the number of vehicles increases, communication interference intensifies and bandwidth allocation decreases, leading to increased latency. When the number of vehicles exceeds 30, the latency of TD3, DDPG, and PPO-RNN rises rapidly, with DDPG approaching 800ms at 40 vehicles. In contrast, the algorithm MP-SAC of this invention maintains low latency across all vehicle scales, demonstrating superior scalability under high load conditions. MP-SAC can adaptively adjust offloading strategies, effectively avoiding interference, optimizing resource allocation, and mitigating latency issues caused by bandwidth limitations, exhibiting stronger robustness and latency optimization capabilities.
[0198] Figure 6 compares the latency performance of various algorithms under two scenarios: insufficient battery power (5 kWh) and sufficient battery power (30 kWh). Experimental results show that as the number of vehicles increases, all algorithms experience a significant increase in latency under low battery conditions. The fundamental mechanism lies in the adjustment factor. When the battery power is below 20%, this factor constrains the local CPU frequency, leading to increased local computation latency and forcing more tasks to be offloaded to RSU or the cloud. Due to V2R interference and bandwidth bottlenecks, the system becomes rapidly congested when the number of vehicles exceeds 15, making the latency curve steeper than that in Figure 4. With multi-step temporal difference learning, priority experience replay, and flexible exploration strategies, MP-SAC can dynamically optimize task offloading decisions. Under the same conditions, even in low battery scenarios, MP-SAC maintains a lower average processing latency compared to TD3, DDPG, and PPO-RNN, demonstrating the robustness advantage of the proposed method in energy-constrained environments.
[0199] Figure 7 shows the average processing latency of four algorithms as the number of vehicles increases from 10 to 40 in the Markov chain-based task generation model. The latency increases sharply with the traffic load. DDPG exhibits the highest latency, especially when the number of vehicles exceeds 30, followed by TD3 and PPO-RNN. MP-SAC consistently maintains the lowest latency, demonstrating its stronger adaptability and scalability in high-density scenarios, effectively mitigating bandwidth contention and interference.
[0200] Figure 8 shows the average latency of TD3, DDPG, PPO-RNN, and MP-SAC in the Markov chain-based task generation model under 5kW power conditions as the number of vehicles increases from 10 to 40. The overall latency increases sharply with the increase of traffic load. DDPG exhibits the highest latency, followed by TD3 and PPO-RNN. MP-SAC consistently achieves the lowest latency, especially under heavy load conditions, demonstrating the stronger robustness and more efficient offloading decision-making capability of the proposed method in low-power scenarios.
[0201] Figure 9 compares the success rates of the five algorithms for tasks with 10 to 40 vehicles. As the number of vehicles increases, the success rates of all algorithms gradually decrease due to increased network load and intensified resource contention. MP-SAC consistently maintains the highest success rate, demonstrating strong stability even under heavy traffic loads. SAC and PPO-RNN have moderate performance, while TD3 shows a rapid decline. DDPG exhibits the lowest success rate, particularly noticeable when the number of vehicles exceeds 25, indicating its weak adaptability in high-density environments.
[0202] Figure 10 compares the resource utilization of five algorithms across a range of 10 to 40 vehicles. As the number of vehicles increases, the resource utilization of all algorithms steadily rises. The proposed method, MP-SAC, consistently maintains the highest utilization, indicating more efficient scheduling and fuller utilization of available computing resources. SAC and PPO-RNN follow closely with moderate utilization, while TD3 is slightly lower. DDPG exhibits the lowest utilization across all scenarios, reflecting its weak adaptability and low resource allocation efficiency when network load increases.
[0203] Figure 11 compares the system overhead of four algorithms in the Rice and Nakagami-m hybrid fading channel model as the number of vehicles increases from 10 to 40. As traffic load increases, the system cost of all algorithms steadily rises. DDPG incurs the highest overhead, followed by PPO-RNN and TD3. MP-SAC consistently maintains the lowest system cost, demonstrating stronger adaptability and more efficient task offloading capabilities in complex fading environments.
[0204] Figure 12 shows a comparison of the average latency of four algorithms in the Rice and Nakagami-m hybrid fading channel model as the number of vehicles increases from 10 to 40. As the number of vehicles increases, the latency of all algorithms shows an upward trend. MP-SAC maintains the lowest latency across all vehicle numbers, demonstrating superior transmission efficiency. DDPG, TD3, and PPO-RNN exhibit relatively weak latency management capabilities in fading channels.
[0205] Figure 13 compares the average latency of four algorithms (TD3, DDPG, PPO-RNN, and MP-SAC) based on the Rice and Nakagami-m hybrid fading channel model under a 5kW power condition as the number of vehicles increases from 10 to 40. As the number of vehicles increases, the latency of all algorithms shows an increasing trend. MP-SAC consistently maintains the lowest latency, while DDPG experiences the highest latency across all vehicle numbers.
[0206] To verify the innovation of the MP-SAC algorithm, the following ablation experiments were conducted. Figure 14 compares the system costs of MP-SAC and SAC under different vehicle scales (10–40 vehicles). The results show that as the number of vehicles increases, MP-SAC consistently achieves better cost control than SAC, and the cost difference widens significantly in the 20–40 vehicle range. This advantage stems from three key improvements: multi-step temporal difference learning reduces the estimation error of latency rewards in the vehicle-to-everything (V2X) environment; priority experience replay improves the learning efficiency for high-value samples; and the flexible exploration strategy enables dynamic adjustment of action perturbations, helping to avoid local optima. Combined with an edge-cloud collaborative architecture, MP-SAC continuously monitors vehicle battery status, task load, and network conditions to dynamically optimize task offloading and resource scheduling, effectively reducing computational and communication energy consumption while maintaining processing performance. Experimental results verify that in high-density scenarios with intensified resource competition, this algorithm achieves a balance between system performance and energy cost through precise policy updates and multi-dimensional environmental awareness, demonstrating excellent adaptability and policy robustness.
[0207] Figure 15 compares the processing latency under different vehicle sizes (10–40 vehicles), verifying the advantages of the MP-SAC algorithm in electric vehicle networks. Experiments show that as the number of vehicles increases, MP-SAC consistently maintains lower latency compared to SAC, and the performance gap widens significantly when the number of vehicles reaches or exceeds 25. This advantage stems from its multi-step temporal differential learning mechanism, which improves the accuracy of latency reward estimation, and its priority experience replay mechanism, which enhances the ability to identify high latency risks. Combined with an edge-cloud collaborative architecture, the algorithm achieves intelligent dynamic scheduling of local and remote computing resources by comprehensively optimizing RSU bandwidth, V2R interference, and dynamic CPU frequency. Results show that MP-SAC, through precise task offloading decisions and real-time resource coordination, exhibits superior policy adaptability and system responsiveness in complex, high-density vehicle networking scenarios, effectively ensuring the provision of low-latency services.
[0208] Figure 16 compares the average processing latency of SAC and MP-SAC algorithms under different vehicle scales when the battery power is low (5 kWh). Compared to the scenario in Figure 7 (30 kWh), the system latency increases significantly, and when the number of vehicles exceeds 25, the latency of both algorithms shows an accelerated upward trend. This is because when the electric vehicle's battery power is limited, the dynamic adjustment factor reduces the local CPU frequency, forcing more tasks to be offloaded to the RSU or the cloud, exacerbating the computing and communication congestion of the edge network. Although MP-SAC also experiences latency growth in low-battery scenarios, it still maintains a performance advantage over SAC. This is attributed to MP-SAC's multi-dimensional environmental awareness capabilities (including battery status, task load, bandwidth resources, etc.), which allows it to dynamically adjust the task offloading priority strategy, effectively mitigating latency accumulation by prioritizing latency-sensitive tasks and avoiding network hotspots. Experimental results show that MP-SAC not only performs well in resource-rich scenarios but also maintains superior response performance under extreme conditions of limited energy and high load through adaptive strategies, verifying the algorithm's robustness and collaborative optimization capabilities in complex dynamic environments.
[0209] Table 1 Parameter Settings
[0210]
[0211] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking, characterized in that, include: Step S1: Modeling the scenario of offloading tasks in the Internet of Vehicles (IoV) cloud-edge collaborative process, including system architecture design, communication model construction, definition of decision variables, setting of constraints, construction of objective function, and modeling of total task processing delay; Step S2: The vehicle layer dynamically generates computational tasks, including task type, data size, generation timestamp, remaining battery power, local CPU frequency, and current task load. Step S3: The cloud layer collects the task data set, edge server information, and channel state information, and generates a state space vector based on the collected information. Step S4: A deep reinforcement learning agent based on an actor-critic network solves the state space vector, outputs an initial decision action, and evaluates the value of the initial decision action to obtain a value estimate. Step S5: Based on the initial decision action, the vehicle layer offloads tasks to the edge layer, the edge layer forwards tasks to the cloud layer, and allocates link bandwidth. The vehicle layer, edge layer, and cloud layer each execute their corresponding computational tasks, obtaining task execution results, including total computational load, task transmission delay, energy consumption cost, and task success rate. Step S6: The current time slot is calculated based on the task execution results. System processing costs Receive instant rewards Using system state space vectors Initial decision-making actions, immediate rewards Next time slot System state space vector Construct experience samples, store them in the priority experience replay pool, and calculate the initial priority based on the value estimate described in S4; when the number of experience samples in the priority experience replay pool reaches the batch size, sample a small batch of experience samples from the experience replay buffer, calculate the multi-step temporal difference target value, update the parameters of the decision network and value network through the gradient descent method, perform a soft update on the target network, and update the priority of the experience samples at the same time. The updated network will be used for decision reasoning in the next time slot.
2. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 1, characterized in that, The decision variables include the proportion of tasks offloaded from vehicles to the edge layer. The proportion of tasks forwarded from the edge layer to the cloud layer and the bandwidth allocation ratio of each link; the proportion of tasks offloaded from the vehicle to the edge layer. Represented as: ,when When this occurs, it means that all computational tasks generated by the vehicle are executed locally at the vehicle layer and are not offloaded to the edge layer. when When this occurs, it indicates that all computational tasks generated by the vehicle are offloaded to the edge layer; when Values At this time, it indicates that the computational tasks generated by the vehicle are partially offloaded to the edge layer; the proportion of tasks forwarded from the edge layer to the cloud layer. Represented as: ,when When this occurs, it means that all computing tasks at the edge layer are executed locally at the edge layer and are not offloaded to the cloud layer; when When this occurs, it indicates that all computing tasks at the edge layer are offloaded to the cloud layer; when Values When this occurs, it indicates that some computing tasks from the edge layer are being offloaded to the cloud layer.
3. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 2, characterized in that, With minimizing task processing cost as the objective function, the formula is: In the formula, The decision variables represent the process of offloading tasks from the vehicle layer to the edge layer. Decision variables for offloading tasks from the edge layer to the cloud layer; This indicates the bandwidth allocated to the roadside units in the edge layer; This represents the total number of time slots, i.e., the length of the time window considered in the optimization problem; This represents the time slot index, which is the number of the discrete time unit during system operation; Indicates time slot The system processing cost.
4. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 3, characterized in that, The system processing cost Including computing costs Japanese correspondence book : Calculate the cost: In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic computing resource pricing cost of the edge server; This indicates the dynamic pricing cost of computing resources for cloud servers; Indicates roadside unit Vehicle collection within the covered two-way road area In the time slot Total computational load generated within the system; This represents energy consumption costs, including processing energy costs and communication energy costs; Indicates energy consumption trade-off parameters; communication cost: In the formula, Indicates roadside unit Vehicle collection within the covered two-way road area Task unloading decision variables; This represents the dynamic rental cost based on real-time bandwidth contention. This indicates the total available bandwidth for communication between the current edge layer and the cloud layer; Indicates roadside unit bandwidth; Indicates the delay penalty factor; This indicates the task transmission delay.
5. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 4, characterized in that, The constraints are: , , , , In the formula, Indicates the index of the roadside unit in the edge layer; Indicates the number of roadside units; Indicates in time slot Internal allocation to the first Bandwidth ratio coefficient of each roadside unit; Indicates in time slot The Each roadside unit and the vehicle collection within the covered two-way road area. Average transmission rate between; This indicates the minimum transmission rate required for the communication link between the vehicle and the roadside unit. Indicates in time slot The Average transmission rate between each roadside unit and the base station in the cloud layer; This indicates the minimum transmission rate required for the communication link between the roadside unit and the base station; Indicates energy consumption cost; This represents the upper limit threshold for energy consumption costs.
6. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 5, characterized in that, Edge server information includes: currently available computing resources, CPU load status, task queue length, and dynamic pricing of edge computing resources; channel status information includes: channel status information between the vehicle-mounted unit and the roadside unit, channel information between the roadside unit and the base station, real-time interference level and available bandwidth resources of each channel link; the channel status information between the vehicle-mounted unit and the roadside unit includes channel gain, Rice factor, and line-of-sight and non-line-of-sight status; the channel information between the roadside unit and the base station includes channel gain, obstacle probability, shape parameters of the Nakagami-m distribution, and average power.
7. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 6, characterized in that, Step S4 describes the process of outputting the initial decision action, which includes: inputting the system state space vector into the decision network, the decision network using a Gaussian strategy to output the mean of the action distribution; adaptively adjusting the standard deviation of the action exploration noise through an exponential decay mechanism; constructing a Gaussian distribution based on the mean and the decayed standard deviation, and sampling the initial decision action from this distribution.
8. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking according to claim 7, characterized in that, The formula for calculating the attenuated standard deviation is: In the formula, Indicates the initial noise. Indicates the attenuation coefficient. Indicates minimum noise. Indicates the decay time interval.
9. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking as described in claim 8, characterized in that, Step S6 involves sampling a small batch of empirical samples from the empirical playback buffer using an adaptive sampling mechanism based on TD error. The probability of being drawn is: , In the formula, Indicates the index number of the experience sample in the priority experience replay buffer; Indicates the first Temporal difference error of an empirical sample; Indicate the value network for the first The spatial state vector of an empirical sample and action vectors Value estimation; It is a very small positive number; Indicates the first Each empirical sample corresponds to Step returns the target value.
10. The cloud-edge collaborative task offloading multi-strategy SAC joint optimization method based on electric vehicle networking according to claim 9, characterized in that, The formula for calculating the multi-step time-difference target value is as follows: In the formula, Indicates time slot Calculated The step returns the target value, i.e., from the current time slot. Start from the beginning Step-by-step cumulative discount rewards plus Value estimation of the state after the step; Indicates the summation index; Indicates the discount factor; Indicates in time slot The instant reward received; The identifier representing the value network; Indicates will The weighting coefficients for converting post-step value estimates to the current time slot; Represents a value network; This represents the system state n steps ahead from the current time slot; The target policy network is state-based. The generated target action; Represents the output of the value network Post-step state - estimated action value corresponding to the action; This indicates the step size for multi-step temporal difference learning.