Calculation unloading and caching joint optimization method and device based on improved A3C algorithm
Through the improved A3C algorithm and maximum entropy concept, the SA3C algorithm is proposed, which solves the problems of calculation offloading and cache optimization in on-board edge computing, realizes efficient resource utilization and priority processing of tasks, and improves the overall performance of the system.
Patent Information
- Application Number
- CN202510201348.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Traditional vehicle-mounted edge computing technology is difficult to effectively solve the problems of computing offloading and cache optimization, especially when resources are limited and task execution order is unconstrained, resulting in RSU overload, low task completion rate and large system delay.
The improved A3C algorithm is adopted and combined with the concept of maximum entropy, and the Soft-A3C algorithm (SA3C) is proposed to optimize the computational offload and cache joint problem. By building V2V and V2R models, using the computing resources of road mobile vehicles and roadside units, formulating task priorities and preemption plans, and optimizing cache strategies.
Effectively reduce server load, reduce resource waste, improve task hit rate, reduce task failure rate, optimize the total average delay of the system, and make the joint optimization of calculation and unloading and cache more in line with the actual scenario.
Smart Images

Figure CN120075897A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle-mounted edge computing, and particularly relates to a joint optimization method and device for computing offloading and caching based on an improved A3C algorithm. Background Art
[0002] With the popularization of 5G and future 6G technologies, various new vehicle-mounted services have emerged continuously, such as intelligent navigation, voice assistants, and driverless driving. Although these emerging mobile services have greatly enriched people's lives, they also occupy huge computing and storage resources of mobile devices. Vehicle-mounted edge computing emerges as an efficient solution. The VEC system allows user vehicles to offload computationally intensive tasks to the RSU equipped with an edge server that is relatively close, thereby reducing latency, alleviating the vehicle burden, and improving efficiency. At the same time, caching technology is widely applied in edge computing to store frequently accessed data, further reducing network transmission latency and improving the user experience. However, due to limited resources and unconstrained task execution order, when transmitting tasks with a large amount of data, problems such as RSU overload, low task completion rate, and large system latency will occur.
[0003] In most vehicle-mounted edge computing research, only the task processing of local vehicles is considered, resulting in resource waste. For emergency tasks that need to be executed in a timely manner, task failures also occur due to poor queue sorting rules, greatly reducing the system efficiency. By utilizing the computing resources of other moving vehicles on the road and formulating reasonable task priorities and task preemption schemes, the above problems can be effectively solved. However, for the optimal solution of such problems, traditional algorithms are difficult to adapt to complex, dynamic, and high-dimensional task offloading and caching optimization problems. For example, classic heuristic algorithms, game theory models, optimization methods, etc. can theoretically provide the optimal solution to optimization problems, but they often assume that the network environment is static or that task requirements are known and will not change, making it difficult for them to handle sudden changes, resulting in relatively fixed decision-making strategies and a lack of flexibility. Deep Reinforcement Learning (DRL) has characteristics such as adaptive ability, long-term reward optimization, no need for an accurate model, and real-time online learning, showing great advantages in task offloading and caching optimization. Through the interaction between the agent and the environment, DRL continuously optimizes the strategy through a trial-and-error mechanism without a clear model, making it particularly suitable for the dynamic and complex scenario of task offloading and caching optimization in MEC. DRL can optimize decisions in real time under these dynamic changes, avoiding the inefficiency or misjudgment caused by static assumptions in traditional methods. For example, Chinese Patent CN119300089A discloses a method and device for mobile edge computing offloading and service caching based on deep reinforcement learning. This application transforms the task offloading and service caching optimization problem into a Markov decision process (MDP) and uses deep reinforcement learning (DRL) technology to explore and optimize the optimal strategy. DRL well solves the challenges in this aspect. It can use the highly variable communication state as the state input and output the computing offloading and service caching strategy through a DNN (deep neural network), and reduce the computational complexity. However, this method has problems such as insufficient exploration, poor policy stability, and difficulties in adapting to dynamic environments, resulting in an inability to fit the actual situation of joint optimization of computing offloading and caching. Summary of the Invention
[0004] To solve the problems existing in the background technology, the present invention proposes a joint optimization method and device for computing offloading and caching based on an improved A3C algorithm. By constructing V2V models and V2R models through multiple roadside units equipped with MEC servers, multiple terminal vehicles, and a cloud server, these V2V models and V2R models, as the target network, can reduce the server load and resource waste; by introducing the concept of maximum entropy, an improved Soft-A3C algorithm, namely the SA3C algorithm, is obtained. Soft-A3C optimizes the policy by introducing entropy regularization, enabling the agent to pay more attention to exploration during training, thus avoiding over-reliance on local optimal solutions and improving the stability and convergence speed of the policy. This soft policy method can better handle complex and dynamic environments, improve learning efficiency, and enhance stability; this approach can be used to solve optimization problems to minimize the total average delay of the system.
[0005] One aspect of the present invention provides a joint optimization method for computing offloading and caching based on an improved A3C algorithm, which is applied to a target network. The target network at least includes multiple roadside units equipped with MEC servers, multiple terminal vehicles, and a cloud server. The method includes:
[0006] Based on the processing delay generated by each terminal vehicle for locally processing the current task in the current time slot and the transmission delay generated by each terminal vehicle for offloading the current task to a corresponding other terminal vehicle in the current time slot, obtain the total delay of all terminal vehicles in the current time slot;
[0007] Based on the transmission delay generated by each roadside unit for receiving the current task from the terminal vehicle in the current time slot and the processing delay of the task during the preemption process in the current time slot for each roadside unit, obtain the total delay of all roadside units in the current time slot;
[0008] Based on the processing delay of each roadside unit or / and each terminal vehicle for the previous task, obtain the waiting delay of the current task in the current time slot;
[0009] Based on the transmission rate at which the cloud server sends the current task to the roadside unit and the task popularity of the current task, obtain the caching delay of the current task in the current time slot;
[0010] Based on the total delay of all tasks, with the goal of minimizing the average total delay, determine the joint optimization problem model of all terminal vehicles and all roadside units;
[0011] Use the improved A3C algorithm to solve the joint optimization problem model to obtain the preemption indication, offloading policy, and caching policy of the computing tasks to be processed in each roadside unit and each terminal vehicle;
[0012] Among them, the policy loss function of the improved A3C algorithm is determined by the sum of the product of the advantage at each time step and the logarithmic probability gradient of the policy network, and the product of the control entropy parameter at each time step and the policy entropy.
[0013] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the joint optimization of computing offloading and caching based on the improved A3C algorithm.
[0014] Another aspect of the present invention provides a joint optimization device for computing offloading and caching based on the improved A3C algorithm, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the joint optimization device for computing offloading and caching based on the improved A3C algorithm executes the joint optimization of computing offloading and caching based on the improved A3C algorithm.
[0015] The present invention has at least the following beneficial effects
[0016] The present invention uses other moving vehicles on the road and RSU equipped with an MEC server to form a V2V model and a V2R model respectively, reducing the server load and resource waste; the present invention formulates task priorities to enable tasks to be executed in an orderly manner, and designs a task preemption scheme to enable current urgent tasks to be processed preferentially, reducing the task failure rate; the present invention is based on the Zipf distribution and uses the cloud server to cache services for RSU according to task popularity, effectively improving the task hit rate; the present invention combines the Asynchronous Advantage Actor Critic (A3C) algorithm based on the Actor-Critic architecture, introduces the concept of maximum entropy, and proposes the Soft-A3C algorithm, i.e., the SA3C algorithm, for solving optimization problems, minimizing the total average delay of the system. The present invention takes into account factors such as the needs of users in the actual scenario, making the joint optimization of computing offloading and caching more in line with the actual situation while realizing intelligence and autonomy. Brief Description of the Drawings
[0017] Figure 1 It is a schematic diagram of a VEC system based on task preemption according to an embodiment of the present invention;
[0018] Figure 2 It is a schematic diagram of the method flow according to an embodiment of the present invention;
[0019] Figure 3 It is a flowchart of the preemption scheme according to an embodiment of the present invention;
[0020] Figure 4This is the SA3C algorithm framework based on the Actor-Critic architecture in the embodiments of the present invention. Detailed implementation manners
[0021] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0022] The method and device for jointly optimizing computing offloading and caching based on the improved A3C algorithm in the embodiments of the present invention are implemented by a target network, and the target network at least includes multiple roadside units equipped with MEC servers, multiple terminal vehicles, and a cloud server.
[0023] Please refer to Figure 1 , exemplarily, the entire target network system is a cloud-edge-end architecture, including a cloud server, three roadside units (RSUs), and multiple vehicle users. Vehicles can communicate with each other and with RSUs. The cloud server can cache services to RSUs. The uplink transmission channel between vehicles and RSUs uses OFDMA technology to eliminate interference between users.
[0024] Among them, the cloud server serves as a remote computing resource center in the target network and can provide large-scale data storage and complex computing services. The roadside units are distributed along the road. As roadside infrastructure, they are responsible for communicating with vehicles, can collect traffic information, send road condition warnings, etc., and can also serve as relay nodes for communication between vehicles and the cloud server or MEC server. The MEC cloud server is connected to the RSU, and there is an MEC server beside each RSU. The MEC server deploys computing and storage resources at the network edge, close to the vehicle, can provide low-latency computing services, reduce the latency caused by vehicle-cloud communication, and can quickly process the data uploaded by the vehicle, such as real-time traffic data analysis, vehicle status monitoring, etc. The red vehicle represents the task request vehicle, that is, the vehicle with computing task requirements; the blue vehicle is an ordinary vehicle that can assist in task processing or communication. In the actual process, the identities of the task request vehicle and the ordinary vehicle can be interchanged, that is, vehicle A can be used as a task request vehicle at time t1, and vehicle A can be used as an ordinary vehicle at time t2. Vehicles communicate directly with each other through the V2V link to achieve information sharing, such as exchanging vehicle speed, position and other information. Vehicles and roadside units can communicate directly through the V2R link within the coverage area of the roadside unit. Vehicles can send task requests, location information, etc. to the RSU through the V2R link, and the RSU can also send traffic rules, road conditions and other information to the vehicles.
[0025] Please refer to Figure 2 , one aspect of the present invention provides a joint optimization method for computing offloading and caching based on an improved A3C algorithm, and the method includes:
[0026] S1. Based on the processing delay generated by each terminal vehicle locally processing the current task in the current time slot and the transmission delay generated by each terminal vehicle offloading the current task to the corresponding other terminal vehicle in the current time slot, obtain the total delay of all terminal vehicles in the current time slot;
[0027] In the embodiment of the present invention, within the current time slot, the delay of the terminal vehicle is determined by the local processing delay and the processing delay on other terminal vehicles. Each terminal vehicle generates a processing delay when locally processing the current task in the current time slot. This depends on factors such as the vehicle's own computing power and the complexity of the task. For example, if the vehicle has a low processor performance, or the task involves a large amount of complex calculations (such as high-definition image recognition, complex path planning calculations, etc.), then the local processing delay will be relatively long. The local processing delay reflects the time spent by the vehicle in processing the task relying on its own resources. When the terminal vehicle selects to offload the current task to the corresponding other terminal vehicle in the current time slot, a transmission delay will be generated. This is mainly related to the communication distance between vehicles, the quality of the communication link (such as signal strength, interference situation, etc.), and the data transmission rate. For example, if the distance between two vehicles is far, or there is strong interference in the communication frequency band, resulting in a decrease in the data transmission rate, then the time required to transmit the task data, that is, the transmission delay, will increase.
[0028] Specifically, in the embodiment of the present invention, the calculation process of the total delay of all terminal vehicles within the current time slot may include the following steps:
[0029] S11: Define vehicle n ∈ {1, 2,..., N}, the speed of vehicle n is v n , the vehicle position is (x n , y n ),
[0030] The speed of the remaining mobile vehicle j is v j , the position is (x j , y j ), the communication range between vehicles is According to the above data, calculate the communication time between vehicles as:
[0031]
[0032] Through the above method, the connection time between vehicles can be obtained. This connection time can be used as the processing delay limit for tasks at the vehicle to ensure that task offloading can be completed at the vehicle.
[0033] Taking each vehicle as a node and the communication time between vehicles as the edge weight, construct a topological path, and use the depth-first search algorithm to find all paths from the target vehicle to which the task needs to be offloaded to the current task vehicle, denoted as the hop count H. Then the single-hop transmission delay of V2V communication is:
[0034]
[0035] Among them, D n is the task size, V n,j is the transmission rate from vehicle n to j, Wn,j is the transmission bandwidth between vehicles n and j, P n,j is the transmission power of vehicle n, h n,j is the channel gain, σ 2 is the noise power. Then the total transmission delay of V2V communication is:
[0036]
[0037] According to the task size D n and the number of CPU cycles X required for the user to calculate 1 bit n , the computing resource D n X n is obtained. Then, according to the vehicle computing power f n , the processing delay of the task at the vehicle is calculated as:
[0038]
[0039] S12: According to the above data, calculate the total local vehicle task processing delay as:
[0040]
[0041] Among them, is the total delay of V2V communication, t n,l is the delay of the vehicle computing task.
[0042] Through the above formula, the delay of each terminal vehicle in the current time slot can be obtained. Assuming there are N terminal vehicles in total, then the delays of these N terminal vehicles can be added up to obtain the total delay of all terminal vehicles in the current time slot.
[0043] S2. Based on the transmission delay generated by each roadside unit receiving the current task from the terminal vehicle in the current time slot and the processing delay of the task during the preemption process of each roadside unit in the current time slot, obtain the total delay of all roadside units in the current time slot;
[0044] In the embodiments of the present invention, the delay of each roadside unit within the current time slot is determined by the transmission delay and the processing delay related to task preemption. Each roadside unit receiving the current task from the terminal vehicle within the current time slot will generate a transmission delay. This transmission delay is affected by various factors. For example, the distance between the terminal vehicle and the roadside unit. The farther the distance, the longer the signal transmission path, and the greater the possible transmission delay. The quality of the communication link. If there are problems such as interference and signal attenuation, it will reduce the data transmission rate and thus increase the transmission delay. And the size of the transmitted data. The larger the task data volume, the longer the transmission time usually is. Whether each roadside unit preempts a subsequent task within the current time slot will generate a processing delay. If the roadside unit preempts within the current time slot, then it needs to process the preempted subsequent task, and the current task waits for it to complete the processing and then continues to be processed. The amount of calculation and complexity of processing the task will determine the length of the processing delay. This effectively ensures that burst emergency tasks can be preferentially processed; if there is no preemption, the current task continues to be processed.
[0045] Specifically, in the embodiments of the present invention, the calculation process of the total delay of all roadside units within the current time slot may include the following steps:
[0046] S21: Define vehicle n ∈ {1, 2, …, N}, roadside unit RSU r ∈ {1, 2, …, R}, the speed of vehicle n is v n , the vehicle position is (x n , y n ), the RSU position is (x r , y r ), the maximum communication distance of the RSU is ψ. According to the above data, calculate the residence time of the vehicle within the RSU coverage range as:
[0047]
[0048] Through the above method, the total delay of a task processed between vehicles cannot exceed the connection time between vehicles. This connection time can be used as the processing delay limit of tasks at the roadside unit to ensure that task offloading can be completed at the roadside unit.
[0049] S22: Calculate the transmission delay of V2R communication:
[0050]
[0051] Among them, D n is the task size, V n,r is the transmission rate from vehicle n to RSU r, W n,r is the transmission bandwidth between vehicle n and RSU r, P n,r is the transmission power of vehicle n, h n,r is the channel gain, σ2 is the noise power. According to the task size D n and the number of CPU cycles X required by the user to calculate 1 bit n , the computing resource D n X n is obtained. Then, according to the computing power f of the roadside unit r , the processing delay of the task on the RSU is calculated as follows:
[0052]
[0053] S23: According to the above data, the total processing delay of the task on the RSU is calculated as follows:
[0054]
[0055] where represents the transmission delay from the vehicle to the RSU, t n,r represents the processing delay of the current task on the RSU, t n+1,r represents the processing delay of the successor task on the RSU, m n ∈{0,1} represents the preemption indication. When m n =1, preemption occurs; otherwise, preemption is not allowed.
[0056] Through the above formula, the delay of each RSU in the current time slot can be obtained. Assuming there are a total of R RSUs, then the delays of these R RSUs can be added up to obtain the total delay of all roadside units in the current time slot.
[0057] S3. Based on the processing delays of the predecessor tasks by each roadside unit and / or each terminal vehicle, obtain the waiting delay of the current task in the current time slot;
[0058] In the embodiment of the present invention, the waiting delay of the current task in the current time slot is mainly determined by the processing delay of the predecessor task by the roadside unit and the processing delay of the predecessor task by the terminal vehicle. When multiple tasks arrive at the roadside unit in sequence, the processing of the predecessor task by the roadside unit will cause a delay. This delay depends on the computing power of the roadside unit, the complexity of the task, and the situation of the task queue, etc. For example, if the computing resources of the roadside unit are limited and there are multiple complex tasks queuing up for processing at the same time, then the processing delay of the predecessor task will be longer, which will in turn affect the waiting time of the subsequent current task. When a vehicle has multiple tasks to be executed in sequence, the processing delay of the predecessor task will also have an impact. If the processor performance of the vehicle is average and the predecessor task involves a large amount of data processing or complex algorithm operations, then the time spent on processing the predecessor task will increase, resulting in a longer waiting time for the current task at the vehicle end.
[0059] Specifically, in the embodiments of the present invention, the calculation process of the waiting delay of the current task in the current time slot may include the following steps:
[0060] S31: Calculate the waiting delay of task n as:
[0061]
[0062] Wherein, represents the processing delay of the predecessor task j.
[0063] Through the above formula, the processing delay of the predecessor task j in the current time slot can be obtained. By superimposing the processing delays of these predecessor tasks, the waiting delay of each current task is obtained.
[0064] S4. Based on the transmission rate at which the cloud server sends the current task to the roadside unit and the task popularity of the current task, obtain the caching delay of the current task in the current time slot;
[0065] In the embodiments of the present invention, the caching delay of the current task in the current time slot is determined by the transmission rate at which the cloud server sends the current task to the roadside unit and the task popularity of the current task. If the network bandwidth is sufficient and there is no congestion, the transmission rate may be relatively high, and the task data can be quickly sent to the roadside unit, and the caching delay may be relatively small accordingly; on the contrary, if the network is congested or the link quality is poor, the transmission rate decreases, and the time for the task data to wait for transmission at the cloud server end becomes longer, and the caching delay will increase. Task popularity reflects the frequency at which the current task is demanded or requested. If a task has a high popularity, it means that there may be more terminal vehicles or other system components that may request this task. In order to meet possible subsequent large-scale requests, the cloud server may cache the task data in advance to reduce subsequent transmission pressure and response time.
[0066] Specifically, in the embodiments of the present invention, the calculation process of the caching delay of the current task in the current time slot may include the following steps:
[0067] S41: Define the Zipf slope ∈, and calculate the task popularity using the Zipf distribution as:
[0068]
[0069] Wherein, ρ is a normalization factor used to ensure that the sum of probabilities is 1, and N f is the number of types of all tasks in the cloud server;
[0070] S42: Calculate the caching delay from the cloud server to the RSU as:
[0071]
[0072] Wherein, βn For the offloading decision, For the caching decision, D n For the task size, p n For the task popularity, R is the transmission rate from the cloud server to the RSU;
[0073] Through the above formula, the caching delay from the cloud server to the RSU can be obtained. By superimposing the caching delay from the cloud server to the RSU, the waiting delay for each current task is obtained.
[0074] In a preferred embodiment of the present invention, the present invention formulates task priorities to enable tasks to be executed in an orderly manner, and designs a task preemption scheme so that current urgent tasks can be processed preferentially, reducing the task failure rate; the method further includes:
[0075] S4A. Determine the first priority constraint according to the normalized constraint of the task processing delay; the first priority constraint is:
[0076]
[0077] Wherein, D n Represents the task size, X n Represents the number of CPU cycles required for the user to calculate 1 bit, f r Represents the computing power of the RSU, Represents the minimum threshold of the processing delay, Represents the maximum threshold of the processing delay; x n,1 Represents the normalized constraint of the processing delay.
[0078] S4B. Determine the second priority constraint according to the normalized constraint of the residence time of the terminal vehicle; the second priority constraint is:
[0079]
[0080] Wherein, Represents the vehicle residence time, Represents the minimum residence time threshold of the vehicle, Represents the maximum residence time threshold of the vehicle; x n,2 Represents the normalized constraint of the vehicle residence time.
[0081] S4C. Determine the third priority constraint according to the normalized constraint of the energy consumption; the third priority constraint is:
[0082]
[0083] Wherein, D n Represents the task size, X n Represents the number of CPU cycles required for the user to calculate 1 bit, f rRepresents the computing power of the RSU, Represents the minimum threshold of processing energy consumption, Represents the maximum threshold of processing energy consumption; x n,3 Represents the normalization constraint of processing energy consumption.
[0084] S4D. Perform weighted summation on the first priority constraint, the second priority constraint, and the third priority constraint to obtain the priority of the task; the calculated value of the task priority is expressed as:
[0085]
[0086] Among them, λ 1 、λ 2 、λ 3 respectively represent the corresponding weights of the three normalization constraints; Represents the calculated value of the task priority.
[0087] S4E. Execute the task according to the priority of the task.
[0088] In the preferred embodiment of the present invention, tasks can be executed according to the priority of the tasks, so that the current urgent tasks can be processed preferentially, reducing the task failure rate.
[0089] As Figure 3 shown, in the embodiment of the present invention, the judgment method for whether preemption occurs in each roadside unit in the current time slot includes:
[0090] S4a. If the waiting delay of a certain successor task does not exceed the threshold, no preemption is performed, and the processing delay without preemption is obtained according to the processing delay of the current task on the roadside unit;
[0091] S4b. If the waiting delay of a certain successor task exceeds the threshold, it is judged whether the processing delay of the successor task is less than or equal to 0.5 times the maximum waiting delay of the current task. If it is less than or equal to, preemption is performed, and the processing delay after preemption is obtained according to the processing delay of the successor task on the roadside unit.
[0092] In the embodiment of the present invention, a task preemption scheme is formulated. When the waiting delay of a certain successor task exceeds the threshold ξ, a judgment is entered. When its processing delay is less than or equal to 0.5 times the maximum waiting delay of the current task
[0093] S5. Based on the total delay of all tasks, with the goal of minimizing the average total delay, determine the joint optimization problem model for all terminal vehicles and all roadside units; in the embodiments of the present invention, calculate the total average delay expression of the communication system according to the system model of the communication system, and use the total average delay of the communication system as the optimization goal; in this embodiment, according to the above-set system model of the communication system, the process of determining the optimization goal and constraints includes:
[0094] S51: According to the above data, calculate the total system delay as:
[0095]
[0096] where, β n represents the offloading decision. When β n = 1, offload to the RSU for processing; otherwise, execute on the vehicle. T n represents the total delay required to execute task n;
[0097] S52: Define the energy consumption coefficient δ, and calculate the processing energy consumption on the vehicle and the RSU respectively:
[0098]
[0099] where, D n is the task size, X n is the number of CPU cycles required for the user to calculate 1 bit, f n is the vehicle computing power, f r is the RSU computing power;
[0100] S53: According to the above data, calculate the total system energy consumption as:
[0101]
[0102] S54: Considering the total average delay of the system, the optimization goal and constraints are:
[0103]
[0104] where, N represents the number of tasks; T n represents the total delay of the nth task; T represents the total system delay, T max represents the maximum delay threshold, C1 represents that the delay cannot exceed the maximum tolerance value; E represents the total system energy consumption, E max represents the maximum energy consumption threshold, C2 represents that the energy consumption cannot exceed the maximum tolerance value; represents the caching decision, D nLet \(T\) denote the task size, \(C\) denote the maximum storage capacity, and \(C_3\) denote that the amount of data cached at an RSU cannot exceed the RSU storage capacity; \(C_4\), \(C_5\), and \(C_6\) respectively denote the caching decision, offloading decision, and task preemption indication; Let \(T_s\) denote the residence time of the vehicle, and \(C_7\) denote that the total delay for a task to be transmitted to an RSU for processing cannot exceed the residence time of the vehicle within the range of the RSU; Let \(T_c\) denote the connection time between vehicles, \(C_8\) denote that the total delay for a task to be processed between vehicles cannot exceed the connection time between vehicles; \(C_9\) denotes that the sum of the corresponding weights of the three normalization constraints is 1.
[0105] It should be noted that if the execution according to task priorities is not considered, then in the above process, the priority constraint \(C_9\) can be not considered.
[0106] S6. Use the improved A3C algorithm to solve the joint optimization problem model, and obtain the preemption indication, offloading strategy, and caching strategy of the computing tasks to be processed in each roadside unit and each terminal vehicle;
[0107] Please refer to Figure 4 to determine the state space, action space, and reward function;
[0108] The state space expression is:
[0109] \(s(\tau)=\{C\) V , \(C\) R \}\)
[0110]
[0111] where \(C\) V denotes the computing resources of the vehicle, and \(C\) R denotes the computing resources of the RSU; \(s(\tau)\) denotes the state space at time \(\tau\);
[0112] The action space expression is:
[0113] \(a(\tau)=\{\beta, m, c\) r \}\)
[0114] \(\beta = \{\beta\) 1 , \(\beta\) 2 , \(\cdots\), \(\beta\) N \}\)
[0115] \(m = \{m\) 1 , \(m\) 2 , \(\cdots\), \(m\) N \}\)
[0116]
[0117] where \(\beta\) denotes the offloading decision, \(m\) denotes the preemption indication, that is, whether preemption occurs, and \(c\)r Denote the caching decision; a(τ) represents the action space at time τ;
[0118] The expression of the reward function is:
[0119]
[0120] where N represents the number of tasks, T represents the total system delay; r(τ) represents the reward function at time τ.
[0121] In this embodiment, the process of using the improved A3C algorithm, that is, the Soft-A3C algorithm, to solve the optimization problem includes:
[0122] S61: Initialize the policy network π(a(τ)|s(τ);θ) and the value network V(s(τ);θ v ), set the entropy regularization coefficient β and the learning rate η; the policy network is used to output the probability distribution of the current action a(τ) according to the current state s(τ), and the value network is used to evaluate the value of the current state s(τ); the current state s(τ) includes the computing resources of the terminal vehicle and the computing resources of the roadside unit; the current action a(τ) includes the preemption indication, offloading strategy, and caching strategy of the computing task to be processed;
[0123] S62: Multiple parallel threads interact with the environment to generate trajectory data (s(τ),a(τ),r(τ),s(τ+1)) and store them in the experience pool; where r(τ) represents the currently obtained reward, and s(τ+1) represents the predicted state;
[0124] S63A: Calculate the cumulative discounted reward using the trajectory data as:
[0125]
[0126] where γ represents the discount factor, k represents the k-th step, r(τ+i) represents the reward function at the τ+i step, and V(s(τ+k);θ V ) represents the value network value function at the τ+k step;
[0127] S63B: Calculate the advantage function by combining the value function as:
[0128]
[0129] S64A: According to the maximum entropy concept, introduce entropy regularization to calculate the policy loss function L π (θ) and its gradient are respectively:
[0130] L π (θ) = logπ(a(τ)|s(τ);θ)A(s(τ),a(τ);θ,θ v) + βH(π(s(τ); θ))
[0131]
[0132] Among them, β represents the hyperparameter that controls the strength of the entropy regularization term, and H(π(s(τ); θ)) represents the policy entropy. The purpose of introducing the entropy term is to encourage policy exploration and thus avoid prematurely falling into local optimal solutions.
[0133] The calculation formula of the policy entropy includes:
[0134]
[0135] Among them, π(a n (τ)|s(τ)) represents the probability output of the policy network for the nth action a n (τ), N represents the number of tasks to be calculated, and each action corresponds to a preemption indication, offloading policy, and caching policy for the task to be calculated.
[0136] It should be noted that the higher the policy entropy value, the more uniform the probability distribution of different actions taken by the policy in this state and the greater the uncertainty; the lower the policy entropy value, the more concentrated the probability distribution and the smaller the uncertainty. In this embodiment, introducing the entropy term into the optimization objective of the policy is equivalent to encouraging the policy to have a higher entropy, that is, to more evenly select different actions while maximizing the reward.
[0137] S64B: According to the TD error, calculate the estimated value loss function and its gradient as:
[0138] L v (θ v ) = (R(τ) - V(s(τ); θ v )) 2
[0139]
[0140] S65A: Use the RMSProp (Root Mean Square Propagation) optimizer to calculate the exponentially weighted average of the gradient for τ steps as:
[0141] g(τ) = αg(τ - 1) + (1 - α)Δθ 2
[0142] Among them, α represents the decay coefficient, and Δθ represents the current gradient of the policy loss function or the estimated value loss function for τ steps;
[0143] S65B: Use the policy gradient and value gradient to update the policy parameter θ and the value network parameter θ v respectively as:
[0144]
[0145] Among them, η represents the learning rate, and ∈′ represents a very small positive number to prevent the denominator from being zero;
[0146] S66: Asynchronously upload the gradients of each thread to the global network, update the global parameters, and finally synchronize them to the local network;
[0147] S67: Repeat the above steps until the training reaches the preset termination condition.
[0148] In summary, the present invention uses other moving vehicles on the road and RSU equipped with an MEC server to form a V2V model and a V2R model respectively, which reduces the server load and resource waste; the present invention formulates task priorities to enable the tasks to be executed in an orderly manner, and designs a task preemption scheme to enable the current urgent tasks to be processed preferentially, reducing the task failure rate; the present invention is based on the Zipf distribution and caches services to the RSU by the cloud server according to the task popularity, effectively improving the task hit rate; the present invention combines the Asynchronous Advantage Actor-Critic (A3C) algorithm based on the Actor-Critic architecture, introduces the concept of maximum entropy, and proposes the Soft-A3C algorithm, i.e., the SA3C algorithm, for solving optimization problems to minimize the total average delay of the system. The present invention considers factors such as the needs of users in the actual scenario, and while realizing intelligence and autonomy, makes the joint optimization of computing offloading and caching more in line with the actual situation.
[0149] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the above-mentioned method for jointly optimizing computing offloading and caching based on an improved A3C algorithm.
[0150] Another aspect of the present invention provides a device for jointly optimizing computing offloading and caching based on an improved A3C algorithm, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the device for jointly optimizing computing offloading and caching based on an improved A3C algorithm executes the above-mentioned method for jointly optimizing computing offloading and caching based on an improved A3C algorithm.
[0151] Specifically, the memory includes various media such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc that can store program codes.
[0152] Preferably, the processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0153] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for joint optimization of computation offloading and cache based on an improved A3C algorithm, characterized in that: Applied to a target network, the target network includes at least a plurality of roadside units equipped with MEC servers, a plurality of terminal vehicles and a cloud server, the method comprising: Based on the processing delay generated by each terminal vehicle locally processing the current task in the current time slot and the transmission delay generated by each terminal vehicle unloading the current task to corresponding other terminal vehicles in the current time slot, the total delay of all terminal vehicles in the current time slot is obtained; Based on the transmission delay generated by each roadside unit receiving the current task from the terminal vehicle in the current time slot and the processing delay of the task in the preemption process of each roadside unit in the current time slot, the total delay of all roadside units in the current time slot is obtained; Based on the processing delay of each roadside unit and / or each terminal vehicle for the previous task, obtaining the waiting delay of the current task in the current time slot; Based on the transmission rate at which the cloud server sends the current task to the roadside unit and the task popularity of the current task, obtaining the cache delay of the current task in the current time slot; Based on the total delay of all tasks, with the goal of minimizing the average total delay, determine the joint optimization problem model of all terminal vehicles and all roadside units; The joint optimization problem model is solved by using an improved A3C algorithm to obtain preemption indication, unloading strategy, and caching strategy of the computing tasks to be processed in each roadside unit and each terminal vehicle; The policy loss function of the improved A3C algorithm is determined by the product of the advantage at each time step and the logarithmic probability gradient of the policy network and the sum of the product of the control entropy parameter at each time step and the policy entropy.
2. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to claim 1 is characterized in that: The transmission channel in the target network adopts Orthogonal Frequency Division Multiple Access.
3. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to claim 1, characterized in that: The method for determining whether each roadside unit has preempted in the current time slot includes: If the waiting delay of a subsequent task does not exceed the threshold, no preemption is performed, and the processing delay without preemption is obtained based on the processing delay of the current task on the roadside unit; If the waiting delay of a successor task exceeds the threshold, it is determined whether the processing delay of the successor task is less than or equal to 0.5 times the maximum waiting delay of the current task. If so, preemption is performed, and the processing delay after preemption is obtained based on the processing delay of the successor task on the roadside unit.
4. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to claim 1, characterized in that: The methods for obtaining the cache delay of the current task in the current time slot include: If a current task is unloaded at a roadside unit and there is no current task in the roadside unit, the current task is cached to the roadside unit through the cloud server; According to the size of the current task, the task popularity of the current task, and the transmission rate from the cloud server to the roadside unit, the cache delay of the current task in the current time slot is calculated; The task popularity is obtained by Zipf distribution.
5. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to any one of claims 1 to 4, characterized in that: The method further comprises: According to the normalized constraint of the task processing delay, a first priority constraint is determined; according to the normalized constraint of the terminal vehicle's stay time, a second priority constraint is determined; According to the energy consumption normalization constraint, a third priority constraint is determined; Performing a weighted summation on the first priority constraint, the second priority constraint, and the third priority constraint to obtain the priority of the task; The tasks are executed according to their priorities.
6. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to claim 5, characterized in that: The joint optimization problem model includes: st C1:0 <T≤T max C2:0<E≤E max C4:c r ∈{0,1} C5:β∈{0,1} C6:m∈{0,1} C9:λ1+λ2+λ3=1 Where N represents the number of tasks; T n represents the total delay of the nth task; T represents the total system delay, T max represents the maximum delay threshold, C1 means the delay cannot exceed the maximum tolerance value; E represents the total energy consumption of the system, E max Indicates the maximum energy consumption threshold, C2 means that the energy consumption cannot exceed the maximum tolerance value; represents the cache decision, D n represents the task size, C represents the maximum storage capacity, C3 represents the amount of data cached at an RSU cannot exceed the RSU storage capacity; C4, C5, and C6 represent the cache decision c respectively. r , offloading decision β, task preemption indication m; Indicates the residence time of the vehicle. C7 indicates that the total delay of transmitting a task to the RSU for processing cannot exceed the residence time of the vehicle within the RSU range; represents the connection time between vehicles, C8 indicates that the total delay of processing a task between vehicles cannot exceed the connection time between vehicles; C9 indicates that the sum of the corresponding weights of the three normalized constraints is 1.
7. The method for joint optimization of computation offloading and cache based on improved A3C algorithm according to claim 1, characterized in that: The improved A3C algorithm is used to solve the joint optimization problem model, and the preemption indication, unloading strategy, and caching strategy of the pending computing tasks in each roadside unit and each terminal vehicle are obtained, including: Initialize the policy network and the value network, set the entropy regularization coefficient β and the learning rate η; the policy network is used to output the probability distribution of the current action a(τ) according to the current state s(τ), and the value network is used to evaluate the value of the current state S(τ); the current state s(τ) includes the computing resources of the terminal vehicle and the computing resources of the roadside unit; the current action a(τ) includes the preemption indication, unloading strategy and cache strategy of the computing task to be processed; Multiple parallel threads interact with the environment to generate trajectory data (s(τ), a(τ), r(τ), s(τ+1)) and store them in the experience pool; where r(τ) represents the current reward and s(τ+1) represents the predicted state; The trajectory data is used to calculate the cumulative discounted reward R(τ), and then the advantage function A(s(τ), a(τ); θ, θ is calculated by combining the value function. v ); Entropy regularization is used to calculate the policy loss function L π (θ) and its gradient, calculate the estimated value loss function and its gradient based on the TD error; Use policy gradient and value gradient to update policy parameters θ and value network parameters θ respectively v ; Upload the gradient of each thread asynchronously to the global network, update the global parameters, and finally synchronize them to the local network; Repeat the above steps until the training reaches the preset termination condition.
8. The method for joint optimization of computation offloading and cache based on the improved A3C algorithm according to claim 7, characterized in that: The calculation formula of the strategy entropy includes: Among them, H(π(s(τ); θ)) represents the policy entropy, π(s(τ); θ) represents the probability output under the current state under the policy network θ, and π(a n (τ)|s(τ)) represents the policy network’s response to the nth action a n (τ), N represents the number of tasks to be calculated, and each action corresponds to the preemption indication, offloading strategy, and caching strategy of a task to be calculated.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the calculation offloading and cache joint optimization method based on the improved A3C algorithm according to any one of claims 1 to 8.
10. A computing offloading and cache joint optimization device based on an improved A3C algorithm, characterized in that: It comprises a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the computing offloading and cache joint optimization device based on the improved A3C algorithm executes the computing offloading and cache joint optimization method based on the improved A3C algorithm as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Mobile edge computing unloading and service caching method and device based on deep reinforcement learning
CN119300089A
Internet of vehicles task unloading method and system and electronic equipment
CN116112525A
Internet-of-vehicles task unloading and content caching method based on cooperation of block chain and edge computing
CN116566838A
MEC task unloading decision-making method based on discrete soft actor-commentator algorithm
CN118870433A
Techniques to shape network traffic for server-based computational storage
US20230403236A1
Cited By
Distributed task unloading and service caching joint optimization method and device
CN122054237A