A task offloading method for electric vehicles
Through the edge-end system offloading framework and multi-objective near-end strategy optimization algorithm, edge servers and electric vehicle resources are used to solve the problem of limited computing resources of electric vehicles, and the task offloading efficiency and cost reduction are achieved.
Patent Information
- Application Number
- CN202510815902.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Due to limited computing resources, electric vehicles are difficult to effectively deal with the needs of low latency and low cost task offloading. The existing technology has failed to effectively solve the problem of limited resources of vehicle computing equipment, resulting in inefficient task offloading.
Design an edge-end system offload framework, utilize edge servers and electric vehicles to train agents through multi-objective near-end strategy optimization algorithm (MOPPO), dynamically optimize task offload strategies, considering time delay, energy consumption and payment costs.
It reduces the total cost of the system, improves task offload efficiency, reduces terminal load pressure, and achieves multi-objective optimization in time delay, energy consumption and payment costs.
Smart Images

Figure CN120343028B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communication networks, and in particular to a task offloading method for electric vehicles. Background Art
[0002] With the development of the automotive industry, many vehicles are equipped with features such as autonomous driving and entertainment audio and video, which generate tasks that require low latency and low cost. Due to the limited resources of various computing devices equipped in vehicles, it is necessary to offload tedious tasks to edge servers or idle computing resources in nearby vehicles to alleviate the local execution load of the vehicle.
[0003] By simulating the task offloading process in the connected vehicle network, the tasked vehicle is considered an intelligent agent, which makes decisions based on its local observation data (such as network status, task type, and current load). Through a deep reinforcement learning model, the agent can obtain rewards through interaction with the environment, learn optimal offloading strategies, and dynamically adjust task allocation. This algorithm not only optimizes traditional performance metrics (such as latency and energy consumption), but also considers the needs and preferences of vehicles and systems to provide more personalized and intelligent offloading strategies.
[0004] The present application is compared with the prior art as follows:
[0005] Technical comparison with the application document China Publication No. CN117749797A "Differentially private task offloading method based on Markov chain in the Internet of Vehicles environment";
[0006] 1. Application text China Publication No. CN117749797A discloses a differential privacy task offloading method based on Markov chain in the Internet of Vehicles environment. The method includes constructing and initializing the discrete Markov state space of each vehicle; updating the state transition probability matrix according to the current vehicle's moving speed, the number of surrounding vehicles and the number of edge servers to obtain the privacy parameters of the current vehicle; performing local differential privacy protection on the position of the current vehicle and generating the obfuscated position of the current vehicle; calculating the delay generated by the transmission task between the current vehicle and the edge server, and constructing a minimized system task offloading delay objective function; using the whale algorithm to process the system task offloading delay objective function, searching for the optimal offloading solution, and performing differential privacy task offloading. The present invention reduces the risk of data leakage, and provides an effective and flexible solution for privacy protection in the Internet of Vehicles while ensuring the Internet of Vehicles task offloading delay. The focus is on the Internet of Vehicles and privacy protection.
[0007] 2. The application text proposes a solution to a task offloading algorithm for electric vehicles in the Internet of Vehicles. The algorithm first designs an offloading model for the Internet of Vehicles, taking into account the time delay, energy consumption, and payment cost of the task vehicle. Then, the problem of writing the model of the Internet of Vehicles system is formalized, and the edge unloading scenario of the task vehicle is modeled as an intelligent agent interaction scenario. The proximal strategy is used to optimize the reinforcement learning algorithm to ensure the stability of training and reduce conflicts. The task vehicle is regarded as an intelligent agent, and the unloading decision of the task vehicle is trained through the MOPPO algorithm. The algorithm allows the intelligent agent to learn in a simulation environment, outputs a better unloading strategy through training data, and minimizes the time delay, energy consumption and payment cost generated by the execution of tasks in the system. This application focuses on the unloading strategy obtained by multi-objective optimization in the Internet of Vehicles environment.
[0008] There are essential differences between the two in terms of scheduling objectives, system structure and technical routes.
[0009] Technical comparison with the application document China Publication No. CN113810878A "A macro base station placement method based on vehicle network task offloading decision";
[0010] 1. Application text China Publication No. CN113810878A discloses a macro base station placement method based on the Internet of Vehicles task offloading decision, specifically: Step 1: Establish Y*Y coding matrices and combine the rows and columns, Step 2: Establish a digital twin network and simulate each combination; Step 3: Calculate the optimal task offloading decision for each macro base station in each combination, Step 4: Calculate the total energy consumption of each macro base station under each combination; Step 5: Establish the minimum objective function of total energy consumption and use the particle swarm algorithm to solve it; thereby obtaining the optimal combination. This application considers the placement problem of macro base stations in edge computing, and on the premise of meeting the task computing requirements, reduces the idle time of edge servers in macro base stations and further reduces energy consumption. The focus is on the deployment location of base stations and task offloading efficiency in the Internet of Vehicles.
[0011] 2. The application text proposes a multi-objective task offloading optimization algorithm for electric vehicle Internet of Vehicles scenarios. The solution first constructs an Internet of Vehicles task offloading model that includes multiple dimensions such as time delay, energy consumption, and payment cost, and transforms the complex edge computing offloading problem into a quantifiable optimization problem. Based on the reinforcement learning framework, the multi-objective proximal policy optimization (MOPPO) algorithm is innovatively adopted to model each task vehicle as an independent intelligent agent, and dynamically optimize the offloading decision through continuous interactive learning between the intelligent agent and the simulation environment. The algorithm focuses on solving the problem of unstable training in traditional methods, effectively reducing decision conflicts, and ultimately outputs a task offloading strategy that is balanced and optimized in multiple objective dimensions such as time delay, energy consumption, and cost. This application focuses on obtaining an offloading strategy through multi-objective optimization in the Internet of Vehicles environment.
[0012] There are essential differences between the two in terms of system architecture, scheduling goals and technical paths. Summary of the Invention
[0013] The present invention proposes a task offloading method for electric vehicles. In the scenario of electric vehicle movement, it takes into account multiple complex factors such as time delay, energy consumption, payment cost, etc., reduces the total cost of the system, improves system efficiency and reduces terminal load pressure.
[0014] To achieve the above object, the technical solution adopted by the present invention is:
[0015] A task offloading method for electric vehicles includes the following specific steps:
[0016] Step 1: Considering the vehicle mobility scenario, design an edge-to-end system offloading framework;
[0017] Step 2: The mission vehicle generates intensive tasks and designs the communication model, energy model, time model, and payment cost model;
[0018] Step 3: Based on the problem model, design the system objective function and optimize the system time delay, energy consumption and payment cost with multiple objectives;
[0019] Step 4: Design the vehicle-agent interaction environment based on the system model and reinforcement learning method;
[0020] Step 5: Design a multi-objective proximal strategy optimization algorithm and determine the unloading target after training the agent.
[0021] As a further improvement of the present invention, in step 1, the system model of the edge-end system offloading framework is divided into three parts: electric vehicles, mission vehicles, and edge servers;
[0022] The edge-end system offloading framework has a three-layer architecture;
[0023] The first layer is the edge server, which quickly processes offloaded tasks and provides services for intensive tasks;
[0024] The second layer is electric vehicles, which are composed of vehicles parked around the road and have idle computing resources for mission vehicles to choose to unload;
[0025] The third layer is the mission vehicle. Mission vehicles generate data to be processed and are equipped with small computing devices. Intensive tasks need to be offloaded to edge servers or electric vehicles. Tasks with different service quality requirements need to be offloaded based on the surrounding computing resources.
[0026] In step 1, in the vehicle mobility scenario of the Internet of Vehicles environment, considering the mobility of the vehicle, the vehicle needs to complete the task within the coverage service range of the edge server. When the vehicle user cannot process all the tasks, he needs to offload the data to the edge server or electric vehicle to meet his own task requirements. The edge server, electric vehicle and task vehicle are defined as follows:
[0027] (1) Edge Server:
[0028] The edge server has corresponding computing resources and can provide offloading services for intensive tasks. It can also formulate payment strategies to obtain benefits based on the multi-faceted needs of the task vehicles. The edge server set is defined as , define the number of edge servers as ;
[0029] (2) Electric vehicles:
[0030] The task vehicle selects the computing resources of idle electric vehicles to complete the unloading and execution of the task, and charges the task vehicle for the unloading service provided through the formulated payment strategy to encourage more vehicles to provide unloading services. The electric vehicle set is defined ;
[0031] (3) Mission vehicles:
[0032] The mission vehicle will generate tasks locally and offload tasks that cannot be met by local computing resources to edge servers or electric vehicles. It will select the offloading target based on the various requirements of the tasks to be offloaded and pay the fees according to the payment strategy, defining the time slot of each mission vehicle. Both The task needs to be unloaded, and the task set generated by the task vehicle is defined as , Submit bid for mission vehicle, For the corresponding task, vehicle unloading task By tuple composition, Indicates vehicle The size of the task data to be executed; Indicates the computing resources required for task processing per unit size of data; The unloading target represents the unloading indicator. The task vehicle unloads the task request to the edge server. , among which When the task is processed locally; when When it is unloaded to the corresponding server; when When , it means unloading to the corresponding electric vehicle. To indicate the time sensitivity of the task, the larger the value, the higher the completion time requirement of the task. Indicates the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task Reduced the quality of service required of mission vehicles;
[0033] The mission vehicle needs to complete the unloading within the communication coverage of the service. The execution time of the mission unloading should be within the expected completion time threshold of the mission, and a task can only be unloaded to one unloading target. If the mission vehicle chooses to execute on the local computing device, the unloading target ; If the task is offloaded to the edge server, the target ; If the task is offloaded to the electric vehicle, then the target , where s is the edge server and For electric vehicles.
[0034] As a further improvement of the present invention, the communication model designed in step 2 is as follows:
[0035] The free space path loss model is used to describe the process of unloading the task from the mission vehicle to the edge server, ignoring the delay caused by the downlink transmission. It is assumed that for different mobile vehicles, orthogonal frequency division multiple access is used, and the channel bandwidth of each mobile vehicle is B. The channel interference between mobile vehicles is ignored. According to the Shannon formula, the mobile vehicle unloads the task from the vehicle. The uplink transmission rate offloaded to the edge server s is Similarly, the mobile vehicle can be used to complete the task Unloading to electric vehicles The uplink transmission rate is ,in represents the transmission power of the mission vehicle, and represent the mission vehicles and edge servers or electric vehicles, respectively The channel gain between the two channels is inversely proportional to the distance due to the free space path loss model. and They represent the variance of Gaussian white noise in the corresponding cases.
[0036] As a further improvement of the present invention, the time model in step 2 is as follows:
[0037] After obtaining the uplink transmission rate between the mobile vehicle and the server or electric vehicle, without considering the speed loss during the transmission process, the task vehicle unloads the task Transmission time delay to edge server s and electric vehicle g and They are expressed as follows:
[0038] ;
[0039] ;
[0040] Indicates vehicle The size of the task data to be executed, represents the transmission power of the mission vehicle, B is the channel bandwidth owned by each mobile vehicle, Unloading the vehicle for the mobile vehicle The uplink transmission rate offloaded to the edge server s, Unloading the vehicle for the mobile vehicle Unloading to electric vehicles The uplink transmission rate is and Respectively represent the variance of Gaussian white noise in the corresponding cases;
[0041] The total completion time of a task consists of transmission time and execution time. According to the model, the task vehicle can offload part of the task to the edge server, electric vehicle, or local execution. Due to different offloading targets, the execution time of the task will be calculated according to different offloading targets;
[0042] Define the computing resource set of each terminal computing device , where the computing device subscript 0 represents the local computing resource size of the task vehicle, the subscript between 1 and s represents the computing resource size of the edge server computing device, and the subscript between s+1 and s+g represents the computing resource size of the electric vehicle. The execution time under different offloading targets is given below:
[0043] (1) The mission vehicle performs the mission locally, and the local execution time is given by the following formula:
[0044] ;
[0045] in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the size of the local computing resources of the mission vehicle, To uninstall the target;
[0046] (2) When the task vehicle offloads the task to the edge server or electric vehicle, it needs to be transmitted through the communication link. Compared with local execution, there will be a transmission delay. The execution time of task offloading is obtained by the following formula:
[0047] ;
[0048] in Indicates the computing resource size of the edge server computing device, represents the computing resource size of the electric vehicle, and Represents the task vehicle unloading task The transmission time delay to the edge server s and the electric vehicle g.
[0049] As a further improvement of the present invention, the energy consumption model in step 2 is as follows:
[0050] When the vehicle performs tasks locally The energy consumption model in step 2 is as follows:
[0051] When the vehicle performs the vehicle unloading task locally When the energy consumption is less than the task energy consumption threshold set by the vehicle, the vehicle unloading task performed by the vehicle is defined. The energy consumption threshold is ;
[0052] (1) Transmission energy consumption:
[0053] The transmission energy consumption of the mission vehicle comes from the process of transmitting mission data to the edge server or electric vehicle through the wireless network. During this process, the mission vehicle needs to adjust the transmission power according to the transmission distance, signal strength and network bandwidth factors. When the mission vehicle chooses to offload the task to the edge server or electric vehicle, the task transmission delay is obtained. The transmission power of the mission vehicle is defined as , the energy consumed by offloading the task to the edge server is: ; The energy consumed when offloading the task to the electric vehicle is: ,in and They are the transmission delay of offloading tasks to edge servers and electric vehicles, represents the transmission power of the mission vehicle;
[0054] (2) Local energy consumption:
[0055] When the resources required for task processing are small, the network conditions are not ideal, or the task processing latency requirement is high, the task vehicle chooses to perform the task locally. The task vehicle needs to make a trade-off between computing resource consumption and network transmission consumption and choose the most appropriate processing method. The power consumption of the computing device of the task vehicle is defined as follows: , the local execution time of the task is ,Therefore, the energy consumption of the task that needs to be processed locally can be obtained by the following formula: ;
[0056] (3) Energy consumption of electric vehicles:
[0057] When the task vehicle chooses to offload the task to the electric vehicle, the electric vehicle is responsible for processing the offloaded task. The computing energy consumption of the electric vehicle is related to its processing power and the complexity of the task. Electric vehicles are usually in a low-power state and have relatively limited computing resources. The power consumption of the computing device of the electric vehicle is defined as follows: ,in , the energy consumed by the electric vehicle computing device when processing tasks is: ,in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the computing resource size of the electric vehicle;
[0058] (4) Edge server energy consumption:
[0059] Edge computing servers are deployed at the edge of the network, close to user devices. The power consumption of the computing equipment of the edge server depends on the performance of the server and the complexity of the task. When making task offloading decisions, it is necessary to consider the computing requirements of the task and the load capacity of the server. If the edge server load is too high, it will lead to increased energy consumption and affect the response time of the task. The energy required by the edge server to process the task can be calculated by the following formula: ,in represents the computing device power consumption of the edge server, Indicates the computing resource size of the edge server computing device.
[0060] As a further improvement of the present invention, the payment cost model in step 2 is set as follows:
[0061] After the vehicle offloads the task to the edge server or electric vehicle, the edge server or electric vehicle will charge the task vehicle a corresponding fee based on its own payment strategy. This fee serves as the edge server's or electric vehicle's revenue and also as one of the task vehicle's costs.
[0062] Edge servers and electric vehicles charge fees for the memory usage, computing resource requirements, and security costs of tasks generated by task vehicles to ensure their own operating costs. Assuming that edge servers and electric vehicles are priced according to their own computing resources and charged according to the amount of data offloaded by the task, the specific calculation method for the fee that task vehicle m needs to pay for offloading tasks is as follows:
[0063] ;
[0064] in It represents the price charged by the edge server for processing each unit of data. represents the price charged by the electric vehicle for processing each unit of data. Assuming that the fee that the task vehicle needs to pay is related to the data size of the task and the sensitivity of the task time delay, the total fee that the task vehicle m needs to pay for obtaining the service is calculated as follows:
[0065] ;
[0066] Indicates vehicle The size of the task data to be executed, represents the time sensitivity of the task.
[0067] As a further improvement of the present invention, in step 3, based on the problem model, the system objective function is designed to obtain the task offloading cost of the task vehicle, which is composed of the fees paid by the task vehicle, the time delay of the task vehicle offloading the task, and the energy consumption. Since the data offloaded remotely by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, the task offloading problem is formalized to obtain the following optimization problem:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] in represents the payment cost of the mission vehicle, represents the mission time delay of the mission vehicle, e represents the energy consumption of the entire system, Indicates the maximum time delay threshold of the task, Indicates the task's uninstall indicator, is the local execution time of the task, is the execution time of the task offload, is the transmission delay of offloading the task to the electric vehicle, The energy consumption required to process the task locally, The energy consumed by offloading tasks to edge servers, The energy required to offload the task to the electric vehicle;
[0076] Constraint 1 indicates that the local execution time of the task should be less than the delay threshold of the task vehicle. Constraints 2 and 3 indicate that the time delay of offloading the task to the edge server or electric vehicle should be less than the expected time delay threshold of the task. Constraint 4 indicates that the offloading target limit of the task is ; Constraints 5 and 6 indicate that the energy consumption of local execution should be less than the mission expected energy consumption threshold set by the mission vehicle.
[0077] As a further improvement of the present invention, in step 4, the proximal strategy is used to optimize the reinforcement learning method, constraining the amplitude of the policy update to maximize the agent's cumulative reward while ensuring training stability. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by limiting the amplitude of the policy update;
[0078] MOPPO combines the Actor-Critic architecture, uses the actor network to generate action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, thereby optimizing the direction of policy update;
[0079] The task vehicle can be regarded as an intelligent agent. The unloading strategy of the task vehicle is trained by the MOPPO algorithm and described by a Markov decision process represented by the four-tuple S, A, R, P, where S is the state space, A is the action space, R is the reward function, and P is the state transition probability.
[0080] State Space:
[0081] It is necessary to consider task requirements, network conditions and available computing resources, and define the state of time slot k as , unloading tasks by vehicles of task tuples and the vehicle edge computing environment, where represents the size of the task data that vehicle m needs to process, Represents the various computing resources required for the task, To express the time delay sensitivity of the task, represents the expected time delay threshold of the task, represents the expected energy consumption threshold of the task, and F represents the computing resource set of the task vehicle's surrounding environment;
[0082] Action Space:
[0083] In the unloading scenario, after receiving feedback from the state space, the task vehicle will make a decision on the unloading target. The unloading decision is designed as a discrete vector , in the alternative model Subscript the task's offloading target;
[0084] Reward function:
[0085] In each time slot k, the agent receives a reward when it reaches the next state after executing the action. In general, the reward function should be positively correlated with the objective function. Since the objective function of the problem is to minimize the time delay of the task vehicle, energy consumption, and payment cost, the reward function is negatively correlated with the task vehicle cost. According to the problem model described in the previous section, the reward function can be designed as follows, where Represents the weight of each target parameter:
[0086] ;
[0087] in represents the payment cost of the mission vehicle, represents the mission time delay of the mission vehicle, and e represents the energy consumption of the entire system.
[0088] As a further improvement of the present invention, in step 5, a multi-objective proximal policy optimization algorithm is designed, Actor-Critic architecture: MOPPO adopts Actor-Critic architecture, the Actor network is responsible for generating strategies, the Critic network is responsible for evaluating the state value function, the Actor network outputs the probability distribution of actions, and the Critic network estimates the value of the current state. , used to calculate the advantage function , that is, the probability distribution of the output action given a specific state. The Actor network decides the action based on the current state and gradually improves the strategy by optimizing the objective function, thereby guiding the strategy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean square error of the value function, and the advantage function The calculation formula is as follows:
[0089] ;
[0090] in represents the discount factor, To receive the reward, Represents state k+1;
[0091] Importance sampling mechanism:
[0092] Importance sampling is a method of adjusting data weights. The MOPPO algorithm uses importance sampling to use data in the chronological order of the results of the interaction between the agent and the environment.
[0093] Strategy Optimization:
[0094] The core of the MOPPO algorithm is to optimize the strategy parameters To maximize the expected cumulative reward, the MOPPO algorithm introduces proximal constraints to limit the magnitude of policy updates and uses a clipping objective function to ensure that the difference between the new policy and the old policy is not too large. The clipping objective function is defined as follows:
[0095] ;
[0096] in, The probability of the new strategy to the old strategy is the strategy update ratio, which is defined as the current strategy in the state Select Action Probability The probability of choosing the same state as the old strategy If the ratio is less than or greater than , then crop it to and between, among is the advantage function, is the clipping range parameter.
[0097] As a further improvement of the present invention, in step 5, the agent interaction environment is trained, and the training process is as follows:
[0098] First, the input data required by the algorithm is initialized, including the tasks that need to be unloaded by the vehicle, the computing requirements, delay requirements and energy limits of the tasks, the edge nodes involved in the task vehicle unloading calculation, and the output of the task unloading results; as well as the cost of task unloading, including energy consumption and time delay; clarify the scale of the task set and its multi-dimensional performance requirements, which include computing size, energy consumption and delay sensitivity. The initialization strategy network is used to generate the task unloading action strategy, the value function network evaluates the state value to guide the strategy optimization, and the experience replay buffer stores historical interaction data. The interaction data includes old state, action, reward, and new state. The initial state is obtained at each time sequence, including various attributes of the task, the resource status of the edge server and the collaborative ability of the vehicle cluster. Based on the current state , policy network Output Action , select the target server for offloading or local execution, optimize the action probability distribution through policy gradient, calculate the cost, execution time and energy consumption of the task in this time sequence, observe the new state after executing the action, and calculate the immediate reward , calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , the algorithm line 10 converts the quad Store in buffer D for subsequent policy updates and calculation of generalized advantage function , quantify the pros and cons of actions relative to the average policy, estimate the state value difference through the critic network, solve the credit allocation problem under sparse rewards, enhance the accuracy of policy gradient estimation, and calculate the objective function according to the definition of the clipping objective function , used to update the Actor network and the Critic network, control the changes in the strategy to ensure the stability of training, balance the exploration and utilization of the intelligent agent, calculate the loss value of the Actor network and the Critic network, and then update the network parameters of the Actor network according to the importance sampling results , the critic network updates parameters through the generalized advantage function GAE .
[0099] Compared with the prior art, the present invention has the following beneficial effects:
[0100] In the problem of task offloading of task vehicles in the Internet of Vehicles, the present invention reduces system cost, time delay, energy consumption and payment cost by utilizing computing resources around the task vehicles, such as edge servers and electric vehicle computing devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 This is a diagram of the architecture model of the uninstallation system of the present invention;
[0102] Figure 2 It is the algorithm framework diagram of the present invention;
[0103] Figure 3 This is a flow chart of the MOPPO algorithm of the present invention. DETAILED DESCRIPTION
[0104] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0105] This embodiment provides a task offloading method for electric vehicles, comprising the following steps:
[0106] In step 1, the system model is as follows Figure 1The system can be divided into three parts: electric vehicles, mission vehicles, and edge servers. A three-tier architecture is then designed. The first tier is the edge server, which has powerful computing equipment that can quickly process offloaded tasks and provide services for intensive tasks. The second tier is the electric vehicles, which are composed of vehicles parked around the road. They have a large amount of idle computing resources that can be selected for offloading by mission vehicles, but their computing resources are relatively small compared to edge servers. The third tier is the mission vehicles, which generate data to be processed and are equipped with small computing equipment. Intensive tasks need to be offloaded to edge servers or electric vehicles. Tasks with different service quality requirements need to be offloaded based on the surrounding computing resources.
[0107] In the connected vehicle (IoV) environment, due to vehicle mobility and real-time requirements, the computing power and resources of the onboard units (OBUs) in smart vehicles vary, and server computing resources fluctuate with each time slot. Given the high mobility of vehicles, they must complete tasks within the coverage area of edge servers. When tasks cannot be fully processed, vehicle users need to offload data to edge servers or electric vehicles to meet their own task requirements.
[0108] (1) Edge server: The edge server has corresponding computing resources and can provide offloading services for intensive tasks. It can also formulate payment strategies based on the multi-faceted requirements of the task vehicles to obtain benefits. The edge server set is defined as , define the number of edge servers as ;
[0109] (2) Electric vehicles: The task vehicle selects the computing resources of idle electric vehicles to complete the unloading and execution of the task, and charges the task vehicle for the unloading service provided through the established payment strategy to encourage more vehicles to provide unloading services. Define the electric vehicle set ;
[0110] (3) Mission vehicle: The mission vehicle will generate tasks locally and offload tasks that cannot be met by local computing resources to edge servers or electric vehicles. It will select the offloading target based on the various requirements of the tasks to be offloaded and pay the fees according to the payment strategy. It defines the time slot of each mission vehicle. Both The task needs to be unloaded, and the task set generated by the task vehicle is defined as , Submit the task for the mission vehicle, For the corresponding tasks, tasks By tuple composition, Indicates vehicle The size of the task data to be executed; Indicates the computing resources required for task processing per unit size of data; The unloading target represents the unloading indicator. The task vehicle unloads the task request to the edge server. , among which When the task is processed locally; when When it is unloaded to the corresponding server; when When , it means unloading to the corresponding electric vehicle. To indicate the time sensitivity of the task, the larger the value, the higher the completion time requirement of the task. Indicates the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task Reduced the quality of service required of mission vehicles;
[0111] The mission vehicle needs to complete the unloading within the communication coverage of the service. The execution time of the mission unloading should be within the expected completion time threshold of the mission, and a task can only be unloaded to one unloading target. If the mission vehicle chooses to execute on the local computing device, the unloading target ; If the task is offloaded to the edge server, the target ; If the task is offloaded to the electric vehicle, then the target .
[0112] In step 2, the communication model of the system model is designed as follows: Considering that the edge computing model with multiple electric vehicles and multiple edge servers randomly distributed is often used in large shopping malls and complex commercial districts, the free space path loss model is used to describe the process of unloading the task vehicle to the edge server, ignoring the time delay caused by the downlink transmission. It is assumed that for different mobile vehicles, orthogonal frequency division multiple access is used, and the channel bandwidth of each mobile vehicle is B. Ignoring the channel interference between mobile vehicles, it can be obtained from the Shannon formula that the mobile vehicle will transfer the task The uplink transmission rate offloaded to the edge server s is Similarly, the mobile vehicle can be used to complete the task Unloading to electric vehicles The uplink transmission rate is ,in represents the transmission power of the mission vehicle, and represent the mission vehicles and edge servers or electric vehicles, respectively Since the free space path loss model is used, the channel gain is inversely proportional to the distance. and They represent the variance of Gaussian white noise in the corresponding cases.
[0113] The time model for designing the system model is as follows:
[0114] After obtaining the uplink transmission rate between the mobile vehicle and the server or electric vehicle, without considering the speed loss during the transmission process, the task vehicle unloads the task Transmission time delay to edge server s and electric vehicle g and They are expressed as follows:
[0115]
[0116]
[0117] Indicates vehicle The size of the task data to be executed, represents the transmission power of the mission vehicle, B is the channel bandwidth owned by each mobile vehicle, To move the vehicle The uplink transmission rate offloaded to the edge server s, To move the vehicle Unloading to electric vehicles The uplink transmission rate is and They represent the variance of Gaussian white noise in the corresponding cases.
[0118] The total completion time of a task is usually composed of transmission time and execution time. According to the model, the task vehicle can offload part of the task to the edge server, electric vehicle or local execution. Due to different offloading targets, the execution time of the task will be calculated according to different offloading targets;
[0119] Define the computing resource set of each terminal computing device , where when the computing device subscript is 0, it represents the computing resource size of the local task vehicle, and the subscript between 1 and s represents the computing resource size of the edge server computing device. arrive The interval represents the computing resource size of the electric vehicle. The execution time under different offloading targets is given below:
[0120] (1) The mission vehicle performs the mission locally, and the local execution time is given by the following formula:
[0121]
[0122] in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the size of the local computing resources of the mission vehicle.
[0123] (2) When the task vehicle offloads the task to the edge server or electric vehicle, it needs to be transmitted through the communication link. Compared with local execution, there will be a transmission delay. The execution time of task offloading is obtained by the following formula:
[0124]
[0125] in Indicates the computing resource size of the edge server computing device, represents the computing resource size of the electric vehicle, and Represents the task vehicle unloading task The transmission time delay to the edge server s and the electric vehicle g.
[0126] The energy consumption model of the designed system model is as follows:
[0127] When a task vehicle is waiting to be executed, it will generate a lot of energy consumption, such as local execution energy consumption, transmission energy consumption, in addition, the edge server and electric vehicle in the entire unloading system also have energy consumption. When the vehicle is in a state of low energy consumption, the energy consumption should be less than the mission energy consumption threshold set by the vehicle, which defines the mission of the vehicle. The energy consumption threshold is .
[0128] (1) Transmission energy consumption: The transmission energy consumption of the mission vehicle mainly comes from the process of transmitting mission data to the edge server or electric vehicle through the wireless network. During this process, the mission vehicle needs to adjust the transmission power according to factors such as transmission distance, signal strength, and network bandwidth. Although higher transmission power can reduce transmission delay, it will also lead to more energy consumption. When the mission vehicle chooses to offload the task to the edge server or electric vehicle, it must take into account the stability and transmission rate of the network to avoid unnecessary energy waste or increased delay due to poor network conditions. According to the above formula, the delay of task transmission can be obtained, and the transmission power of the mission vehicle is defined as , the energy consumed when offloading the task to the server is: ; Offload tasks to electric vehicles The energy required is Among them and They are the transmission delay of offloading tasks to edge servers and electric vehicles, respectively.
[0129] (2) Local energy consumption: When the resources required for task processing are small, the network conditions are not ideal, or the task processing delay requirements are high, the task vehicle chooses to perform the task locally. Although local processing may reduce the energy consumption of task transmission, it will increase the computing burden of the vehicle itself, resulting in more energy consumption. Therefore, the task vehicle needs to make a trade-off between computing resource consumption and network transmission consumption and choose the most appropriate processing method. The following defines the power consumption of the computing device of the task vehicle as , the energy consumption of the task that needs to be processed locally can be obtained by the following formula:
[0130] (3) Electric vehicle energy consumption: When a task vehicle chooses to offload a task to an electric vehicle, the electric vehicle is responsible for processing the offloaded task. The computing energy consumption of the electric vehicle is related to its processing power and the complexity of the task. Electric vehicles are usually in a low-power state and have relatively limited computing resources. The computing device power consumption of an electric vehicle is defined as ,in The energy consumed by the electric vehicle computing device when processing tasks is: .in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the computing resource size of the electric vehicle.
[0131] (4) Edge server energy consumption: Edge computing servers are usually deployed at the edge of the network, close to user devices, and can provide relatively powerful computing resources. The main advantage of offloading tasks to edge servers is that they can utilize the high-performance computing capabilities of the server, thereby significantly reducing the processing time of the task. However, this also brings about the problem of computing energy consumption. The power consumption of the computing equipment of the edge server depends on the performance of the server and the complexity of the task. Usually, the computing requirements of the task and the load capacity of the server need to be considered when making task offloading decisions. If the edge server load is too high, it may lead to increased energy consumption and affect the response time of the task. The energy consumed by the edge server when processing a task can be calculated by the following formula: ,in represents the computing device power consumption of the edge server, Indicates the computing resource size of the edge server computing device.
[0132] The payment cost model for the designed system model is as follows: Similar to carrier-based fees for data traffic, after a vehicle offloads tasks to an edge server or electric vehicle, the edge server or electric vehicle will charge the tasking vehicle a corresponding fee based on its own payment strategy. This serves as revenue for the edge server or electric vehicle and also as a cost for the tasking vehicle. Because computing power varies across heterogeneous edge environments, the fees charged by edge devices are assumed to be proportional to their various computing resources.
[0133] Edge servers and electric vehicles charge fees for the memory usage, computing resource requirements, and security costs of tasks generated by task vehicles to ensure their own operating costs. To facilitate the calculation of charging standards, it is assumed that edge servers and electric vehicles are priced according to their own computing resources and charged according to the amount of data offloaded by the task. The specific calculation method for the fee that task vehicle m needs to pay for offloading tasks is as follows:
[0134]
[0135] in It represents the price charged by the edge server for processing each unit of data. represents the price charged by the electric vehicle for processing each unit of data. Assuming that the fee paid by the mission vehicle is related to the data size of the mission and the mission's sensitivity to time delays, a higher fee can be paid for more efficient services. The total fee paid by mission vehicle m can be defined as:
[0136]
[0137] In step 3, based on the problem model, the system objective function is designed. In summary, it can be concluded that the task offloading cost of the task vehicle is composed of the fees paid by the task vehicle, the time delay of the task vehicle offloading the task, and the energy consumption. Since the data offloaded remotely by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, after formalizing the task offloading problem, the following optimization problem can be obtained:
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145] in represents the payment cost of the mission vehicle, represents the mission time delay of the mission vehicle, represents the energy consumption of the entire system, Indicates the maximum time delay threshold of the task, Indicates the uninstall indicator of the task, Indicates a task Constraint 1 indicates that the local execution time of the task should be less than the delay threshold of the task vehicle. Constraints 2 and 3 indicate that the time delay of the task offloaded to the edge server or electric vehicle should be less than the expected time delay threshold of the task. Constraint 4 indicates that the offloading target limit of the task is ; Constraints 5 and 6 indicate that the energy consumption of local execution should be less than the mission expected energy consumption threshold set by the mission vehicle.
[0146] In step 4, the proximal strategy is used to optimize the reinforcement learning method such as Figure 2 As shown in Figure 1, the magnitude of the policy update is constrained to maximize the agent's cumulative reward while ensuring training stability. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by limiting the magnitude of the policy update.
[0147] MOPPO combines the Actor-Critic architecture, uses the actor network to generate action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, thereby optimizing the direction of policy update.
[0148] The task vehicle can be considered an intelligent agent, and its unloading policy is trained using the MOPPO algorithm. The unloading problem can typically be described using a Markov decision process represented by a four-tuple (S, A, R, P), where S is the state space, A is the action space, R is the reward function, and P is the state transition probability.
[0149] State space: Reinforcement learning continuously learns strategies from historical information, so a comprehensive state definition is crucial for decision-making efficiency. It is necessary to consider task requirements, network conditions, and available computing resources. The state of time slot k is defined as , by the task tuple and the vehicle edge computing environment, where Indicates vehicle The size of the task data that needs to be processed, Represents the various computing resources required for the task, To express the time delay sensitivity of the task, represents the expected time delay threshold of the task, represents the expected energy consumption threshold of the task, and F represents the computing resource set of the task vehicle's surrounding environment.
[0150] Action space: In the unloading scenario, the task vehicle will make a decision on the unloading target after receiving feedback from the state space. The unloading decision is designed as a discrete vector , in the alternative model Subscript the task's offload target.
[0151] Reward function: In each time slot k, the agent receives a reward when it reaches the next state after performing an action In general, the reward function should be positively correlated with the objective function. Since the objective function of the problem is to minimize the time delay of the task vehicle, energy consumption, and payment cost, the reward function is negatively correlated with the cost of the task vehicle. According to the problem model described in the previous section, the reward function can be designed as follows, where Represents the weight of each target parameter:
[0152]
[0153] In step 5, a multi-objective proximal policy optimization algorithm is designed. The Actor-Critic architecture is shown in Figure 3: MOPPO adopts the Actor-Critic architecture. The Actor network is responsible for generating policies, and the Critic network is responsible for evaluating the state value function. The Actor network outputs the probability distribution of actions, while the Critic network estimates the value of the current state. , used to calculate the advantage function , that is, the probability distribution of the output action given a specific state. The Actor network decides the action based on the current state and gradually improves the strategy by optimizing the objective function, thereby guiding the strategy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean square error of the value function, and the advantage function The calculation formula is as follows:
[0154] ;
[0155] in represents the discount factor, To receive the reward, Represents state k+1;
[0156] Importance sampling mechanism: Importance sampling is a method for adjusting data weights. The MOPPO algorithm uses importance sampling, but it does not use the random sampling and elimination mechanism used in experience replay. Instead, it uses data in chronological order of the results of the agent's interaction with the environment to measure the difference between the data distribution and the distribution of the current policy.
[0157] Strategy optimization: The core of the MOPPO algorithm is to optimize the strategy parameters To maximize the expected cumulative reward, the MOPPO algorithm introduces proximal constraints to limit the magnitude of policy updates and uses a clipping objective function to ensure that the difference between the new policy and the old policy is not too large. The clipping objective function is defined as follows:
[0158]
[0159] in, is the probability ratio of the new strategy to the old strategy, defined as the current strategy in state Select Action Probability The probability of choosing the same state as the old strategy If the ratio is less than or greater than , then crop it to and between, among is the advantage function, is the clipping range parameter.
[0160] Train the agent to interact with the environment. The training process is as follows:
[0161] First, the input data required by the algorithm is initialized, including the tasks that the vehicle needs to unload, and the computing requirements, delay requirements and energy limits of the tasks. The edge nodes involved in the task vehicle unloading calculation output the unloading results of the tasks, as well as the cost of task unloading, including energy consumption and time delay. The scale of the task set and its multi-dimensional performance requirements are clarified. The multi-dimensional performance requirements include computing size, energy consumption and delay sensitivity. The initialization strategy network is used to generate the task unloading action strategy. The value function network evaluates the state value to guide the strategy optimization. The experience replay buffer stores historical interaction data. The interaction data includes old state, action, reward, and new state. The initial state is obtained at each time sequence, including various attributes of the task, the resource status of the edge server and the collaborative ability of the vehicle cluster. Based on the current state , policy network Output Action , select the target server to offload or execute locally, optimize the action probability distribution through policy gradient, and calculate the cost, execution time and energy consumption of the task in this time sequence. After executing the action, observe the new state and calculate the immediate reward , calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , the algorithm line 10 converts the quad Store in buffer D for subsequent policy updates and calculation of generalized advantage function , quantify the pros and cons of actions relative to the average policy, estimate the state value difference through the critic network, solve the credit allocation problem under sparse rewards, enhance the accuracy of policy gradient estimation, and calculate the objective function according to the definition of the clipping objective function , used to update the Actor network and the Critic network, control the changes in the strategy to ensure the stability of training, balance the exploration and utilization of the intelligent agent, calculate the loss value of the Actor network and the Critic network, and then update the network parameters of the Actor network according to the importance sampling results , the critic network updates parameters through the generalized advantage function GAE .
[0162] The following is an introduction to the settings of various parameters in the experiment. All neural networks in the experiment are implemented using PyTorch 2.0.0 and Python 3.8.1 platforms, and Adam is used to optimize the network. A large number of simulation experiments are conducted to evaluate the performance of the algorithm proposed in this application. Suppose the amount of data generated by the mission vehicle follows a normal distribution, The generated data should be processed in real time and deleted after the processing results are returned to avoid storage overflow of local or edge vehicles. Assume that the transmission power of the task vehicle is 10Kwh, and the computing resources of the local computing device of the task vehicle in time slot t are , the computing resources of the edge server default to , the average computing resources of the vehicle cooperation cluster default to .
[0163] In the experimental simulation of the application, the Actor network has two hidden fully connected layers, each containing 128 nodes. The Critic network also has two hidden fully connected layers, each containing 128 nodes. Each task vehicle simulates a real-world scenario and calculates the unloading results of the task vehicle based on its own task requirements using an unloading algorithm based on multi-objective proximal policy gradient optimization. The experiment selected three baseline methods for comparison with the algorithm proposed in this application: the Q-learning algorithm, the random unloading algorithm, and the fully local algorithm. The main experimental parameter settings of the algorithm in this application are shown in Table 1:
[0164] Table 1 Main experimental parameter settings
[0165]
[0166] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.
Claims
1. A task offloading method for electric vehicles, characterized by: The specific steps include: Step 1: Considering the vehicle mobility scenario, design an edge-to-end system offloading framework; In step 1, the system model of the edge-end system offloading framework is divided into three parts: electric vehicles, mission vehicles, and edge servers; The edge-end system offloading framework has a three-layer architecture; The first layer is the edge server, which quickly processes offloaded tasks and provides services for intensive tasks; The second layer is electric vehicles, which are composed of vehicles parked around the road and have idle computing resources for mission vehicles to choose to unload; The third layer is the mission vehicle. Mission vehicles generate data to be processed and are equipped with small computing devices. Intensive tasks need to be offloaded to edge servers or electric vehicles. Tasks with different service quality requirements need to be offloaded based on the surrounding computing resources. In step 1, in the vehicle mobility scenario of the Internet of Vehicles environment, considering the mobility of the vehicle, the vehicle needs to complete the task within the coverage service range of the edge server. When the vehicle user cannot process all the tasks, he needs to offload the data to the edge server or electric vehicle to meet his own task requirements. The edge server, electric vehicle and task vehicle are defined as follows: (1) Edge Server: The edge server has corresponding computing resources and can provide offloading services for intensive tasks. It can also formulate payment strategies to obtain benefits based on the multi-faceted needs of the task vehicles. The edge server set is defined as , define the number of edge servers as ; (2) Electric vehicles: The task vehicle selects the computing resources of idle electric vehicles to complete the unloading and execution of the task, and charges the task vehicle for the unloading service provided through the formulated payment strategy to encourage more vehicles to provide unloading services. The electric vehicle set is defined ; (3) Mission vehicles: The mission vehicle will generate tasks locally and offload tasks that cannot be met by local computing resources to edge servers or electric vehicles. It will select the offloading target based on the various requirements of the tasks to be offloaded and pay the fees according to the payment strategy, defining the time slot of each mission vehicle. Both The task needs to be unloaded, and the task set generated by the task vehicle is defined as , Submit bid for mission vehicle, For the corresponding task, vehicle unloading task By tuple composition, Indicates vehicle The size of the task data to be executed; Indicates the computing resources required for task processing per unit size of data; The unloading target represents the unloading indicator. The task vehicle unloads the task request to the edge server. , among which When the task is processed locally; when When it is unloaded to the corresponding server; when When , it means unloading to the corresponding electric vehicle. To indicate the time sensitivity of the task, the larger the value, the higher the completion time requirement of the task. Indicates the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task Reduced the quality of service required of mission vehicles; The mission vehicle needs to complete the unloading within the communication coverage of the service. The execution time of the mission unloading should be within the expected completion time threshold of the mission, and a task can only be unloaded to one unloading target. If the mission vehicle chooses to execute on the local computing device, the unloading target ; If the task is offloaded to the edge server, the target ; If the task is offloaded to the electric vehicle, then the target , where s is the edge server and For electric vehicles; Step 2: The mission vehicle generates intensive tasks and designs the communication model, energy model, time model, and payment cost model; The communication model designed in step 2 is as follows: The free space path loss model is used to describe the process of unloading the task from the mission vehicle to the edge server, ignoring the delay caused by the downlink transmission. It is assumed that for different mobile vehicles, orthogonal frequency division multiple access is used, and the channel bandwidth of each mobile vehicle is B. The channel interference between mobile vehicles is ignored. According to the Shannon formula, the mobile vehicle unloads the task from the vehicle. The uplink transmission rate offloaded to the edge server s is Similarly, the mobile vehicle can be used to complete the task Unloading to electric vehicles The uplink transmission rate is ,in represents the transmission power of the mission vehicle, and represent the mission vehicles and edge servers or electric vehicles, respectively The channel gain between the two channels is inversely proportional to the distance due to the free space path loss model. and Respectively represent the variance of Gaussian white noise in the corresponding cases; The time model in step 2 is as follows: After obtaining the uplink transmission rate between the mobile vehicle and the server or electric vehicle, without considering the speed loss during the transmission process, the task vehicle unloads the task Transmission time delay to edge server s and electric vehicle g and They are expressed as follows: ; ; Indicates vehicle The size of the task data to be executed, represents the transmission power of the mission vehicle, B is the channel bandwidth owned by each mobile vehicle, Unloading the vehicle for the mobile vehicle The uplink transmission rate offloaded to the edge server s, Unloading the vehicle for the mobile vehicle Unloading to electric vehicles The uplink transmission rate is and Respectively represent the variance of Gaussian white noise in the corresponding cases; The total completion time of a task consists of transmission time and execution time. According to the model, the task vehicle can offload part of the task to the edge server, electric vehicle, or local execution. Due to different offloading targets, the execution time of the task will be calculated according to different offloading targets; Define the computing resource set of each terminal computing device , where the computing device subscript 0 represents the local computing resource size of the task vehicle, the subscript between 1 and s represents the computing resource size of the edge server computing device, and the subscript between s+1 and s+g represents the computing resource size of the electric vehicle. The execution time under different offloading targets is given below: (1) The mission vehicle performs the mission locally, and the local execution time is given by the following formula: ; in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the size of the local computing resources of the mission vehicle, To uninstall the target; (2) When the task vehicle offloads the task to the edge server or electric vehicle, it needs to be transmitted through the communication link. Compared with local execution, there will be a transmission delay. The execution time of task offloading is obtained by the following formula: ; in Indicates the computing resource size of the edge server computing device, represents the computing resource size of the electric vehicle, and Represents the task vehicle unloading task The transmission time delay to the edge server s and the electric vehicle g; The energy consumption model in step 2 is as follows: When the vehicle performs the vehicle unloading task locally When the energy consumption is less than the task energy consumption threshold set by the vehicle, the vehicle unloading task performed by the vehicle is defined. The energy consumption threshold is ; (1) Transmission energy consumption: The transmission energy consumption of the mission vehicle comes from the process of transmitting mission data to the edge server or electric vehicle through the wireless network. During this process, the mission vehicle needs to adjust the transmission power according to the transmission distance, signal strength and network bandwidth factors. When the mission vehicle chooses to offload the task to the edge server or electric vehicle, the task transmission delay is obtained. The transmission power of the mission vehicle is defined as , the energy consumed by offloading the task to the edge server is: ; The energy consumed when offloading the task to the electric vehicle is: ,in and are the transmission delays of offloading tasks to edge servers and electric vehicles, represents the transmission power of the mission vehicle; (2) Local energy consumption: When the resources required for task processing are small, the network conditions are not ideal, or the task processing latency requirement is high, the task vehicle chooses to perform the task locally. The task vehicle needs to make a trade-off between computing resource consumption and network transmission consumption and choose the most appropriate processing method. The power consumption of the computing device of the task vehicle is defined as follows: , the local execution time of the task is ,Therefore, the energy consumption of the task that needs to be processed locally can be obtained by the following formula: ; (3) Energy consumption of electric vehicles: When the task vehicle chooses to offload the task to the electric vehicle, the electric vehicle is responsible for processing the offloaded task. The computing energy consumption of the electric vehicle is related to its processing power and the complexity of the task. Electric vehicles are usually in a low-power state and have relatively limited computing resources. The power consumption of the computing device of the electric vehicle is defined as follows: ,in , the energy consumed by the electric vehicle computing device when processing tasks is: ,in Indicates vehicle The size of the task data to be executed, Indicates the computing resources required for task processing per unit size of data, Indicates the computing resource size of the electric vehicle; (4) Edge server energy consumption: Edge computing servers are deployed at the edge of the network, close to user devices. The power consumption of the computing equipment of the edge server depends on the performance of the server and the complexity of the task. When making task offloading decisions, it is necessary to consider the computing requirements of the task and the load capacity of the server. If the edge server load is too high, it will lead to increased energy consumption and affect the response time of the task. The energy required by the edge server to process the task can be calculated by the following formula: ,in represents the power consumption of computing equipment of edge server, Indicates the computing resource size of the edge server computing device; In step 2, the payment cost model is set as follows: After the vehicle offloads the task to the edge server or electric vehicle, the edge server or electric vehicle will charge the task vehicle a corresponding fee based on its own payment strategy. This fee serves as the edge server's or electric vehicle's revenue and also as one of the task vehicle's costs. Edge servers and electric vehicles charge fees for the memory usage, computing resource requirements, and security costs of tasks generated by task vehicles to ensure their own operating costs. Assuming that edge servers and electric vehicles are priced according to their own computing resources and charged according to the amount of data offloaded by the task, the specific calculation method for the fee that task vehicle m needs to pay for offloading tasks is as follows: ; in It represents the price charged by the edge server for processing each unit of data. represents the price charged by the electric vehicle for processing each unit of data. Assuming that the fee that the task vehicle needs to pay is related to the data size of the task and the sensitivity of the task time delay, the total fee that the task vehicle m needs to pay for obtaining the service is calculated as follows: ; Indicates vehicle The size of the task data to be executed, To indicate the time sensitivity of the task; Step 3: Based on the problem model, design the system objective function and optimize the system time delay, energy consumption and payment cost with multiple objectives; Step 4: Design the vehicle-agent interaction environment based on the system model and reinforcement learning method; Step 5: Design a multi-objective proximal strategy optimization algorithm and determine the unloading target after training the intelligent agent.
2. The method for offloading tasks for electric vehicles according to claim 1, characterized in that: In step 3, the system objective function is designed based on the problem model to obtain the task offloading cost of the task vehicle. The cost is composed of the fees paid by the task vehicle, the time delay of the task vehicle offloading the task, and the energy consumption. Since the data offloaded remotely by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, the task offloading problem is formalized to obtain the following optimization problem: ; ; ; ; ; ; ; in represents the payment cost of the mission vehicle, represents the mission time delay of the mission vehicle, e represents the energy consumption of the entire system, Indicates the maximum time delay threshold of the task, Indicates the uninstall indicator of the task, is the local execution time of the task, is the execution time of the task offload, is the transmission delay of offloading the task to the electric vehicle, The energy consumption required to process the task locally, The energy consumed by offloading tasks to edge servers, The energy required to offload the task to the electric vehicle; Constraint 1 indicates that the local execution time of the task should be less than the delay threshold of the task vehicle. Constraints 2 and 3 indicate that the time delay of offloading the task to the edge server or electric vehicle should be less than the expected time delay threshold of the task. Constraint 4 indicates that the offloading target limit of the task is ; Constraints 5 and 6 indicate that the energy consumption of local execution should be less than the mission expected energy consumption threshold set by the mission vehicle.
3. The method for offloading tasks for electric vehicles according to claim 1, characterized in that: In step 4, the proximal strategy is used to optimize the reinforcement learning method and constrain the amplitude of the policy update to maximize the agent's cumulative reward while ensuring training stability. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by limiting the amplitude of the policy update. MOPPO combines the Actor-Critic architecture, uses the actor network to generate action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, thereby optimizing the direction of policy update; The task vehicle can be regarded as an intelligent agent. The unloading strategy of the task vehicle is trained by the MOPPO algorithm and described by a Markov decision process represented by the four-tuple S, A, R, P, where S is the state space, A is the action space, R is the reward function, and P is the state transition probability. State Space: It is necessary to consider task requirements, network conditions and available computing resources, and define the state of time slot k as , unloading tasks by vehicles of task tuples and the vehicle edge computing environment, where represents the size of the task data that vehicle m needs to process, Represents the various computing resources required for the task, To express the time delay sensitivity of the task, represents the expected time delay threshold of the task, represents the expected energy consumption threshold of the task, and F represents the computing resource set of the task vehicle's surrounding environment; Action Space: In the unloading scenario, after receiving feedback from the state space, the task vehicle will make a decision on the unloading target. The unloading decision is designed as a discrete vector , in the alternative model Subscript the task's offloading target; Reward function: In each time slot k, the agent receives a reward when it reaches the next state after executing the action. In general, the reward function should be positively correlated with the objective function. Since the objective function of the problem is to minimize the time delay of the task vehicle, energy consumption, and payment cost, the reward function is negatively correlated with the task vehicle cost. According to the problem model described in the previous section, the reward function can be designed as follows, where Represents the weight of each target parameter: ; in represents the payment cost of the mission vehicle, represents the mission time delay of the mission vehicle, and e represents the energy consumption of the entire system.
4. The method for offloading tasks for electric vehicles according to claim 3, characterized in that: In step 5, a multi-objective proximal policy optimization algorithm is designed, Actor-Critic architecture: MOPPO adopts the Actor-Critic architecture, where the Actor network is responsible for generating strategies, the Critic network is responsible for evaluating the state value function, the Actor network outputs the probability distribution of actions, and the Critic network estimates the value of the current state. , used to calculate the advantage function , that is, the probability distribution of the output action given a specific state. The Actor network decides the action based on the current state and gradually improves the strategy by optimizing the objective function, thereby guiding the strategy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean square error of the value function, and the advantage function The calculation formula is as follows: ; in represents the discount factor, To receive the reward, Indicates state k+1; Importance sampling mechanism: Importance sampling is a method of adjusting data weights. The MOPPO algorithm uses importance sampling to use data in the chronological order of the results of the interaction between the agent and the environment. Strategy Optimization: The core of the MOPPO algorithm is to optimize the strategy parameters To maximize the expected cumulative reward, the MOPPO algorithm introduces proximal constraints to limit the magnitude of policy updates and uses a clipping objective function to ensure that the difference between the new policy and the old policy is not too large. The clipping objective function is defined as follows: ; in, The probability of the new strategy to the old strategy is the strategy update ratio, which is defined as the current strategy in the state Next select action Probability The probability of choosing the same state as the old strategy If the ratio is less than or greater than , then crop it to and between, among is the advantage function, is the clipping range parameter.
5. A task offloading method for electric vehicles according to any one of claims 1 to 4, characterized in that: In step 5, the agent interaction environment is trained. The training process is as follows: First, the input data required by the algorithm is initialized, including the tasks that need to be unloaded by the vehicle, the computing requirements, delay requirements and energy limits of the tasks, the edge nodes involved in the task vehicle unloading calculation, and the output of the task unloading results; as well as the cost of task unloading, including energy consumption and time delay; clarify the scale of the task set and its multi-dimensional performance requirements, which include computing size, energy consumption and delay sensitivity. The initialization strategy network is used to generate the task unloading action strategy, the value function network evaluates the state value to guide the strategy optimization, and the experience replay buffer stores historical interaction data. The interaction data includes old state, action, reward, and new state. The initial state is obtained at each time sequence, including various attributes of the task, the resource status of the edge server and the collaborative ability of the vehicle cluster. Based on the current state , policy network Output Action , select the target server for offloading or local execution, optimize the action probability distribution through policy gradient, calculate the cost, execution time and energy consumption of the task in this time sequence, observe the new state after executing the action, and calculate the immediate reward , calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , the algorithm line 10 converts the quad Store in buffer D for subsequent policy updates and calculation of generalized advantage function , quantify the pros and cons of actions relative to the average policy, estimate the state value difference through the critic network, solve the credit allocation problem under sparse rewards, enhance the accuracy of policy gradient estimation, and calculate the objective function according to the definition of the clipping objective function , used to update the Actor network and the Critic network, control the changes in the strategy to ensure the stability of training, balance the exploration and utilization of the intelligent agent, calculate the loss value of the Actor network and the Critic network, and then update the network parameters of the Actor network according to the importance sampling results , the critic network updates parameters through the generalized advantage function GAE .
Citation Information
Patent Citations
Markov chain-based differential privacy task unloading method in Internet of Vehicles environment
CN117749797A
Macro base station placement method based on Internet of Vehicles task unloading decision
CN113810878A
Calculation unloading and privacy protection joint optimization method for mobile block chain network
CN116437341A