Task unloading method for electric vehicle

By using the MOPPO algorithm in the Internet of Vehicles environment, the task vehicle is modeled as an intelligent body, combined with edge servers and the computing resources of electric vehicles, the task offload decision is optimized, and the time delay, energy consumption and payment costs of task offloading electric vehicles in the Internet of Vehicles is solved, and the system efficiency and cost reduction are improved.

CN120343028AActive Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510815902.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the Internet of Vehicles environment, the task offloading algorithm of electric vehicles is difficult to effectively optimize time delay, energy consumption and payment costs, and the existing methods have problems of training instability and decision-making conflicts.

Method used

Multi-objective near-end strategy optimization (MOPPO) algorithm is used to model the task vehicle as an agent, and task offload decisions are made through a reinforcement learning framework. Combining the computing resources of edge servers and electric vehicles, a multi-objective optimization system is designed to consider time delay, energy consumption and payment costs.

Benefits of technology

The total cost of the system is reduced, the system efficiency is improved, the terminal load pressure is reduced, the task offload strategy is optimized, the time delay and energy consumption is reduced, and the payment cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343028A_ABST
    Figure CN120343028A_ABST
Patent Text Reader

Abstract

The invention provides a task unloading method for an electric vehicle. The method comprises the following steps: firstly, designing an unloading system model of the Internet of Vehicles, considering time delay, energy consumption and payment cost of a task vehicle, and modeling an edge unloading scene of the task vehicle into an intelligent agent interaction scene after formalizing an unloading model problem of an Internet of Vehicles system; a multi-objective optimization method is designed in combination with a deep reinforcement learning algorithm to solve the problem. Finally, experimental results of the method in multiple aspects of time delay, energy consumption, unloading cost and the like are analyzed through simulation experiments, and the results show that compared with a traditional unloading algorithm, the time delay is reduced by about 43%, and the energy consumption is reduced by about 38%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless communication networks, and specifically to a task offloading method for electric vehicles. Background Art

[0002] With the development of the automotive industry, many vehicles are equipped with functions such as autonomous driving and entertainment audio and video, which will generate tasks with low-latency requirements and low-cost requirements. Since the various resources of the computing devices equipped on the vehicles are limited, it is necessary to offload cumbersome tasks to edge servers or the computing resources of surrounding idle vehicles, thereby alleviating the execution load on the vehicle local.

[0003] By simulating the task offloading process in the vehicle-to-everything (V2X) network, the task vehicle is regarded as an agent, and the agent makes decisions based on its local observation data (such as network status, task type, current load, etc.). Through the deep reinforcement learning model, the agent can obtain rewards through interaction with the environment, learn the optimal offloading strategy, and achieve dynamic adjustment of task allocation, so that the algorithm can not only optimize traditional performance metrics (such as latency, energy consumption, etc.), but also take into account the requirements and preferences of vehicles and systems to provide a more personalized and intelligent offloading strategy.

[0004] The comparison between the present application and the prior art is as follows:

[0005] Technical comparison with the application text with Chinese publication number CN117749797A, "Differential Privacy Task Offloading Method Based on Markov Chain in Vehicle-to-Everything Environment";

[0006] 1. The application text with Chinese publication number CN117749797A discloses a differential privacy task offloading method based on Markov chain in vehicle-to-everything environment. The method includes constructing and initializing the discrete Markov state space of each vehicle; updating the state transition probability matrix according to the moving speed of the current vehicle, the number of surrounding vehicles and the number of edge servers to obtain the privacy parameter of the current vehicle; performing local differential privacy protection on the position of the current vehicle and generating the confused position of the current vehicle; calculating the delay generated by transmitting tasks between the current vehicle and the edge server, and constructing a system task offloading delay objective function to be minimized; using the whale algorithm to process the system task offloading delay objective function, searching for the optimal offloading scheme, and performing differential privacy task offloading. The present invention reduces the risk of data leakage and provides an effective and flexible solution for privacy protection in vehicle-to-everything while ensuring the task offloading delay in vehicle-to-everything. The focus is on vehicle-to-everything and privacy protection.

[0007] 2. The application text proposes a task offloading algorithm for electric vehicles in the vehicle-to-everything (V2X) network. The algorithm first designs an offloading model for the V2X network, considering the time delay, energy consumption, and payment cost of task vehicles. Then, it formalizes the problem of the V2X system in the model form, models the edge offloading scenario of task vehicles as an agent interaction scenario, and uses the proximal policy optimization reinforcement learning algorithm to ensure the stability of training and reduce conflicts. Regarding task vehicles as agents, it trains the offloading decisions of task vehicles through the MOPPO algorithm. The algorithm enables agents to learn in a simulation environment, outputs an optimal offloading strategy through training data, and simultaneously minimizes the time delay, energy consumption, and payment cost generated by task execution in the system. The focus of this application is to obtain an offloading strategy through multi-objective optimization in the V2X environment.

[0008] There are essential differences between the two in terms of scheduling objectives, system structure, and technical routes.

[0009] Technical comparison with the application text, Chinese Patent Publication No. CN113810878A, "A Macro Base Station Placement Method Based on Task Offloading Decision in Vehicle-to-Everything Network";

[0010] 1. Chinese Patent Publication No. CN113810878A of the application text discloses a macro base station placement method based on task offloading decision in the vehicle-to-everything network, specifically as follows: Step 1: Establish Y*Y coding matrices and combine rows and columns; Step 2: Establish a digital twin network and simulate each combination; Step 3: Calculate the optimal task offloading decision for each macro base station in each combination; Step 4: Calculate the total energy consumption of each macro base station under each combination; Step 5: Establish a minimum objective function for the total energy consumption and use the particle swarm algorithm to solve it; thus obtaining the optimal combination. This application considers the placement problem of macro base stations in edge computing, reduces the idle time of edge servers in macro base stations on the premise of meeting task computing requirements, and further reduces energy consumption. The focus is on the deployment location of base stations and task offloading efficiency in the vehicle-to-everything network.

[0011] 2. The application text proposes a multi-objective task offloading optimization algorithm for the vehicle-to-everything scenario of electric vehicles. This solution first constructs a vehicle-to-everything task offloading model that includes multiple dimensions such as time delay, energy consumption, and payment cost, and transforms the complex edge computing offloading problem into a quantifiable optimization problem. Based on the reinforcement learning framework, it innovatively adopts the multi-objective proximal policy optimization (MOPPO) algorithm, models each task vehicle as an independent agent, and dynamically optimizes offloading decisions through continuous interaction learning between the agent and the simulation environment. The algorithm mainly solves the problem of unstable training in traditional methods, effectively reduces decision conflicts, and finally outputs a task offloading strategy that is balanced and optimized in multiple objective dimensions such as time delay, energy consumption, and cost. The focus of this application is to obtain an offloading strategy through multi-objective optimization in the vehicle-to-everything environment.

[0012] There are essential differences between the two in terms of system architecture, scheduling objectives, and technical paths. Summary of the Invention

[0013] The present invention proposes a task offloading method for electric vehicles. In the scenario of electric vehicle movement, various complex factors such as time delay, energy consumption, and payment cost are considered, reducing the total cost of the system, improving system efficiency, and reducing the load pressure on the terminal.

[0014] To achieve the above object, the technical solution adopted by the present invention is:

[0015] A task offloading method for electric vehicles includes the following specific steps: Step 1: Design an edge-end system offloading framework considering the vehicle movement scenario; Step 2: Task vehicles generate intensive tasks, and design a communication model, an energy model, a time model, and a payment cost model; Step 3: According to the problem model, design a system objective function to multi-objectively optimize system time delay, energy consumption, and payment cost; Step 4: Design a vehicle agent interaction environment according to the system model and the reinforcement learning method; Step 5: Design a multi-objective proximal policy optimization algorithm, and determine the offloading target after training the agent. Offloading target.

[0016] As a further improvement of the present invention, in Step 1, the system model of the edge-end system offloading framework is divided into three parts: electric vehicles, task vehicles, and edge servers; The edge-end system offloading framework has a three-layer architecture; The first layer is the edge server, which quickly processes the offloaded tasks and provides services for intensive tasks; The second layer is the electric vehicle, which is composed of vehicles parked around the road and has idle computing resources for task vehicles to choose for offloading; The third layer is the task vehicle, which will generate data to be processed, and the task vehicle itself is equipped with a computing device with small computing power. Intensive tasks need to be offloaded to the edge server or electric vehicle, and different tasks with different service quality requirements need to be offloaded according to the surrounding computing resources; In Step 1, in the vehicle movement scenario of the vehicle networking environment, considering the mobility of the vehicle, the vehicle needs to complete tasks within the coverage service range of the edge server. When the vehicle user cannot process all tasks, data needs to be offloaded to the edge server or electric vehicle to meet its own task requirements. The edge server, electric vehicle, and task vehicle are defined as follows: (1) Edge server: The edge server has corresponding computing resources, can provide offloading services for intensive tasks, and formulate payment strategies according to the multi-faceted requirements of the task vehicle to obtain benefits. The edge server set is defined as , define the number of edge servers as ; (2) Electric vehicles: The task vehicle selects the computing resources of idle electric vehicles to complete the unloading and execution of the task. Through the payment strategy formulated, the task vehicle is charged for the unloading service provided to encourage more vehicles to provide unloading services. The electric vehicle set is defined ; (3) Mission vehicles: The mission vehicle will generate tasks locally and offload tasks that cannot be met by local computing resources to edge servers or electric vehicles. It will select the offloading target according to the various requirements of the tasks to be offloaded, pay the fees according to the payment strategy, and define the time slot of each mission vehicle. Both The task needs to be unloaded, and the task set generated by the task vehicle is defined as , Submit bid for the mission vehicle, For the corresponding task, vehicle unloading task By tuple composition, Indicates vehicle The size of the task data to be executed; Indicates the computing resources required for task processing per unit size of data; The unloading target represents the unloading indicator. The task vehicle unloads the task request to the edge server. , among which When , it indicates that the task is processed locally; when When When , it means unloading to the corresponding electric vehicle. To indicate the time sensitivity of a task, the larger the value, the higher the completion time requirement of the task. Indicates the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task Reduced the quality of service required of mission vehicles; The task vehicle needs to complete the unloading within the communication coverage of the service. The execution time of the task unloading should be within the expected completion time threshold of the task, and a task can only be unloaded to one unloading target. If the task vehicle chooses to execute on the local computing device, the unloading target ; If the task is offloaded to the edge server, then the target ; If the task is offloaded to an electric vehicle for execution, the offloading target , where s is the edge server and is the electric vehicle.

[0017] As a further improvement of the present invention, the communication model designed in step 2 is as follows: The free space path loss model is adopted to describe the process of task offloading from a task vehicle to an edge server, ignoring the delay generated by downlink transmission. Assuming orthogonal frequency division multiple access is used for different mobile vehicle transmissions, and each mobile vehicle has a channel bandwidth of B, ignoring the channel interference between mobile vehicles, according to the Shannon formula, the uplink transmission rate of a mobile vehicle for offloading its vehicle offloading task to the edge server s is . Similarly, the uplink transmission rate of a mobile vehicle for offloading its task to the electric vehicle is , where represents the transmission power of the task vehicle, and respectively represent the channel gains between the task vehicle and the edge server s or the electric vehicle . Since the free space path loss model is adopted, the channel gain is inversely proportional to the distance. and respectively represent the variances of Gaussian white noise in the corresponding cases.

[0018] As a further improvement of the present invention, the time model in step 2 is as follows: After obtaining the uplink transmission rate between the mobile vehicle and the server or the electric vehicle, without considering the speed loss during the transmission process, the transmission time delays of the task vehicle for offloading its task to the edge server s and the electric vehicle g are respectively expressed as the following formulas: ; ; represents the size of the task data that the vehicle needs to execute, represents the transmission power of the task vehicle, B is the channel bandwidth owned by each mobile vehicle, is the uplink transmission rate of the mobile vehicle for offloading its vehicle offloading task to the edge server s, is the uplink transmission rate of the mobile vehicle for offloading its vehicle offloading task to the electric vehicle . and respectively represent the variances of Gaussian white noise in the corresponding cases; The total completion time of the task consists of the transmission time and the execution time. According to the model, the task vehicle can offload part of the task to the edge server, electric vehicle or execute it locally. Since the offloading targets are different, the execution time of the task will be calculated according to different offloading targets; Define the set of computing resources of each terminal computing device , where when the subscript of the computing device is 0, it represents the size of the computing resources of the task vehicle locally, the subscript between 1 and s represents the size of the computing resources of the edge server computing device, and the subscript between s + 1 and s + g represents the size of the computing resources of the electric vehicle. The following gives the execution time under different offloading targets: (1) The task vehicle executes the task locally, and the local execution time is obtained by the following formula: ; where represents the vehicle needs to execute the size of the task data, represents the computing resources required to process each unit size of data for the task, represents the size of the computing resources of the task vehicle locally, is the offloading target; (2) When the task vehicle offloads the task to the edge server or electric vehicle, it needs to be transmitted through the communication link. Compared with local execution, there will be a transmission delay. The execution time of task offloading is obtained by the following formula: ; where represents the size of the computing resources of the edge server computing device, represents the size of the computing resources of the electric vehicle, and respectively represent the task vehicle offloading the task to the transmission time delay of edge server s and electric vehicle g.

[0019] As a further improvement of the present invention, the energy consumption model in step 2 is as follows: When the vehicle executes the task locally The energy consumption model in step 2 is as follows: When the vehicle executes the vehicle offloading task At this time, the energy consumption should be smaller than the energy consumption threshold of the task set by the vehicle. Define the energy consumption threshold of the vehicle offloading task as ; (1) Transmission energy consumption: The transmission energy consumption of the mission vehicle comes from the process of transmitting mission data to the edge server or electric vehicle through the wireless network. In this process, the mission vehicle needs to adjust the transmission power according to factors such as transmission distance, signal strength, and network bandwidth. When the mission vehicle chooses to offload the mission to the edge server or electric vehicle, the transmission delay of the mission is obtained. Define the transmission power of the mission vehicle as , and the energy required to offload the mission to the edge server is: ; the energy required to offload the mission to the electric vehicle is: , where and are the transmission delays for offloading the mission to the edge server and electric vehicle respectively, represents the transmission power of the mission vehicle; (2) Local energy consumption: When the resources required for mission processing are small, the network conditions are not ideal, or the mission processing delay requirement is high, the mission vehicle chooses to execute the mission locally. The mission vehicle needs to make a trade-off between computing resource consumption and network transmission consumption and select the most appropriate processing method. Define the power consumption of the computing device of the mission vehicle as , and the local execution time of the mission is , so the energy consumption required for the mission to be processed locally can be obtained by the following formula: ; (3) Electric vehicle energy consumption: When the mission vehicle chooses to offload the mission to the electric vehicle, the electric vehicle is responsible for processing the offloaded mission. The computing energy consumption of the electric vehicle is related to its processing capacity and the complexity of the executed mission. The electric vehicle is usually in a low-power state and has relatively limited computing resources. Define the power consumption of the computing device of the electric vehicle as , where , and the energy required to process the mission on the computing device of the electric vehicle is: , where represents the vehicle the size of the mission data that needs to be executed, represents the computing resources required to process each unit size of data for the mission, represents the size of the computing resources of the electric vehicle; (4) Edge server energy consumption: The edge computing server is deployed at the network edge, close to the user device. The power consumption of the computing device of the edge server depends on the performance of the server and the complexity of the mission. When making a mission offloading decision, the computing requirements of the mission and the load capacity of the server need to be considered. If the edge server is overloaded, it will lead to an increase in energy consumption and affect the response time of the mission. The energy required for the edge server to process the mission can be calculated by the following formula: , where represents the power consumption of the computing device of the edge server represents the computing resource size of the computing device of the edge server

[0020] As a further improvement of the present invention, in step 2, the payment cost model is set as follows: After the vehicle offloads the task to the edge server or the electric vehicle, the edge server or the electric vehicle will charge the task vehicle the corresponding fee according to its own payment strategy, which is the income of the edge server or the electric vehicle and also one of the costs of the task vehicle; The edge server and the electric vehicle will charge fees for the tasks generated by the task vehicle in terms of memory occupancy, computing resource requirements, and security costs to provide services to ensure their own operating costs. Assuming that the edge server and the electric vehicle price according to their own computing resources and charge according to the amount of data offloaded by the task, the specific calculation method of the fee that the task vehicle m needs to pay for offloading the task is as follows: ; where represents the price charged by the edge server for processing each unit of data volume represents the price charged by the electric vehicle for processing each unit of data volume. Assuming that the fee that the task vehicle needs to pay is related to the data size of the task and the time delay sensitivity of the task, and in order to obtain services and pay the corresponding fees, the total fee that the task vehicle m needs to pay is calculated and defined as: ; represents the vehicle the data size of the task that needs to be executed is the time sensitivity of the task

[0021] As a further improvement of the present invention, in step 3, according to the problem model, the system objective function is designed, and the task offloading cost of the task vehicle is composed of the fee paid by the task vehicle, the time delay of the task vehicle offloading the task, and the energy consumption. Since the data remotely offloaded by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, after formalizing the task offloading problem, the following optimization problem is obtained: ; ; ; ; ; ; ; where represents the payment cost of the task vehicle, represents the task time delay of the task vehicle, e represents the energy consumption of the entire system, represents the maximum time delay threshold of the task, represents the task offloading indicator, is the local execution time of the task, is the execution time of the task offloading, is the transmission time delay for offloading the task to the electric vehicle, is the energy consumption required for the task to be processed locally, is the energy required to offload the task to the edge server, the energy required when offloading the task to the electric vehicle; Constraint 1 means that the local execution time of the task should be less than the time delay threshold of the task vehicle. Constraints 2 and 3 respectively mean that the time delay for offloading the task to the edge server or the electric vehicle should be less than the expected time delay threshold of the task; Constraint 4 means that the offloading target of the task is limited to ; Constraints 5 and 6 mean that the energy consumption of local execution should be less than the expected energy consumption threshold of the task set by the task vehicle.

[0022] As a further improvement of the present invention, in step 4, the proximal policy optimization reinforcement learning method is used to constrain the amplitude of policy update, ensuring training stability while maximizing the cumulative reward of the agent. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by restricting the amplitude of policy update; MOPPO combines the Actor-Critic architecture, uses the actor network to generate the action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, thereby optimizing the direction of policy update; The task vehicle can be regarded as an agent, and the offloading policy of the task vehicle is trained through the MOPPO algorithm, which is described by the Markov decision process represented by the quadruple S, A, R, P, where S is the state space, A is the action space, R is the reward function, and P is the state transition probability; State space: It is necessary to consider the task requirements, network conditions, and available computing resources. Define the state at time slot k as , which consists of the vehicle offloading task of the task tuple and the vehicle edge computing environment, where represents the size of the task data that vehicle m needs to process, represents various types of computing resources required by the task, To represent the time delay sensitivity of a task, represent the expected time delay threshold of the task, represent the expected energy consumption threshold of the task, and F represents the set of computing resources in the environment around the task vehicle; Action space: In the offloading scenario, after receiving the feedback of the state space, the task vehicle will make a decision on the offloading target. The offloading decision is designed as a discrete vector through model design , replacing in the model is the subscript of the offloading target of the task; Reward function: At each time slot k, after the agent executes an action, it obtains a reward when reaching the next state , in general, the reward function should be positively correlated with the objective function. Since the problem objective function is to minimize multiple objectives such as the time delay of the task vehicle in terms of energy consumption and payment cost, the reward function is negatively correlated with the cost of the task vehicle. According to the problem model described in the previous section, the reward function can be designed as follows, where represents the weights of each objective parameter: ; where represents the payment cost of the task vehicle, represents the task time delay of the task vehicle, and e represents the energy consumption of the entire system.

[0023] As a further improvement of the present invention, in step 5, a multi-objective proximal policy optimization algorithm is designed, Actor-Critic architecture: MOPPO adopts the Actor-Critic architecture. The Actor network is responsible for generating policies, and the Critic network is responsible for evaluating the state value function. The Actor network outputs the probability distribution of actions, while the Critic network estimates the value of the current state , which is used to calculate the advantage function , that is, it outputs the probability distribution of actions given a specific state. The Actor network determines actions based on the current state and gradually improves the policy by optimizing the objective function, thereby guiding the policy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean square error of the value function. The advantage function is calculated as follows: ; where represents the discount factor, is the obtained reward, represents state k + 1; Importance sampling mechanism: Importance sampling is a method for adjusting data weights. The MOPPO algorithm uses importance sampling and uses data in the chronological order of the results of the interaction between the agent and the environment; Policy optimization: The core of the MOPPO algorithm is to maximize the expected cumulative reward by optimizing the policy parameters By introducing a proximal constraint to limit the amplitude of the policy update and using a clipped objective function to ensure that the difference between the new policy and the old policy is not too large. The clipped objective function is defined as follows: ; where, The probability ratio policy update ratio of the new policy and the old policy is defined as the probability that the current policy selects an action in the state divided by the selection probability of the old policy in the same state . If this ratio is less than or greater than , then it is clipped to between and , where is the advantage function and is the clipping range parameter. As a further improvement of the present invention, in step 5, the agent is trained to interact with the environment, and the training process is as follows:

[0024] First, initialize the input data required by the algorithm, including the tasks that the vehicle needs to offload, the computing requirements, latency requirements, and energy limitations of the tasks, the edge nodes participating in the offloading calculation of the task by the vehicle, and output the offloading results of the tasks; as well as the costs of task offloading, including energy consumption and time delay; clarify the scale of the task set and its multi-dimensional performance requirements, where the multi-dimensional performance requirements include computing size, energy consumption, and latency sensitivity. Initialize the policy network to generate task offloading action policies, the value function network to evaluate the state value to guide policy optimization, and the experience replay buffer to store historical interaction data. The interaction data includes the old state, action, reward, and new state. Obtain the initial state at each time step, including various aspects of the task attributes, the resource status of the edge server, and the collaborative ability of the vehicle cluster. Based on the current state ,, the policy network outputs an action , selects the offloading target server or local execution. The action probability distribution is optimized through policy gradients, calculates the cost, execution time, and energy consumption of the task at this time step, observes the new state after executing the action, and calculates the immediate reward . Calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , and calculate the reward , in line 10 of the algorithm, the quadruple is stored in buffer D for subsequent policy updates and to calculate the Generalized Advantage Estimation (GAE), quantify the advantages and disadvantages of the quantized actions relative to the average policy, estimate the state value difference through the Critic network, solve the credit assignment problem under sparse rewards, enhance the accuracy of the policy gradient estimation, and calculate the objective function according to the definition formula of the clipped objective function , which is used to update the Actor network and the Critic network, control the degree of policy change to ensure the stability of training, balance the exploration and exploitation of the agent. After calculating the loss values of the Actor network and the Critic network, the Actor network updates the network parameters according to the importance sampling results , and the Critic network updates the parameters through the Generalized Advantage Estimation (GAE) .

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] In the task offloading problem of task vehicles in the vehicle-to-everything (V2X) network, by utilizing the computing resources around the task vehicle, such as edge servers and electric vehicle computing devices, the present invention reduces the system cost, time delay, energy consumption, and payment cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is the architecture model diagram of the offloading system of the present invention;

[0028] Figure 2 is the algorithm framework diagram of the present invention;

[0029] Figure 3 is the flowchart of the MOPPO algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The present invention will be further described in detail below in conjunction with the drawings and the specific embodiments:

[0031] This embodiment provides a task offloading method for electric vehicles, including the following steps:

[0032] In step 1, the system model is as Figure 1As shown, it can be divided into three parts: electric vehicles, task vehicles, and edge servers. Then, a three-layer architecture is designed. The first layer is the edge server, which has powerful computing devices and can quickly process the offloaded tasks, providing services for intensive tasks. The second layer is the electric vehicle, which is composed of vehicles parked around the road and has relatively more idle computing resources for task vehicles to choose for offloading, but the computing resources are relatively smaller than those of the edge server. The third layer is the task vehicle, which generates data to be processed and is equipped with computing devices with small computing capabilities. Intensive tasks need to be offloaded to the edge server or electric vehicle, and different tasks with different service quality requirements need to select for offloading according to the surrounding computing resources.

[0033] In the vehicle-to-everything (V2X) environment, due to the mobility and real-time nature of vehicles, the computing capabilities and computing resources of in-vehicle units equipped in intelligent vehicles are different, and the computing resources of the server in each time slot are also changing. Considering the high mobility of vehicles, vehicles need to complete tasks within the coverage service range of the edge server. When vehicle users cannot process all tasks, they need to offload data to the edge server or electric vehicle to meet their own task requirements.

[0034] (1) Edge server: The edge server has corresponding computing resources, can provide offloading services for intensive task bodies, and formulates a payment strategy according to the various task requirements of task vehicles to obtain benefits. Define the edge server set as and define the number of edge servers as ;

[0035] (2) Electric vehicle: Task vehicles select the computing resources of idle electric vehicles to complete task offloading and execution, and charge task vehicles for the provided offloading services through the formulated payment strategy to encourage more vehicles to provide offloading services. Define the electric vehicle set ;

[0036] (3) Task vehicle: Task vehicles generate tasks locally, offload tasks that cannot be met by local computing resources to the edge server or electric vehicle, select offloading targets according to the various requirements of the tasks to be offloaded, and pay fees according to the payment strategy. Define that each time slot of task vehicles all have tasks that need to be offloaded. Define the task set generated by task vehicles as , is the task subscript of the task vehicle, is the corresponding task, and task consists of the tuple and represents the size of the task data that vehicle needs to execute; Represents the computing resources required for the task to process data per unit size; Represents the offloading target of the offloading indicator. The task vehicle offloads the task request to the edge server, and the offloading target , where when it represents local task processing; when it represents offloading to the corresponding server; when it represents offloading to the corresponding electric vehicle, Represents the time sensitivity of the task. The larger the value, the higher the requirement for the completion time of the task, Represents the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task reduces the quality of service required by the task vehicle;

[0037] The task vehicle needs to complete offloading within the communication coverage range of providing services. The execution time of task offloading should be within the expected completion time threshold of the task, and a task can only be offloaded to one offloading target. If the task vehicle chooses to execute on the local computing device, the offloading target ; if the task is offloaded to the edge server for execution, the offloading target ; if the task is offloaded to the electric vehicle for execution, the offloading target .

[0038] In step 2, the communication model of the system model is designed as follows: Considering that the edge computing model with randomly distributed multiple electric vehicles and multiple edge servers is often applied to large shopping malls and complex commercial district scenarios, a free space path loss model is adopted to describe the process of the task vehicle offloading tasks to the edge server, ignoring the delay generated by the downlink transmission. Assuming that for different mobile vehicle transmissions, an orthogonal frequency division multiple access method is adopted, and each mobile vehicle has a channel bandwidth of B, ignoring the channel interference between mobile vehicles. According to the Shannon formula, the uplink transmission rate of the mobile vehicle offloading the task to the edge server s is . Similarly, the uplink transmission rate of the mobile vehicle offloading the task to the electric vehicle is , where represents the transmission power of the task vehicle, and respectively represent the channel gains between the task vehicle and the edge server s or the electric vehicle . Since the free space path loss model is adopted, the channel gain is inversely proportional to the distance, and respectively represent the variances of Gaussian white noise in the corresponding cases.

[0039] The time model of the design system model is as follows:

[0040] After obtaining the uplink transmission rate between the moving vehicle and the server or the electric vehicle, without considering the speed loss during the transmission process, the task vehicle unloads the task The transmission time delay to the edge server s and the electric vehicle g and are respectively expressed as follows:

[0041]

[0042]

[0043] represents the vehicle The size of the task data that needs to be executed, represents the transmission power of the task vehicle, B is the channel bandwidth owned by each moving vehicle, For the moving vehicle to transfer the task The uplink transmission rate unloaded to the edge server s, For the moving vehicle to transfer the task Unloaded to the electric vehicle The uplink transmission rate of, and respectively represent the variances of Gaussian white noise in the corresponding cases.

[0044] The total completion time of the task usually consists of the transmission time and the execution time. According to the model, the task vehicle can unload part of the task to the edge server, the electric vehicle or execute it locally. Due to different unloading targets, the execution time of the task will be calculated according to different unloading targets;

[0045] Define the set of computing resources of each terminal computing device , where when the subscript of the computing device is 0, it represents the size of the computing resources local to the task vehicle, and the subscript between 1 and s represents the size of the computing resources of the edge server computing device, and the subscript is between to Represents the size of the computing resources of the electric vehicle. The following gives the execution time under different unloading targets:

[0046] (1) The task vehicle executes the task locally, and the local execution time is obtained from the following formula:

[0047]

[0048] where represents the vehicle The size of the task data that needs to be executed, represents the computing resources required to process each unit size of data for the task, Indicates the size of the computing resources local to the task vehicle.

[0049] (2)When the task vehicle offloads tasks to the edge server or electric vehicle, it needs to be transmitted through a communication link. Compared with local execution, there will be a transmission delay. The execution time of task offloading is obtained by the following formula:

[0050]

[0051] Where Indicates the size of the computing resources of the edge server's computing device, Indicates the size of the computing resources of the electric vehicle, And Respectively represent the transmission time delays when the task vehicle offloads tasks To the edge server s and the electric vehicle g.

[0052] The energy consumption model of the designed system model is as follows:

[0053] When the task vehicle has tasks to be executed, there will be many energy consumptions, such as local execution energy consumption and transmission energy consumption. In addition, there are also energy consumptions in the edge server and electric vehicle in the entire offloading system. When the vehicle executes tasks locally The energy consumption should be smaller than the task energy consumption threshold set by the vehicle. Define the task Energy consumption threshold of the vehicle as .

[0054] (1)Transmission energy consumption: The transmission energy consumption of the task vehicle mainly comes from the process of transmitting task data through the wireless network to the edge server or electric vehicle. In this process, the task vehicle needs to adjust the transmission power according to factors such as transmission distance, signal strength, and network bandwidth. Although a higher transmission power can reduce the transmission delay, it will also cause more energy consumption. When the task vehicle chooses to offload tasks to the edge server or electric vehicle, it must consider the network stability and transmission rate to avoid unnecessary energy waste or increased delay due to poor network conditions. According to the above formula, the transmission delay of task transmission can be obtained. Define the transmission power of the task vehicle as , and the energy required to offload the task to the server is: ; The energy required to offload the task to the electric vehicle Is . Where And Are respectively the transmission time delays of offloading tasks to the edge server and the electric vehicle.

[0055] (2) Local energy consumption: When the resources required for task processing are few, the network condition is not ideal, or the task processing delay requirement is high, the task vehicle selects to execute the task locally. Although local processing may reduce the energy consumption of task transmission, it will increase the computing burden of the vehicle itself, resulting in more energy consumption. Therefore, the task vehicle needs to make a trade-off between computing resource consumption and network transmission consumption and select the most appropriate processing method. The power consumption of the computing device of the task vehicle is defined as follows , and the energy consumption required for the task to be processed locally can be obtained by the following formula:

[0056] (3) Electric vehicle energy consumption: When the task vehicle selects to offload the task to an electric vehicle, the electric vehicle is responsible for processing the offloaded task. The computing energy consumption of the electric vehicle is related to its processing capacity and the complexity of the task executed. Electric vehicles are usually in a low-power state and have relatively limited computing resources. The power consumption of the computing device of the electric vehicle is defined as follows , where . The energy required to process the task on the computing device of the electric vehicle is: . Where represents the size of the task data that vehicle needs to execute, represents the computing resources required to process each unit size of data for the task, represents the size of the computing resources of the electric vehicle.

[0057] (4) Edge server energy consumption: Edge computing servers are usually deployed at the network edge, close to user devices, and can provide relatively powerful computing resources. The main advantage of offloading tasks to edge servers is that they can utilize the high-performance computing capabilities of the servers, thus significantly reducing the task processing time. However, this also brings the problem of computing energy consumption. The power consumption of the computing device of the edge server depends on the performance of the server and the complexity of the task. Usually, the computing requirements of the task and the load capacity of the server need to be considered during the task offloading decision. If the edge server is overloaded, it may lead to an increase in energy consumption and affect the task response time. The energy required for the edge server to process the task can be calculated by the following formula: , where represents the power consumption of the computing device of the edge server, represents the size of the computing resources of the computing device of the edge server.

[0058] The payment cost model of the design system model is as follows: Similar to the operator charging relevant fees according to data traffic, after the vehicle offloads tasks to the edge server or electric vehicle, the edge server or electric vehicle will charge the task vehicle corresponding fees according to its own payment strategy, which is the income of the edge server or electric vehicle and also one of the costs of the task vehicle. Since the computing capabilities of various computing devices in the heterogeneous edge environment are different, it is assumed that the fees charged by the edge device are proportional to various resources of its computing device.

[0059] The edge server and electric vehicle will charge fees for the tasks generated by the task vehicle in terms of memory occupancy, computing resource requirements, and security costs to provide services to ensure their own operating costs. To facilitate the calculation of the charging standard, it is assumed that the edge server and electric vehicle price according to their own computing resources and charge according to the amount of offloaded data. The specific calculation method for the fees that the task vehicle m needs to pay for offloading tasks is as follows:

[0060]

[0061] Among them represents the price charged by the edge server for processing each unit of data volume, represents the price charged by the electric vehicle for processing each unit of data volume. It is assumed that the fees that the task vehicle needs to pay are related to the data size of the task and the task time delay sensitivity. To obtain more efficient services, higher fees can be paid. The total fees that the task vehicle m needs to pay can be defined as:

[0062]

[0063] In step 3, according to the problem model, the system objective function is designed. To sum up, it can be obtained that the task offloading cost of the task vehicle consists of the fees paid by the task vehicle, the time delay and energy consumption of the task vehicle for offloading tasks. Since the data remotely offloaded by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, after formalizing the task offloading problem, the following optimization problem can be obtained:

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071] wherein represents the payment cost of the task vehicle, represents the task time delay of the task vehicle, represents the energy consumption of the entire system, represents the maximum time delay threshold of the task, represents the task offloading indicator, represents the task The maximum time delay threshold. Constraint 1 means that the local execution time of the task should be less than the time delay threshold of the task vehicle. Constraints 2 and 3 respectively mean that the time delay of offloading the task to the edge server or the electric vehicle should be less than the expected time delay threshold of the task. Constraint 4 means that the offloading target of the task is limited to ; Constraints 5 and 6 mean that the energy consumption of local execution should be less than the expected energy consumption threshold of the task set by the task vehicle.

[0072] In step 4, the proximal policy optimization reinforcement learning method is used as shown in Figure 2 to constrain the amplitude of policy update, ensuring training stability while maximizing the cumulative reward of the agent. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by restricting the amplitude of policy update;

[0073] MOPPO combines the Actor-Critic architecture, uses the actor network to generate the action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, thereby optimizing the direction of policy update.

[0074] The task vehicle can be regarded as an agent, and the offloading policy of the task vehicle is trained by the MOPPO algorithm. Generally, the offloading problem can be described by a Markov decision process represented by a quadruple (S, A, R, P), where S is the state space, A is the action space, R is the reward function, and P is the state transition probability.

[0075] State space: Reinforcement learning continuously learns policies from historical information, so a comprehensive state definition is crucial for decision-making efficiency. It is necessary to consider task requirements, network conditions, and available computing resources. Define the state at time slot k as , which consists of the task tuple and the vehicle edge computing environment, where represents the vehicle the size of the task data that needs to be processed, represents various types of computing resources required by the task, is the time delay sensitivity representing the task, represents the expected time delay threshold of the task, represents the expected energy consumption threshold of the task, and F represents the set of computing resources in the environment around the task vehicle.

[0076] Action space: In the offloading scenario, after receiving the feedback of the state space, the task vehicle will make a decision on the offloading target. The offloading decision is designed as a discrete vector through model design , replacing in the model is the subscript of the offloading target of the task.

[0077] Reward function: At each time slot k, after the agent executes an action, it obtains a reward when reaching the next state . Generally, the reward function should be positively correlated with the objective function. Since the problem objective function is to minimize multiple objectives such as the time delay, energy consumption, and payment cost of the task vehicle, the reward function is negatively correlated with the cost of the task vehicle. According to the problem model described in the previous section, the reward function can be designed as follows, where represents the weights of each objective parameter:

[0078]

[0079] In step 5, design a multi-objective proximal policy optimization algorithm. The Actor-Critic architecture is shown in Figure 3: MOPPO adopts the Actor-Critic architecture. The Actor network is responsible for generating policies, and the Critic network is responsible for evaluating the state value function. The Actor network outputs the probability distribution of actions, while the Critic network estimates the value of the current state , which is used to calculate the advantage function , that is, it outputs the probability distribution of actions given a specific state. The Actor network decides actions based on the current state and gradually improves the policy by optimizing the objective function, thereby guiding the policy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean square error of the value function. The advantage function is calculated as follows:

[0080] ;

[0081] where represents the discount factor, is the obtained reward, represents state k + 1;

[0082] Importance Sampling Mechanism: Importance sampling is a method for adjusting data weights. The MOPPO algorithm uses importance sampling, but does not use the random sampling and elimination mechanism in experience replay. Instead, it uses data in the chronological order of the interaction results between the agent and the environment to measure the difference between the distribution of the data and the distribution of the current policy.

[0083] Policy Optimization: The core of the MOPPO algorithm is to maximize the expected cumulative reward by optimizing the policy parameters By introducing proximal constraints to limit the amplitude of policy updates and using a clipped objective function to ensure that the difference between the new policy and the old policy is not too large. The clipped objective function is defined as follows:

[0084]

[0085] where, is the policy update ratio of the probability ratio between the new policy and the old policy, defined as the probability of the current policy choosing action in state divided by the selection probability of the old policy in the same state If this ratio is less than or greater than , it is clipped to between and where is the advantage function, is the clipping range parameter.

[0086] Training the Agent Interaction Environment, the training process is as follows:

[0087] First, initialize the input data required by the algorithm, including the tasks that the vehicle needs to offload, and including the computing requirements, latency requirements, and energy limitations of the tasks, the edge nodes participating in the offloading calculation of the tasks of the vehicle, the offloading results of the output tasks, and the costs of task offloading, including energy consumption and time delay, clarify the scale of the task set and its multi-dimensional performance requirements, the multi-dimensional performance requirements include computing size, energy consumption, and latency sensitivity, initialize the policy network to generate task offloading action policies, the value function network evaluates the state value to guide policy optimization, the experience replay buffer stores historical interaction data, the interaction data includes old states, actions, rewards, new states, obtain the initial state at each time step, including various aspects of the task attributes, the resource status of the edge server, and the collaborative ability of the vehicle cluster, based on the current state , the policy network outputs action , select to unload on the target server or execute locally. The action probability distribution is optimized through policy gradients, and the cost, execution time, and energy consumption of the task at this time step are calculated. After executing the action, observe the new state and calculate the immediate reward , calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , in line 10 of the algorithm, the quadruple is stored in buffer D for subsequent policy updates, and the generalized advantage function is calculated , quantify the pros and cons of the action relative to the average policy, estimate the state value difference through the Critic network, solve the credit assignment problem under sparse rewards, enhance the accuracy of policy gradient estimation, and calculate the objective function according to the definition formula of the clipped objective function , used to update the Actor network and the Critic network, control the degree of policy change to ensure the stability of training, balance the exploration and exploitation of the agent. After calculating the loss values of the Actor network and the Critic network, the Actor network updates the network parameters according to the importance sampling results , the Critic network updates the parameters through the generalized advantage function GAE .

[0088] The following is an introduction to the settings of each parameter in the experiment. All neural networks in the experiment are implemented using the PyTorch 2.0.0 and Python 3.8.1 platforms, and Adam is used to optimize the network. The performance of the algorithm proposed in this application is evaluated through a large number of simulation experiments. Suppose the amount of data generated by a task vehicle follows a normal distribution, which is . The generated data should be processed in real time and deleted after returning the processing results to avoid storage overflow on the local or edge vehicles. Suppose the transmission power of the task vehicle is 10 Kwh, and the computing resources of the local computing device of the task vehicle at time slot t are , and the computing resources of the edge server are defaulted to , and the average computing resources of the vehicle cooperation cluster are defaulted to .

[0089] In the experimental simulation of this application, the Actor network has two hidden fully connected layers, each layer containing 128 nodes respectively. The Critic network also has two hidden fully connected layers, each layer having 128 nodes respectively. Each task vehicle simulates the real scenario according to its own task requirements, and calculates the unloading results of the task vehicle through the unloading algorithm based on multi-objective proximal policy gradient optimization. The experiment selects three baseline methods to compare with the algorithm proposed in this application, namely the Q-learning algorithm, the random unloading algorithm, and the full-local algorithm. The main experimental parameter settings of the algorithm in this application are shown in Table 1:

[0090] Table 1 Main Parameter Settings of the Experiment

[0091] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still falls within the scope of the present invention claimed.

Claims

1. A task offloading method for electric vehicles, characterized in that: It includes the following specific steps: Step 1: Considering the vehicle movement scenario, design an edge-end system offloading framework; Step 2: The task vehicle generates intensive tasks, and design a communication model, an energy model, a time model, and a payment cost model; Step 3: According to the problem model, design a system objective function to multi-objectively optimize the system time delay, energy consumption, and payment cost; Step 4: According to the system model and the reinforcement learning method, design a vehicle agent interaction environment; Step 5: Design a multi-objective proximal policy optimization algorithm, and determine the offloading target after training the agent.

2. A task offloading method for electric vehicles according to claim 1, wherein: In the step 1, the system model of the edge-end system offloading framework is divided into three parts: electric vehicles, task vehicles, and edge servers; The edge-end system offloading framework has a three-layer architecture; The first layer is the edge server, which quickly processes the offloaded tasks and provides services for intensive tasks; The second layer is the electric vehicle, which is composed of vehicles parked around the road and has idle computing resources for the task vehicle to choose for offloading; The third layer is the task vehicle, which will generate data to be processed, and the task vehicle itself is equipped with a computing device with small computing power. Intensive tasks need to be offloaded to the edge server or the electric vehicle, and different tasks with different service quality requirements need to be offloaded according to the surrounding computing resources; In the step 1, in the vehicle movement scenario of the vehicle networking environment, considering the mobility of the vehicle, the vehicle needs to complete tasks within the coverage service range of the edge server. When the vehicle user cannot process all tasks, the data needs to be offloaded to the edge server or the electric vehicle to meet its own task requirements, where the edge server, the electric vehicle, and the task vehicle are defined as follows: (1) Edge server: The edge server has corresponding computing resources, can provide offloading services for intensive task bodies, and formulates a payment strategy according to the various task requirements of the task vehicle to obtain benefits. Define the edge server set as , and define the number of edge servers as ; (2) Electric vehicle: The task vehicle selects the computing resources of idle electric vehicles to complete the unloading and execution of the task. Through the payment strategy formulated, the task vehicle is charged for the unloading service provided to encourage more vehicles to provide unloading services. The electric vehicle set is defined ; (3) Task vehicle: The task vehicle generates tasks locally and offloads tasks that cannot be satisfied by local computing resources to the edge server or electric vehicle, selects the offloading target according to the various requirements of the tasks to be offloaded, and pays the fee according to the payment policy. Define each time slot of the task vehicle All have Tasks need to be offloaded. Define the task set generated by the task vehicle as , is the subscript of the task vehicle, is the corresponding task. The vehicle offloads the task consists of the tuple composed of represents the vehicle The size of the task data that needs to be executed; represents the computing resources required to process each unit size of data for the task; represents the offloading target of the offloading indicator. The task vehicle offloads the task request to the edge server, and the offloading target , where when indicates local task processing; when indicates offloading to the corresponding server; when indicates offloading to the corresponding electric vehicle, represents the time sensitivity of the task. The larger the value, the higher the requirement for the completion time of the task, represents the expected completion time threshold of the task. If the task execution time exceeds this threshold, it means that the task reduces the quality of service required by the task vehicle; The task vehicle needs to complete the offloading within the communication coverage area where the service is provided. The execution time of the task offloading should be within the expected completion time threshold of the task, and a task can only be offloaded to one offloading target. If the task vehicle chooses to execute on the local computing device, the offloading target ; If the task is offloaded to the edge server for execution, the offloading target ; if the task is offloaded to the electric vehicle for execution, the offloading target , where s is the edge server and is the electric vehicle.

3. A task offloading method for electric vehicles according to claim 1, wherein: The communication model designed in the step 2 is as follows: The free space path loss model is adopted to describe the process of task offloading from the task vehicle to the edge server. The delay generated by the downlink transmission is ignored. It is assumed that for different mobile vehicle transmissions, the orthogonal frequency division multiple access method is adopted, and each mobile vehicle has a channel bandwidth of B. The channel interference between mobile vehicles is ignored. According to the Shannon formula, the mobile vehicle offloads the vehicle offloading task The uplink transmission rate of offloading to the edge server s is , and similarly, the mobile vehicle can obtain the task The uplink transmission rate of offloading to the electric vehicle is , where represents the transmission power of the task vehicle, and represent the channel gains between the task vehicle and the edge server s or the electric vehicle respectively. Since the free space path loss model is adopted, the channel gain is inversely proportional to the distance, and represent the variances of the Gaussian white noise in the corresponding cases respectively.

4. A task offloading method for electric vehicles according to claim 1, wherein: The time model in the step 2 is as follows: After obtaining the uplink transmission rate between the moving vehicle and the server or the electric vehicle, without considering the speed loss during the transmission process, the task vehicle unloads the task The transmission time delay to the edge server s and the electric vehicle g and Are respectively expressed as follows: ; ; Represents a vehicle The size of the task data to be executed Represents the transmission power of the task vehicle. B is the channel bandwidth owned by each mobile vehicle For the mobile vehicle to offload the vehicle offloading task The uplink transmission rate for offloading to the edge server s For the mobile vehicle to offload the vehicle offloading task To the electric vehicle The uplink transmission rate And Respectively represent the variances of Gaussian white noise in the corresponding cases The total completion time of the task consists of the transmission time and the execution time. According to the model, the task vehicle can offload part of the task to the edge server, the electric vehicle, or execute it locally. Due to different offloading targets, the execution time of the task will be calculated according to different offloading targets; Define the computing resource sets of each terminal computing device , where when the subscript of the computing device is 0, it represents the computing resource size of the task vehicle locally, the subscript between 1 and s represents the computing resource size of the edge server computing device, and the subscript between s + 1 and s + g represents the computing resource size of the electric vehicle. The execution times under different offloading objectives are given below: (1) When the task vehicle executes the task locally, the local execution time is obtained by the following formula: ; Among them represents the vehicle the size of the task data that needs to be executed, represents the computing resources required to process each unit size of data for the task, represents the size of the computing resources local to the task vehicle, is the offloading target; (2) When the task vehicle offloads the task to the edge server or the electric vehicle, it needs to be transmitted through the communication link, and there will be a transmission delay compared with local execution. The execution time of the task offloading is obtained by the following formula: ; Among them represents the computing resource size of the edge server computing device represents the computing resource size of the electric vehicle and respectively represent the transmission time delay for the task vehicle to offload tasks to the edge server s and the electric vehicle g 5. A task offloading method for electric vehicles according to claim 1, wherein: The energy consumption model in the step 2 is as follows: When the vehicle performs a vehicle unloading task locally the energy consumption should be less than the energy consumption threshold of the vehicle's set task. Define the energy consumption threshold of the vehicle unloading task performed by the vehicle as ; (1) Transmission energy consumption: The transmission energy consumption of the mission vehicle comes from the process of transmitting mission data to the edge server or electric vehicle through the wireless network. In this process, the mission vehicle needs to adjust the transmission power according to factors such as transmission distance, signal strength, and network bandwidth. When the mission vehicle chooses to offload the mission to the edge server or electric vehicle, the transmission delay of the mission is obtained. Define the transmission power of the mission vehicle as , and the energy required to offload the mission to the edge server is: ; the energy required to offload the mission to the electric vehicle is: , where and are the transmission delays for offloading the mission to the edge server and electric vehicle respectively, represents the transmission power of the mission vehicle; (2) Local energy consumption: When the resources required for task processing are less, the network conditions are not ideal, or the task processing delay requirement is high, the task vehicle chooses to execute the task locally. The task vehicle needs to make a trade-off between computing resource consumption and network transmission consumption to select the most appropriate processing method. The power consumption of the computing device of the task vehicle is defined as , and the local execution time of the task is . Therefore, the energy consumption required for local task processing can be obtained by the following formula: ; (3) Electric vehicle energy consumption: When the task vehicle chooses to offload the task to the electric vehicle, the electric vehicle is responsible for processing the offloaded task. The computing energy consumption of the electric vehicle is related to its processing capacity and the complexity of the executed task. The electric vehicle is usually in a low-power state and has relatively limited computing resources. The power consumption of the computing device of the electric vehicle is defined as follows , where . The energy required to process the task on the computing device of the electric vehicle is: , where represents the vehicle the size of the task data that needs to be executed, represents the computing resources required to process each unit size of data for the task, represents the size of the computing resources of the electric vehicle; (4) Edge server energy consumption: Edge computing servers are deployed at the network edge, close to user devices. The power consumption of the computing devices of edge servers depends on the performance of the servers and the complexity of the tasks. When making task offloading decisions, the computing requirements of the tasks and the load capacity of the servers need to be considered. If the edge server is overloaded, it will lead to an increase in energy consumption and affect the response time of the tasks. The energy consumed by the edge server when processing tasks can be calculated by the following formula: , where represents the power consumption of the computing devices of the edge server, represents the computing resource size of the computing devices of the edge server.

6. A task offloading method for electric vehicles according to claim 1, wherein: In step 2, the payment cost model is set as follows: After the vehicle offloads the task to the edge server or electric vehicle, the edge server or electric vehicle will charge the task vehicle corresponding fees according to its own payment strategy, which is the income of the edge server or electric vehicle and also one of the costs of the task vehicle; The edge server and electric vehicle will charge fees for the tasks generated by the task vehicle in terms of memory occupancy, computing resource requirements, and security costs to provide services to ensure their own operating costs. Assuming that the edge server and electric vehicle price according to their own computing resources and charge according to the amount of offloaded task data, the specific calculation method of the fees that the task vehicle m needs to pay for offloading tasks is as follows: ; Among them represents the price charged by the edge server for processing each unit of data volume represents the price charged by the electric vehicle for processing each unit of data volume. Assuming that the cost to be paid by the task vehicle is related to the data size of the task and the task time delay sensitivity, and corresponding fees are paid for obtaining services, the total cost to be paid by the task vehicle m is calculated and defined as: ; Indicates the vehicle The size of the task data that needs to be executed Indicates the time sensitivity of the task 7. A task offloading method for electric vehicles according to claim 1, characterized in that: In step 3, according to the problem model, the system objective function is designed, and the task offloading cost of the task vehicle is obtained, which is composed of the fees paid by the task vehicle, the time delay and energy consumption of the task vehicle for offloading tasks. Since the data remotely offloaded by the task vehicle is completed by the target computing device, it is equivalent to purchasing the energy consumption of the offloading target computing device and saving the task completion time. According to the above model, the following optimization problem is obtained after formalizing the task offloading problem: ; ; ; ; ; ; ; Among them represents the payment cost of the task vehicle represents the task time delay of the task vehicle, and e represents the energy consumption of the entire system represents the maximum time delay threshold of the task represents the task offloading indicator is the local execution time of the task is the execution time of task offloading is the transmission time delay for offloading the task to the electric vehicle is the energy consumption required for the task to be processed locally is the energy consumption required for offloading the task to the edge server is the energy consumption required when offloading the task to the electric vehicle Constraint 1 indicates that the local execution time of the task should be less than the delay threshold of the task vehicle. Constraints 2 and 3 respectively indicate that the time delay of offloading the task to the edge server or the electric vehicle should be less than the expected time delay threshold of the task. Constraint 4 indicates that the offloading target of the task is limited to ; Constraints 5 and 6 indicate that the energy consumption of local execution should be less than the expected energy consumption threshold of the task set by the task vehicle.

8. A task offloading method for electric vehicles according to claim 1, characterized in that: In step 4, the proximal policy optimization reinforcement learning method is used to constrain the amplitude of policy update, ensuring training stability while maximizing the cumulative reward of the agent. PPO-Clip ensures that the difference between the new policy and the old policy is within a controllable range by restricting the amplitude of policy update; MOPPO combines the Actor-Critic architecture, uses the actor network to generate the action distribution, and estimates the state value through the critic network to calculate the generalized advantage function, so as to optimize the direction of policy update; The task vehicle can be regarded as an agent, and the offloading strategy of the task vehicle is trained through the MOPPO algorithm, which is described by a Markov decision process represented by the quadruple S, A, R, P, where S is the state space, A is the action space, R is the reward function, and P is the state transition probability; State space: Considering the task requirements, network conditions, and available computing resources, the state of time slot k is defined as , consisting of the vehicle offloading task of the task tuple and the vehicle edge computing environment, where represents the size of the task data that vehicle m needs to process, represents the various types of computing resources required by the task, is the time delay sensitivity representing the task, represents the expected time delay threshold of the task, represents the expected energy consumption threshold of the task, and F represents the set of computing resources in the environment around the task vehicle; Action space: In the unloading scenario, after receiving the feedback of the state space, the task vehicle will make a decision on the unloading target, and design the unloading decision as a discrete vector through the model , and replace in the model with the subscript of the unloading target of the task; Reward function: At each time slot k, after the agent executes an action, it obtains a reward when reaching the next state , in general, the reward function should be positively correlated with the objective function. Since the objective function of the problem is to minimize multiple objectives such as the time delay, energy consumption, and payment cost of the task vehicle, the reward function is negatively correlated with the cost of the task vehicle. According to the problem model described in the previous section, the reward function can be designed as follows, where represents the weights of each objective parameter: ; Among them represents the payment cost of the task vehicle represents the task time delay of the task vehicle, and e represents the energy consumption of the entire system 9. A task offloading method for electric vehicles according to claim 8, characterized in that: In step 5, a multi-objective proximal policy optimization algorithm, the Actor-Critic architecture, is designed: MOPPO adopts the Actor-Critic architecture. The Actor network is responsible for generating policies, and the Critic network is responsible for evaluating the state value function. The Actor network outputs the probability distribution of actions, while the Critic network estimates the value of the current state , which is used to calculate the advantage function , that is, it outputs the probability distribution of actions given a specific state. The Actor network determines actions based on the current state and gradually improves the policy by optimizing the objective function, thereby guiding the policy update of the Actor network. The Critic network optimizes its parameters by minimizing the mean squared error of the value function. The advantage function is calculated as follows: ; wherein represents a discount factor, to obtain a reward, represents state k + 1; Importance sampling mechanism: Importance sampling is a method to adjust the data weights. The MOPPO algorithm uses importance sampling and uses the data in the chronological order of the interaction results between the agent and the environment; Policy optimization: The core of the MOPPO algorithm is to maximize the expected cumulative reward by optimizing the policy parameters To achieve this, the MOPPO algorithm introduces a proximal constraint to limit the magnitude of policy updates and uses a clipped objective function to ensure that the difference between the new and old policies is not too large. The clipped objective function is defined as follows: ; Among them, The probability ratio policy update ratio of the new policy and the old policy is defined as the probability that the current policy selects an action in the state and the selection probability of the old policy in the same state If the ratio is less than or greater than then it is clipped to between and where is the advantage function and is the clipping range parameter. ​ 10. A task offloading method for electric vehicles according to any one of claims 1-9, characterized in that: In step 5, the agent interaction environment is trained, and the training process is as follows: First, initialize the input data required by the algorithm, including the tasks to be offloaded by the vehicle, the computing requirements, latency requirements, and energy limits of the tasks, the edge nodes participating in the offloading calculation of the tasks by the vehicle, and output the offloading results of the tasks; as well as the costs of task offloading, including energy consumption and time delay; clarify the scale of the task set and its multi-dimensional performance requirements, where the multi-dimensional performance requirements include computing size, energy consumption, and latency sensitivity. Initialize the policy network to generate task offloading action policies, the value function network to evaluate the state value to guide policy optimization, and the experience replay buffer to store historical interaction data, where the interaction data includes old states, actions, rewards, and new states. Obtain the initial state at each time step, including various aspects of the task attributes, the resource status of the edge server, and the collaborative ability of the vehicle cluster. Based on the current state , the policy network outputs an action , selects the offloading target server or local execution. The action probability distribution is optimized through policy gradients, and the cost, execution time, and energy consumption of the task at this time step are calculated. After executing the action, observe the new state and calculate the immediate reward , calculate the reward when the execution time or energy consumption exceeds the expected value of the task vehicle , line 10 of the algorithm stores the quadruple in buffer D for subsequent policy updates, and calculate the generalized advantage function , quantify the pros and cons of the action relative to the average policy, estimate the state value difference through the Critic network, solve the credit assignment problem under sparse rewards, and enhance the accuracy of policy gradient estimation. Calculate the objective function according to the definition formula of the clipped objective function , used to update the Actor network and the Critic network, control the degree of policy change to ensure the stability of training, balance the exploration and exploitation of the agent. After calculating the loss values of the Actor network and the Critic network, the Actor network updates the network parameters according to the importance sampling results , and the Critic network updates the parameters through the generalized advantage function GAE .

Citation Information

Patent Citations

  • Macro base station placement method based on Internet of Vehicles task unloading decision

    CN113810878A

  • Calculation unloading and privacy protection joint optimization method for mobile block chain network

    CN116437341A

  • Method for optimizing dependency task unloading in Internet of Vehicles by using deep reinforcement learning

    CN116455903A

  • Method and system for selecting automobile driving mode and storage medium

    CN116811879A

  • Markov chain-based differential privacy task unloading method in Internet of Vehicles environment

    CN117749797A