An Optimization Method for Task Offloading and Resource Allocation in a Vehicle-Assisted Edge Computing Network

By adopting the combination of multi-agent deep reinforcement learning and particle swarm algorithm in vehicle edge computing, the problems of task complexity and insufficient resource utilization are solved, and the task execution delay is reduced and resource utilization is optimized.

CN119316883BActive Publication Date: 2025-07-01NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411848220.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-07-01
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The prior art fails to fully consider task complexity, vehicle resource utilization flexibility, and cache update strategies in vehicle edge computing, resulting in low task scheduling efficiency and insufficient resource utilization.

Method used

Using a method combining multi-agent deep reinforcement learning and particle swarm algorithm, a task offload and resource allocation optimization model is built, and the sub-task offload decision and resource allocation decision are determined through step-by-step solutions to optimize system performance.

Benefits of technology

The goal of reducing task execution delay, optimizing resource utilization and reducing system costs has been achieved, and the performance of vehicle-assisted edge computing network has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316883B_ABST
    Figure CN119316883B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimization method for task offloading and resource allocation in a vehicle-assisted edge computing network. The method includes: establishing a basic model of the vehicle-assisted edge computing network; establishing a cost evaluation and optimization model for vehicle-assisted edge computing tasks to quantify costs and solve optimization problems; according to system state information, using a multi-agent deep reinforcement learning method based on MADDPG to obtain offloading decisions for subtasks; according to the offloading decisions, using a particle swarm method based on PSO to obtain transmission power and computing resource allocation decisions; after completing the decisions, updating system state information, buffer buffer information, and sampling and training a multi-agent neural network. This method can solve problems such as limited resources of vehicle devices, heavy burden on MEC servers, and complex task offloading decisions, achieve the goals of reducing task execution latency, optimizing resource utilization rate, and reducing system costs, and improve the performance of the vehicle-assisted edge computing network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of Internet of Vehicles and mobile edge computing, and particularly relates to an optimization method for task offloading and resource allocation in a vehicle-assisted edge computing network. Background Art

[0002] With the in-depth development of Internet of Things (IoT) technology in the transportation field, Internet of Vehicles (IoV) has emerged as a focus of common concern in academia and industry. By constructing a network topology among communication entities such as vehicles, pedestrians, and roadside units (RSUs), IoV aims to provide efficient and secure information services for the traffic environment to meet the growing traffic demands. Among them, vehicle edge computing (VEC), as a key technology in IoV, significantly improves the task processing ability of vehicles by offloading computationally intensive and latency-sensitive tasks to mobile edge computing (MEC) servers or other vehicles with computing capabilities, providing strong support for applications such as intelligent driving and augmented reality (AR) navigation.

[0003] In the field of vehicular edge computing, many studies have been dedicated to optimizing task offloading and resource allocation strategies to improve system performance. Some studies focus on the strategy of offloading tasks to RSUs. For example, Dai et al. (P. Dai, K. Hu, X. Wu, H. Xing, F. Teng, and Z. Yu, "A probabilistic approach for cooperative computation offloading in MEC-assisted vehicular networks," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 2, pp. 899-911, 2022.) proposed a joint load balancing and task offloading scheme aiming to maximize system utilization; Wang et al. (H. Wang, T. Lv, Z. Lin, and J. Zeng, "Energy-delay minimization of task migration based on game theory in MEC-assisted vehicular networks," IEEE Transactions on Vehicular Technology, vol. 71, no. 8, pp. 8175 - 8188, 2022.) used game theory methods to enable vehicles to determine task offloading strategies in real time. However, these studies often neglect the computing resources of the vehicles themselves and tend to offload all tasks to edge nodes, which may lead to insufficient resource utilization.

[0004] Although existing studies have made some progress in task offloading and resource allocation in vehicular edge computing, there are still many deficiencies. Many schemes fail to fully consider the complexity of tasks, such as dependencies between tasks, dynamic changes in data transmission and computing resources, etc., resulting in low task scheduling efficiency. Some studies are not flexible enough in using vehicle resources and do not fully exploit the potential of vehicle idle resources, restricting the improvement of the overall system performance. At the same time, the lack of a vehicle cache update strategy makes the data acquisition efficiency during task processing low and increases the task processing delay. Therefore, there is an urgent need for an optimization scheme that comprehensively considers task complexity, vehicle resource utilization flexibility, and cache update to improve the performance of vehicular-assisted edge computing networks and meet the growing demands of intelligent transportation applications. Summary of the Invention

[0005] The object of the present invention is to provide a method for optimizing task offloading and resource allocation in a vehicle-assisted edge computing network, so as to solve the problems existing in the prior art, such as limited resources of vehicle devices, heavy burden on MEC servers, complex task offloading decisions, etc., and achieve the goals of reducing task execution latency, optimizing resource utilization rate, and reducing system costs.

[0006] The technical solution for achieving the object of the present invention is as follows: On the one hand, a method for optimizing task offloading and resource allocation in a vehicle-assisted edge computing network is provided. The method includes the following steps:

[0007] Step 1, construct a basic model of the vehicle-assisted edge computing network, including system architecture, task characteristics, service cache management, and communication capabilities;

[0008] Step 2, construct a task cost evaluation and optimization model, quantify the cost and solve the optimization problem, and clarify the goal, task completion time, and offloading constraints;

[0009] Step 3, according to the system state, calculate the sub-task scheduling order and completion time point, and use the multi-agent deep reinforcement learning method based on MADDPG to determine the sub-task offloading decision, which is realized by defining relevant elements through an algorithm framework;

[0010] Step 4, according to the offloading decision, use the particle swarm method based on PSO to solve the transmission power and computing resource allocation decision, and determine the optimal solution through particle coding, defining the fitness function, and iterative update;

[0011] Step 5, after completing the above decisions, update the system state, buffer buffer information, and train the multi-agent neural network, including the vehicle service program and the update process of the multi-agent reinforcement learning algorithm, to continuously optimize the system performance.

[0012] On the other hand, a system for optimizing task offloading and resource allocation in a vehicle-assisted edge computing network is provided. The system includes:

[0013] The first module is used to construct a basic model of the vehicle-assisted edge computing network, including system architecture, task characteristics, service cache management, and communication capabilities;

[0014] The second module is used to construct a task cost evaluation and optimization model, quantify the cost and solve the optimization problem, and clarify the goal, task completion time, and offloading constraints;

[0015] The third module is used to calculate the sub-task scheduling order and completion time point according to the system state, and use the multi-agent deep reinforcement learning method based on MADDPG to determine the sub-task offloading decision, which is realized by defining relevant elements through an algorithm framework;

[0016] The fourth module is used to solve the transmission power and calculate the resource allocation decision by using the particle swarm method based on PSO according to the offloading decision, and determine the optimal solution through particle coding, defining the fitness function, and iterative update;

[0017] The fifth module is used to update the system state, buffer buffer information, and train the multi-agent neural network after completing the above decisions, including the vehicle service program and the multi-agent reinforcement learning algorithm update process, so as to continuously optimize the system performance.

[0018] Compared with the prior art, the remarkable advantages of the present invention are as follows:

[0019] (1) In task processing, the complexity of tasks is fully considered. The constructed basic model comprehensively covers the system architecture, task characteristics, cache management, and communication capabilities, etc., accurately establishes the cost evaluation and optimization model, comprehensively considers various costs and constraint conditions, and makes up for the deficiencies in the existing research that do not fully consider task dependencies, resource dynamic changes, etc.

[0020] (2) For the task offloading and resource allocation optimization method, it is solved step by step. First, the sub-task offloading decision is determined based on the MADDPG multi-agent deep reinforcement learning algorithm, and the corresponding service program is pulled from the edge server accordingly. Then, the resource allocation decision is made based on the PSO particle swarm algorithm, effectively improving the problems existing in the offloading decision and resource utilization of the prior art.

[0021] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0022] Figure 1 It is a schematic flow chart of the task offloading and resource allocation optimization method for the vehicle-assisted edge computing network.

[0023] Figure 2 It is a schematic diagram of the system architecture model. Detailed Embodiments

[0024] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0025] It should be noted that if the descriptions such as "first" and "second" are involved in the embodiments of the present invention, these descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0026] The object of the present invention is to provide an optimization method for task offloading and resource allocation in a vehicle-assisted edge computing network to solve the problems existing in the prior art, such as limited resources of vehicle devices, heavy burden on MEC servers, complex task offloading decisions, etc., and achieve the goals of reducing task execution latency, optimizing resource utilization rate, and reducing system costs.

[0027] In one embodiment, in combination with Figure 1 and Figure 2 , an optimization method for task offloading and resource allocation in a vehicle-assisted edge computing network is provided. The method includes the following steps:

[0028] Step 1, construct a basic model of the vehicle-assisted edge computing network, covering key elements such as system architecture, task characteristics, service cache management, and communication capabilities, to provide a basis for subsequent optimization.

[0029] Step 2, construct a task cost evaluation and optimization model, quantify the cost, and solve the optimization problem, clarifying the objectives and constraints such as task completion time and offloading.

[0030] Step 3, according to the system state, calculate the scheduling order and completion time points of subtasks, and use the multi-agent deep reinforcement learning method based on MADDPG to determine the subtask offloading decision, which is achieved through the definition of relevant elements in the algorithm framework.

[0031] Step 4, according to the offloading decision, use the particle swarm method based on PSO to solve the transmission power and computing resource allocation decisions, and determine the optimal solution through particle coding, definition of fitness function, and iterative update.

[0032] Step 5, after completing the above decisions, update the system state, buffer buffer information, and train the multi-agent neural network, including the vehicle service program and the update process of the multi-agent reinforcement learning algorithm, to continuously optimize the system performance.

[0033] Further, in one of the embodiments, step 1 specifically includes:

[0034] Step 1-1, construct a vehicle-assisted edge computing network system model, specifically: obtain vehicle resource information in the urban environmental scenario, divide time slots and define the set of task vehicles, clarify the RSU resource information, and establish a decision-making module on the cloud server to process the status information uploaded by vehicles and RSUs and generate decisions; specifically including:

[0035] Step 1-1-1, obtain vehicle resource information in the urban environmental scenario: Consider an urban environmental scenario with parallel lanes and vehicle driving, and the set of vehicles in the system , the i-th vehicle is represented as , where , is the total number of vehicles, is the driving speed, is the maximum transmission power, is the maximum computing frequency, is the maximum cache capacity.

[0036] Step 1-1-2, divide time slots. Divide the time into these equal time slots, and each vehicle can generate at most one task within one time slot. Within the time slot , define the set to represent the set of vehicles that generate tasks, represents the th vehicle.

[0037] Step 1-1-3, obtain RSU resource information in the urban environmental scenario. RSUs are deployed on one side of the road, denoted as , represents the M-th RSU deployed, and adjacent RSUs are connected by optical fibers. According to the coverage range of RSUs, the road is divided into sections , is the M-th section, and M is the total number of sections. All RSUs are equipped with MEC servers to provide computing offloading services for vehicles.

[0038] Here, preferably, for simplicity, an MEC server and the connected RSU are represented by a fixed edge node.

[0039] Step 1-1-4, establish a decision-making module. Deploy an intelligent agent decision-making module on the cloud server to process the status information uploaded by vehicles and RSUs and generate decisions.

[0040] Step 1-2, establish a task model. Include defining the task type and the required cache capacity, determining the RSU service program rental rule; establish a vehicle task model, clarify the number of subtasks, information and dependency relationship; introduce three offloading methods for local execution, offloading to RSU and other vehicles of subtasks.

[0041] Step 1-2-1, Task type definition. Define the set of task types as , and the i-th task type requires a cache capacity of ; assume that the RSU has deployed the service programs corresponding to all task types. After the vehicle makes a sub-task offloading decision, it updates the service program lease according to the sub-task offloading situation to assist itself and other vehicles in sub-task offloading, and the unit price for leasing the service program from the RSU is .

[0042] Step 1-2-2, Establish a task model. The task generated by the vehicle within a time slot can be expressed as , where is the number of sub-tasks in the task; is the sub-task information; is the matrix representing the sub-task dependency relationship; indicates that there is a direct dependency relationship between sub-tasks and ; indicates that there is no direct dependency relationship between sub-tasks and ; is the task type ; is the task submission time point; is the maximum tolerance time of the task.

[0043] Step 1-2-3, Establish a sub-task model. Assume that the sub-task cannot be further divided, and define the -th sub-task of task as , where represents the input data size of the sub-task; represents the number of CPU cycles required to complete the sub-task.

[0044] Step 1-2-4, Sub-task offloading method. In the ITS scenario, after the vehicle generates the task , it can process the sub-task in three ways: local execution, offloading to a fixed RSU node, and offloading to a qualified auxiliary vehicle. For the sub-task , there are offloading decisions, and their task offloading symbol factors are respectively , , .

[0045] Step 1-2-4-1, Offload to local. When the sub-task is executed locally , otherwise .

[0046] Step 1-2-4-2, offload to the RSU. Subtask When offloading to a fixed RSU node , otherwise .

[0047] Step 1-2-4-3, offload to other vehicles. Subtask Offload to other vehicles When , otherwise .

[0048] Step 1-2-4-4, definition of subtask offloading decision. The offloading decision of is defined as , and the relationship between the offloading symbol factor and the offloading decision is: if then , if then , if then , indicating that decides to offload the subtask to .

[0049] Step 1-3, establish a service cache model. It includes the vehicle leasing service programs based on subtask offloading information and the storage space constraint of service program caching.

[0050] Step 1-3-1, leasing service conditions. After the subtask makes an offloading decision, the vehicle updates the service program lease according to the offloading information. When the vehicle processes a subtask, it first checks whether there is a service program for processing the subtask in its own cache space. If there is, it processes directly; if not, it leases from the RSU. The vehicle cache space has an upper limit. When there is no extra space and a service needs to be leased, the FIFO strategy is used to replace the saved service program.

[0051] Step 1-3-2, vehicle storage space limit. For the caching of service programs, there are the following storage capacity limits:

[0052] (1)

[0053] Among them, is a binary variable indicating whether the vehicle caches a service program with a cache capacity of , indicates that the service is cached, otherwise it indicates that it is not cached, and finally forms ​, indicating whether each vehicle has cached a certain service program. Define A binary variable indicating whether it is a service program newly cached within the current time slot, Indicates the number of types of service programs, Indicates the vehicle 's cache space.

[0054] Steps 1 - 4, establish a communication model. It includes vehicle - RSU communication and vehicle - vehicle communication.

[0055] Step 1 - 4 - 1, V2I communication model. Each vehicle in the system uses the V2I communication method to establish a communication connection with the RSU. The communication bandwidth between the vehicle and the RSU is , and multiple subtasks are offloaded to the edge server node . The vehicle and the fixed RSU node The uplink data rate between Indicates the number of types of service programs, Indicates the vehicle The calculation formula for the cache space is:

[0056] (2)

[0057] Where, is 's upload transmit power, is and The instantaneous channel gain between is the instantaneous distance between the two, and respectively represent and The path loss and fading factor between Indicates the background noise power, , is the total number of vehicles generating tasks.

[0058] Step 1 - 4 - 2, V2V communication model. The i - th vehicle and the g - th vehicle The bandwidth resources they own are , and multiple subtasks are offloaded to the vehicle . The achievable transmission rate between the vehicle and the vehicle is expressed as:

[0059] (3)

[0060] Where is and The instantaneous channel gain between is the instantaneous distance between the two, and respectively represent the path loss and fading factor between and

[0061] Furthermore, in one of the embodiments, step 2 is to establish a vehicle-assisted edge computing task cost evaluation and optimization model. The establishment of this model needs to comprehensively consider various factors, clarify the objectives, task completion time, offloading and other constraints. Specifically, it includes:

[0062] Step 2-1, establish a computing model. It mainly elaborates on the execution delay, energy consumption calculation and their formulaic representations under different processing methods of vehicle subtasks.

[0063] Step 2-1-1, local processing. If the vehicle generates a subtask selects local execution, then the execution time is:

[0064] (4)

[0065] where is the computing resource allocated by the vehicle to the subtask , , .

[0066] The energy consumption of local execution:

[0067] (5)

[0068] where is the effective switching capacitance in the chip.

[0069] Step 2-1-2, MEC server processing. If the vehicle generates a subtask that is offloaded to a fixed RSU node for execution, and if the vehicle drives out of the RSU coverage area during task execution, the result of will be communicated and transmitted through the MEC server. Since the amount of returned data is small, the download delay can be ignored. The transmission time of the subtask

[0070] (6)

[0071] Sub-task The execution time on is expressed as:

[0072] (7)

[0073] Wherein, is the computing resources selected and allocated to the sub-task .

[0074] If the vehicle drives out of the road section where the current RSU coverage area is located, the communication time between RSUs is:

[0075] (8)

[0076] Wherein, is the number of road sections passed by the vehicle, is the communication time between adjacent RSUs, which is defined as a constant.

[0077] Therefore, the total time for offloading and executing on the MEC server is:

[0078] (9)

[0079] The energy consumption generated by the vehicle for transmitting tasks is:

[0080] (10)

[0081] Step 2-1-3, processing of SV. If the sub-task generated by the vehicle is offloaded to , the time consists of two parts: transmission time and execution time. and The transmission time between is expressed as:

[0082] (11)

[0083] The execution time of the sub-task is expressed as:

[0084] (12)

[0085] Wherein, is the vehicle the computing resources allocated to the sub-task .

[0086] Therefore, the total time to unload to is: For:

[0087] (13)

[0088] The transmission energy consumption of is expressed as:

[0089] (14)

[0090] Meanwhile, The computing energy consumption generated by the auxiliary computing offloading is expressed as:

[0091] (15)

[0092] Step 2-1-4, offloading formula representation.

[0093] Step 2-1-4-1, latency formula. By analyzing the computing model, the total execution time of the subtasks generated by the vehicle is: For: For:

[0094] (16)

[0095] The task upload time of the subtask is: :

[0096] (17)

[0097] Step 2-1-4-2, energy consumption formula. The energy consumption of the subtask is divided into two parts, the energy consumption generated by the vehicle is: :

[0098] (18)

[0099] The execution energy consumption generated by the vehicle for auxiliary computing is: :

[0100] (19)

[0101] Step 2-2, establish the cost model. It includes the cost calculation of leasing the service program from the RSU and the cost calculation of leasing computing resources between vehicles. Specifically, it includes:

[0102] Step 2-2-1, Charge the RSU rental service program fee. For a vehicle, if it needs to process a task, it is required to cache the service program corresponding to the task, and the rental service program fee to be paid for each vehicle is:

[0103] (20)

[0104] where represents whether the vehicle has newly cached the service program , represents the unit price of renting the service program from the RSU, represents the cache size of the service program.

[0105] Step 2-2-2, Charge the computing resource fee. According to the above calculation model, when the vehicle decides to offload the subtask , it is required to pay the computing resource usage fee to the resource provider. For the vehicle that provides resources, its computing resource unit price is set to :

[0106] (21)

[0107] That is, the unit price is proportional to the computing resource that the vehicle is willing to allocate to the subtask , is the correlation coefficient.

[0108] For the vehicle with task requirements : If more computing resources are allocated to the subtask , the latency cost of can be reduced, but the service cost and energy consumption cost of will increase. Finally, for the subtask , the service fee to be paid to the resource provider can be expressed as :

[0109] (22)

[0110] Step 2-3, Construct and solve the problem of minimizing the average execution cost of the task initiator in the system, specifically including:

[0111] Step 2-3-1: Define the cost and optimization problem of the task initiator, and clarify the system optimization objectives and constraints.

[0112] The cost of the task initiator is defined as:

[0113] (23)

[0114] where , and represent the relative weights of latency, energy consumption, and service cost respectively, and , represents the completion time point of the exit sub-task of task i, i.e., the last sub-task. The total energy consumption generated by the vehicle (including generating tasks and providing services for other vehicles) is:

[0115] (24)

[0116] Finally, the optimization problem is expressed as:

[0117] (25)

[0118] where , , represent the task offloading symbol factors. In addition, the offloading decisions of all tasks are defined as , where . represents the computing resources allocated to all sub-tasks in the system, represents the computing resources allocated to each sub-task; represents the transmission power allocated to all sub-tasks in the system, represents the transmission power allocated to each sub-task.

[0119] Constraints:

[0120] (26)

[0121] (27)

[0122] (28)

[0123] (29)

[0124] (30)

[0125] (31)

[0126] According to formula (25), the objective function is the average execution cost of all tasks in the system. To reduce the average execution cost, it is necessary to minimize the sum of the execution costs of all tasks. Formula (26) represents the completion time constraint of tasks; formula (27) represents the subtask offloading constraint; formula (28) represents the offloading decision constraint; formula (29) represents the transmission power constraint; formulas (30) and (31) respectively represent that the computing resources allocated by the RSU and the vehicle for all subtasks do not exceed their total computing resources.

[0127] Step 2-3-2: Adopt a step-by-step solution strategy to handle the optimization problem, and integrate multi-agent reinforcement learning and particle swarm optimization to optimize task offloading and resource allocation, as follows:

[0128] In the first step, use multi-agent reinforcement learning to solve the subtask offloading decision. Determine the optimal offloading location of subtasks through agent interaction and learning, laying a foundation for resource allocation.

[0129] In the second step, use the particle swarm algorithm to solve the resource allocation decision. With the goal of minimizing the system cost, optimize the transmission power and computing resource allocation to improve resource utilization efficiency.

[0130] In the third step, update information according to the current execution results, including system state information, buffer information, and agent learning buffer data, so that the agent can update the neural network based on the latest information, continuously optimize the decision-making process, and improve system performance. This step-by-step solution strategy effectively addresses the problem complexity and improves the optimization effect of task offloading and resource allocation in the vehicle-assisted edge computing network.

[0131] Furthermore, in one embodiment, step 3 uses the multi-agent reinforcement learning method based on MADDPG to solve the offloading decision, and its implementation process is as follows: First, calculate the subtask scheduling order and completion time point according to the system state, and then determine the subtask offloading decision through the multi-agent deep reinforcement learning method based on MADDPG. Specifically include:

[0132] Step 3-1: Subtask scheduling order.

[0133] Step 3-1-1: According to the internal dependency relationship , obtain the set of predecessor subtasks and the set of successor subtasks of subtask

[0134] Step 3-1-2: Subtask execution delay.

[0135] Step 3-1-2-1, Local execution delay. Substitute into formula (4) to obtain the time for the subtask to be processed locally, where is the maximum computing frequency provided by the local vehicle .

[0136] Step 3-1-2-2, Execution delay on the RSU . Substitute into formula (7), substitute into formula (2), and call formula (6) to obtain the transmission and computing delays of the subtask on the RSU , where is the maximum computing resource provided by . Therefore, the total execution delay

[0137] :

[0138] Step 3-1-2-3, Calculate the execution delay on other vehicles. Substitute into formula (12), substitute into formula (3), and call formula (11) to obtain the transmission and computing delays of the subtask unloaded to the vehicle

[0139] . Therefore, the total execution delay unloaded to other vehicles is obtained from formula (13). , and obtained from formulas (4), (32), and (13) into the following formula:

[0140] (33)

[0141] to obtain the execution cost of the subtask, where represents the number of vehicles in the system.

[0142] Step 3-1-4, Execution priority. Based on the execution cost of the subtask calculated in Step 3-1-3, obtain its execution priority. In the tasks described by the DAG, define the last subtask as the exit subtask, and each subtask has at least one path to the exit subtask. Use the reverse recursive method to calculate the priority of the subtask:

[0143] (34)

[0144] where Denoted as a subtask The set of successor subtasks of is a subtask The priority of is a subtask The priority of

[0145] Step 3-1-5, construct a subtask scheduling queue. According to the subtask priorities obtained in Step 3-1-4, all subtasks are sorted to form a scheduling queue, denoted as .

[0146] Step 3-2, calculate the completion time point of subtasks.

[0147] Define Denote the task completion time point of subtask , Denote the task upload completion time point of subtask ;

[0148] The task completion time point of subtask can be expressed as:

[0149] (35)

[0150] where is the set of predecessor subtasks of subtask , when , .

[0151] The task upload completion time point of subtask is calculated as:

[0152] (36)

[0153] where when , .

[0154] Step 3-3, vehicle-assisted edge computing network task offloading optimization algorithm based on multi-agent reinforcement learning.

[0155] Formula (25) aims to optimize the execution cost of all tasks in the vehicle-assisted edge computing network. Since this problem involves resource competition and cooperation among multiple vehicles and a dynamic environment, traditional single-agent reinforcement learning algorithms cannot solve it. Therefore, a multi-agent reinforcement learning algorithm is proposed to solve the task offloading problem in the vehicle-assisted edge computing network. In this algorithm, the intelligent vehicle that generates tasks is the agent, and the vehicle-assisted edge computing network is the reinforcement learning environment. The agent maximizes its utility function by learning different strategies, thereby optimizing the system performance.

[0156] The following introduces the algorithm framework based on multi-agent reinforcement learning.

[0157] In reinforcement learning, the environment and the agent are the core elements. In each time slot, the agent makes an action decision based on the current environmental state. This action will cause the environment to change, and the environment then feedbacks a reward signal of the new state to the agent. There are three key elements in the reinforcement learning algorithm: state, action, and reward. The state represents the current situation of the system. The action is the decision made by the agent based on the observed state. The reward is the feedback signal from the environment, which is used to evaluate the agent's behavior. The agent maximizes the cumulative reward by continuously adjusting the action strategy. The following will elaborate on the applications of these three key elements in the vehicle-assisted edge computing network in combination with the system model and research questions.

[0158] Step 3-4-1, state space. To enable the agent to make appropriate actions based on environmental information, the state space of the th agent in time slot is defined as follows:

[0159] (37)

[0160] (1) Vehicle state information. Determine the vehicle set , and set the initial driving speed , maximum transmission power , maximum computing frequency , and maximum cache capacity for each vehicle .

[0161] (38)

[0162] (2) RSU state information. Specify the position coordinates , coverage range , and total computing resources of the equipped MEC server of the current RSU .

[0163] (39)

[0164] (3) Task and subtask status information. Define the task set . Among them represents the number of subtasks included in the task, represents the scheduling order of subtasks (calculated from step 3-1-5), represents the entire task 's task type ; represents the submission time point of the task ; represents the maximum tolerance time of the task .

[0165] (40)

[0166] (4) Service program status information, including:

[0167] Determine the status information of all vehicle caches for services within the current time slot, where represents whether this service is cached.

[0168] (41)

[0169] Determine the relevant information of the service program, such as unit price and capacity.

[0170] (42)

[0171] Then the service program status information is:

[0172] (43)

[0173] Step 3-4-2, action space. Mainly solve the offloading decision of subtasks. The th agent's action space in time slot is

[0174] (44)

[0175] In the formula, respectively represent that agent n in time slot t uses three processing methods for each subtask in the tasks it generates: local execution, offloading to the RSU for execution, and offloading to other vehicles for execution;

[0176] Step 3-4-3, reward function. For vehicle , if the current state satisfies the constraint conditions in formulas (26), (27), and (28), then the th agent's reward function in time slot is expressed as:

[0177] (45)

[0178] where is a negative parameter.

[0179] In a multi-agent system, the reward function comprehensively considers the behaviors of all agents, promotes system cooperation or coordination, balances the cooperation and competition among agents, and maximizes the overall performance. Usually, the reward of time slot is defined as the sum of the rewards of all agents. Extending formula (45) to a multi-agent environment, the reward value at the current state of the system can be expressed as:

[0180] (46)

[0181] Furthermore, in one embodiment, in step 4, based on the offloading decision determined in step 3, a particle swarm method based on PSO is used to solve the transmission power and computing resource allocation decisions. This process involves particle encoding, fitness function definition, and iterative update operations to determine the optimal solution. Specifically as follows:

[0182] Step 4-1, resource allocation.

[0183] The present invention uses a particle swarm optimization (PSO) algorithm to solve the transmission power resource allocation problem. The PSO algorithm simulates the foraging behavior of bird flocks in nature, abstracts the bird flock as a particle swarm, regards the food source as the optimal solution to be searched, and within a spatial range, particles share information, and each particle searches the space based on self-awareness and social experience to obtain the optimal solution.

[0184] Step 4-1-1, particle encoding. Let the number of particles be , the particle serial number , the number of tasks , the number of subtasks of task , and use matrices and and to represent the movement positions of particles, where matrix describes the transmission power resource allocation decision, and matrix describes the computing resource allocation decision. The element value of matrix represents the transmission power resource allocation decision of subtask , indicating the amount of transmission power allocation, and the element value of matrix describes the computing resource allocation decision of subtask , indicating the amount of computing resource allocation, with the unit of GHz.

[0185] In the resource allocation matrix and in, the particles all adopt real number coding. Use matrices and to represent the emission power allocation decision and the computing resource allocation decision of the particles respectively. The value of matrix represents the movement trend of allocating emission power for subtask , and the value of matrix represents the movement trend of allocating computing resources for subtask .

[0186] Step 4-1-2, fitness function. The total cost of the system:

[0187] (47)

[0188] Step 4-2, design the particle swarm algorithm, including:

[0189] In the particle swarm algorithm, the position of a particle represents a feasible solution. By calculating the next position of each particle's movement in each iteration until convergence to the optimal position. The update of the particle position is determined by the previous position and the velocity of the particle. The core of the algorithm is the iterative update method of the particle position and velocity. The velocity update is affected by three aspects: inertial velocity, self-cognitive experience, and social experience. Each aspect corresponds to a factor, namely the inertial factor and the learning factor , . By dynamically changing the values of each factor during the iteration process, the purpose of optimizing the particle swarm algorithm is achieved.

[0190] Step 4-2-1, parameter update.

[0191] If the velocity of the particle at the -th iteration is , its update formula is:

[0192] (48)

[0193] Among them, represents the updated iteration number, is a random function, is the optimal solution currently searched by the particle, is the optimal solution currently searched by all particles, is the inertial factor, is the -th iteration velocity; , are the learning factors. Let the inertial factor be dynamically updated, and its update formula is:​​

[0194] (49)

[0195] Among them, and are the current iteration number and the maximum iteration number respectively, is the number of particles, , are all coefficients, which can be adjusted according to 's optimal initial value. This update formula can make the inertia factor relatively large in the early stage of particle movement, giving the algorithm a strong global convergence ability. As the number of iterations increases in the middle and late stages of movement, decreases non-linearly, making the algorithm have a strong local convergence ability. At the same time, the number of particles affects : when the number of particles is large, appropriately reduce the value of to prevent the particle path from repeating; when the number of particles is small, appropriately increase the value of to increase the global convergence ability and prevent the particle path length from being insufficient, resulting in local convergence of the algorithm.

[0196] The learning factors , describe the influence degrees of self-cognitive experience and social experience on the particle velocity respectively. In this part, , are dynamically updated. According to the comparison of the particle fitness values in each iteration, the values of , in the next iteration of the particle are dynamically reduced or increased. Their update formulas are respectively:

[0197] (50)

[0198] (51)

[0199] In the formula, are respectively the learning factors in the th iteration and the th iteration, are respectively the learning factors in the th iteration and the th iteration, , and are respectively the fitness value of the particle at the th update, the fitness value of the current optimal solution of the particle, and the fitness value of the global optimal solution, represents the number of updates, , They are all coefficients used to adjust the increment ratio;

[0200] Step 4-2-2, position update.

[0201] The position update formula of the particle is:

[0202] (52)

[0203] Where, 、 and are the fitness values of the particle at the -th update, the fitness value of the current optimal solution of the particle, and the fitness value of the global optimal solution respectively. represents the number of updates. , They are all coefficients that can adjust the increment ratio. Formula (52) indicates that when the fitness value of the particle at the -th update is smaller than the fitness value of the individual optimal solution of the previous time, appropriately increase the at the -th update of the particle (i.e., ), and vice versa, appropriately decrease . This operation means that excellent particles will increase the influence of their own experience, while other particles rely on the social experience of the particle swarm. Similarly, when the fitness value of the individual optimal solution is less than the fitness value corresponding to the global optimal solution at the -th update of the particle, appropriately decrease the at the -th update of the particle (i.e., ), and vice versa, increase . This operation means that excellent particles will reduce the influence of social experience on the next velocity of the particle, trust their own cognitive experience, while the influence of social experience of other particles will also increase, making the particles approach the global optimal solution. To prevent the particles from being overconfident or overly dependent, set the variation ranges of 、 , that is, 、 、 and .

[0204] Step 4-2-3, the specific algorithm process is as follows:

[0205] Step 4-2-3-1, initialization, randomly generate a particle swarm of particles in the solution space, including the position matrix 、 and the velocity matrix 、 . To avoid the particle velocity calculated by the formula being too large or too small, for the velocity , set the minimum speed limit to , and the maximum speed limit to ; for the speed , set the minimum speed limit to , and the maximum speed limit to . When calculating the next position of the particle according to the speed formula, the position may exceed the boundary. To avoid this phenomenon, for the position , set the minimum position limit to , and the maximum position limit to ; for the position , set the minimum position limit to , and the maximum position limit to .

[0206] Step 4-2-3-2, calculate the fitness value of each particle according to formula (47).

[0207] Step 4-2-3-3, update the individual historical optimal position and the global optimal particle position .

[0208] Step 4-2-3-4, update the particle swarm position matrix and , so that the out-of-bounds element values in the transmission power resource allocation matrix are between and , and make the out-of-bounds element values in be between and .

[0209] Step 4-2-3-5, update the particle swarm velocity matrix and , so that the out-of-bounds element values in the velocity matrix are between and , and the out-of-bounds element values in

[0210] Step 4-2-3-6, let .

[0211] Step 4-2-3-7, if the fitness difference between two iterations is less than a fixed value, or the number of iterations reaches a fixed value, the algorithm ends and outputs the global optimal particle position , otherwise return to Step 4-2-3-2 to continue the iteration.

[0212] Further, in one of the embodiments, in step 5, after the decisions in steps 3 and 4 are completed, update the system state and Buffer buffer information, and train the multi-agent neural network, including the vehicle service program operation and the multi-agent reinforcement learning algorithm update process, to continuously optimize the system performance. Specifically as follows:

[0213] Step 5-1, update the vehicle service program;

[0214] Step 5-1-1, according to the subtask offloading situation (formula (44)), check the service type of the subtasks offloaded to the current vehicle during the offloading process;

[0215] Step 5-1-2, if the current vehicle already has this service program, there is no need to lease the current service from the RSU.

[0216] Step 5-1-3, if the current vehicle does not have this service program:

[0217] Step 5-1-3-1, determine whether the size of the currently pulled service program will exceed the cache capacity limit of the current vehicle (formula (1)). If not, directly cache it into the vehicle;

[0218] Step 5-1-3-2, if it exceeds the cache capacity of the current vehicle, perform FIFO replacement on the service programs in the vehicle cache space;

[0219] Step 5-2, the multi-agent reinforcement learning algorithm update process, specifically including:

[0220] Step 5-2-1, initialize the global policy network parameters , the target network parameters and the corresponding target network parameters and ;

[0221] Step 5-2-2, initialize the global experience replay memory buffer ;

[0222] Step 5-2-3, set the number of iterations , with the initial value of 0;

[0223] Step 5-2-4, initialize the observations of all agents ;

[0224] Step 5-2-5, initialize the time slot to 0;

[0225] Step 5-2-6, each agent generates an independent random number , if the random number is less than the greedy factor , with probability select to explore, that is, randomly select an action , otherwise, with probability select an action according to the policy ; ;

[0226] Step 5-2-7, each agent executes its respective action , and observes the current reward value and the next state ;

[0227] Step 5-2-8, upload the sample to the buffer ;

[0228] Step 5-2-9, update the current state to the next state: ;

[0229] Step 5-2-10, if the number of samples in reaches the experience pool size samples ;

[0230] Step 5-2-11, calculate the target value ;

[0231] Step 5-2-12, minimize the loss function and update the value network parameters ;

[0232] Step 5-2-14, update the policy network parameters by the deterministic policy gradient method ;

[0233] Step 5-2-15, use soft update to update the target network parameters to a smoothed version of the policy network parameters;

[0234] Step 5-3, start a new round of training until the preset end condition is reached.

[0235] In one embodiment, a task offloading and resource allocation optimization system for a vehicle-assisted edge computing network is provided. The system includes:

[0236] The first module is used to construct the basic model of the vehicle-assisted edge computing network, including the system architecture, task characteristics, service cache management, and communication capabilities;

[0237] The second module is used to construct a task cost evaluation and optimization model, quantify the cost, solve the optimization problem, and clarify the goal, task completion time, and offloading constraints;

[0238] The third module is used to calculate the sub-task scheduling order and completion time point according to the system state, and determine the sub-task offloading decision using the multi-agent deep reinforcement learning method based on MADDPG, which is implemented by defining relevant elements through the algorithm framework;

[0239] The fourth module is used to solve the transmission power and computing resource allocation decision using the particle swarm method based on PSO according to the offloading decision, and determine the optimal solution through particle encoding, defining the fitness function, and iterative update;

[0240] The fifth module is used to update the system state, buffer buffer information, and train the multi-agent neural network after completing the above decisions, including the vehicle service program and the multi-agent reinforcement learning algorithm update process, to continuously optimize the system performance.

[0241] For the specific limitations of the task offloading and resource allocation optimization system of the vehicle-assisted edge computing network, reference can be made to the limitations of the task offloading and resource allocation optimization method of the vehicle-assisted edge computing network in the above text, which will not be elaborated here. Each module in the above vehicle-assisted edge computing network task offloading and resource allocation optimization system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0242] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following is implemented:

[0243] Step 1, construct the basic model of the vehicle-assisted edge computing network, including the system architecture, task characteristics, service cache management, and communication capabilities;

[0244] Step 2, construct the task cost evaluation and optimization model, quantify the cost, and solve the optimization problem to clarify the goals, task completion time, and offloading constraints;

[0245] Step 3, calculate the sub-task scheduling order and completion time point according to the system state, and determine the sub-task offloading decision using the multi-agent deep reinforcement learning method based on MADDPG, which is implemented by defining relevant elements through the algorithm framework;

[0246] Step 4, solve the transmission power and computing resource allocation decision using the particle swarm method based on PSO according to the offloading decision, and determine the optimal solution through particle encoding, defining the fitness function, and iterative update;

[0247] Step 5, after completing the above decisions, update the system state, buffer buffer information, and train the multi-agent neural network, including the vehicle service program and the multi-agent reinforcement learning algorithm update process, to continuously optimize the system performance.

[0248] For the specific limitations of each step, reference can be made to the limitations of the task offloading and resource allocation optimization method for the vehicle-assisted edge computing network in the above text, which will not be elaborated here.

[0249] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes:

[0250] Step 1, construct the basic model of the vehicle-assisted edge computing network, including the system architecture, task characteristics, service cache management, and communication capabilities;

[0251] Step 2, construct a task cost evaluation and optimization model, quantify the cost, solve the optimization problem, and clarify the goals, task completion time, and offloading constraints;

[0252] Step 3, according to the system state, calculate the subtask scheduling order and completion time point, and use the multi-agent deep reinforcement learning method based on MADDPG to determine the subtask offloading decision, which is realized through the definition of relevant elements in the algorithm framework;

[0253] Step 4, according to the offloading decision, use the particle swarm method based on PSO to solve the transmission power and computing resource allocation decision, and determine the optimal solution through particle coding, definition of fitness function, and iterative update;

[0254] Step 5, after completing the above decisions, update the system state, buffer buffer information, and train the multi-agent neural network, including the vehicle service program and the multi-agent reinforcement learning algorithm update process, to continuously optimize the system performance.

[0255] For the specific limitations of each step, reference can be made to the limitations of the task offloading and resource allocation optimization method for the vehicle-assisted edge computing network in the above text, which will not be elaborated here.

[0256] The method of the present invention can solve problems such as limited resources of vehicle devices, heavy burden on MEC servers, and complex task offloading decisions, achieve the goals of reducing task execution latency, optimizing resource utilization, and reducing system costs, and improve the performance of the vehicle-assisted edge computing network.

[0257] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for optimizing task offloading and resource allocation in a vehicle-assisted edge computing network, characterized in that: The method comprises the following steps: Step 1: Build a basic model of the vehicle-assisted edge computing network, including system architecture, task characteristics, service cache management and communication capabilities; specifically: Step 1-1, building a vehicle-assisted edge computing network system model, specifically including: obtaining vehicle resource information in urban environment scenarios, dividing time slots and defining task vehicle sets, determining RSU resource information, and establishing a decision module on the cloud server to process vehicle and RSU uploaded status information and generate decisions; Step 1-2, establish a task model, including defining the task type and required cache capacity, and determining the RSU service program leasing rules; establish a vehicle task model, clarify the number, information and dependencies of subtasks; define three unloading methods: local execution of subtasks, unloading to RSU and other vehicles; Step 1-3, establishing a service cache model, including the vehicle leasing service programs based on subtask offloading information, and storage space constraints for service program cache; Step 1-4, establish a communication model, including vehicle-to-RSU communication and vehicle-to-vehicle communication; Step 2: Build a task cost evaluation and optimization model, quantify the cost and solve the optimization problem, clarify the goal and task completion time, and unloading constraints; specifically include: Step 2-1, establishing a calculation model; Step 2-2, establishing a cost model, including the cost calculation of leasing service programs from RSU and the cost calculation of leasing computing resources between vehicles; Step 2-3, construct and solve the problem of minimizing the average execution cost of the task initiator in the system; Step 3: Calculate the subtask scheduling order and completion time based on the system status, and use the multi-agent deep reinforcement learning method based on MADDPG to determine the subtask offloading decision, which is implemented by defining relevant elements through the algorithm framework; specifically, it includes: Step 3-1, calculate the subtask scheduling order; Step 3-2, calculate the completion time of the subtask; Step 3-3, designing a vehicle-assisted edge computing network task offloading optimization algorithm based on multi-agent reinforcement learning to optimize the task offloading problem of the vehicle-assisted edge computing network; in the algorithm, the smart car that generates the task is the agent, and the vehicle-assisted edge computing network is the reinforcement learning environment; Step 4: Based on the unloading decision, the PSO-based particle swarm method is used to solve the transmission power and computing resource allocation decision, and the optimal solution is determined through particle encoding, definition of fitness function and iterative update; Step 5: After completing the above decision, update the system status, buffer information and train the multi-agent neural network, including the vehicle service program and the multi-agent reinforcement learning algorithm update process to achieve continuous optimization of system performance.

2. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 1, characterized in that: Steps 1-4 specifically include: Step 1-4-1, establish a V2I communication model; Each vehicle in the system uses V2I communication to establish a communication connection with the RSU. The communication bandwidth between the vehicle and the RSU is , multiple subtasks are offloaded to the edge server nodes, i.e. fixed RSU nodes ; The i-th vehicle and fixed RSU nodes Uplink data rate between The calculation formula is: ; in, yes The upload transmission power, yes and The instantaneous channel gain between is the instantaneous distance between the two, and Respectively and The path loss and fading factor between represents the background noise power, , The total number of vehicles that generate tasks; Step 1-4-2, establish a V2V communication model; The i-th vehicle and the gth vehicle The bandwidth resources available are , and there are multiple subtasks offloaded to the vehicle , then the vehicle and vehicles The transmission rate between : ; in, yes and The instantaneous channel gain between is the instantaneous distance between the two, and Respectively and The path loss and fading factor between , is the total number of vehicles.

3. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 1, characterized in that: Step 1-1 specifically includes: Step 1-1-1, obtain vehicle resource information in the urban environment scene: Consider an urban environment scene with parallel lanes and vehicles running, the vehicle collection in the system , the i-th vehicle Expressed as ,in, , is the total number of vehicles, is the driving speed, is the maximum transmit power, is the maximum computation frequency, is the maximum cache capacity; Step 1-1-2, divide time slots: divide time into These equal time slots, each vehicle generates at most one task in a time slot; In, define the set represents the set of vehicles that generate tasks, , , Indicates Vehicles; Step 1-1-3, obtain RSU resource information in the urban environment scenario: RSUs are deployed on one side of the road, denoted as , Represents the Mth RSU deployed. Adjacent RSUs are connected by optical fiber. The road is divided into sections according to the coverage of RSUs. , is the Mth road section, where M is the total number of road sections; all RSUs are equipped with MEC servers to provide computing offloading services for vehicles; Step 1-1-4, establish a decision-making module: deploy the intelligent agent decision-making module on the cloud server to process the status information uploaded by the vehicle and RSU and make decisions.

4. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 3, characterized in that: Steps 1-2 specifically include: Step 1-2-1, Task type definition: define the task type set as , the i-th task type The required cache capacity is ; Assume that RSU has deployed service programs corresponding to all task types. After the vehicle makes a decision on subtask unloading, it updates the service program leasing according to the subtask unloading situation to assist itself and other vehicles in unloading subtasks. The unit price for leasing the service program from RSU is service fees; Step 1-2-2, build the mission model: vehicle The tasks generated in a time slot are represented as , here ,in, is the number of subtasks in the task; is the subtask information, For the task No. subtasks, is a matrix representing the subtask dependencies, Represents a subtask and There is a direct dependency between Represents a subtask and There is no direct dependency between them; is the task type, ; Submit the task at a certain time. is the maximum tolerable time for the task; Step 1-2-3, establish a subtask model: assume that subtasks cannot be divided any further, define tasks No. The subtask is ,in Indicates The input data size of each subtask; Indicates completion of The number of CPU cycles required for each subtask, ; Step 1-2-4, subtask offloading method: In the ITS scenario, the vehicle Generate Tasks After that, the subtasks are processed by three methods: local execution, unloading to fixed RSU nodes, and unloading to auxiliary vehicles that meet the conditions; for subtasks ,have offloading decisions, the symbolic factors of the three task offloading methods are , , ; Specifically include: (1) Unloading to local machine: subtask When executing locally ,otherwise ; (2) Offloading to RSU: Subtask When offloading to a fixed RSU node ,otherwise ; (3) Unloading to an auxiliary vehicle that meets the conditions: Subtask Unload to a qualified auxiliary vehicle hour ,otherwise ,in ; (4) Subtask offloading decision definition: The uninstall decision is defined as , the relationship between the task offloading mode symbol factor and the offloading decision is: if but ,like but ;like but ,show Decide on subtasks Uninstall to .

5. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 4, characterized in that: Steps 1-3 specifically include: Step 1-3-1, establish rental service conditions; After the subtask makes an unloading decision, the vehicle updates the service program leasing according to the unloading information; when the vehicle processes a subtask, it first checks whether there is a service program to process the subtask in its own cache space. If so, it processes it directly; if not, it leases it from the RSU; the vehicle cache space has an upper limit. If there is no extra space and a rental service is required, the FIFO strategy is used to replace the saved service program; Step 1-3-2, establish vehicle storage space restrictions; For the service program cache, there are the following storage capacity limitations: ; in, is a binary variable, indicating vehicle Is the cached capacity Service Procedure , Indicates that the service program is cached , otherwise it means it is not cached, and finally forms , indicating whether each vehicle has cached a certain service program; definition Binary variable, indicating whether it is a newly cached service program in the current time slot. Indicates the number of service program types. Indicates vehicle cache space.

6. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 5, characterized in that: In step 2-1, a calculation model is established, which specifically includes: Step 2-1-1, uninstall to local execution; If the vehicle Generated subtasks If you choose local execution, the execution time for: ; in, For vehicles Assign to subtask of computing resources, , ; Execution Energy Consumption for: ; in, is the effective switching capacitance in the chip; Step 2-1-2, the subtask is processed by the MEC server and offloaded to the fixed RSU node; If the vehicle Generated subtasks Offloaded to fixed RSU nodes Execute, if the vehicle When leaving the RSU coverage area during mission execution, The results will be transmitted through MEC server communication, subtask Transmission time It is expressed as: ; Subtask exist Execution time It is expressed as: ; in, for Select Assign to Subtask computing resources; If the vehicle If you leave the road section where the current RSU coverage is located, the communication time between RSUs for: ; in, is the number of road sections the vehicle passes through, is the communication time between adjacent RSUs, which is defined as a constant; Therefore, the total execution time of offloading to the MEC server is for: ; vehicle Energy consumption of transmission tasks for: ; Step 2-1-3, unloading to an auxiliary vehicle that meets the conditions for execution, that is, SV processing; If the vehicle Generated subtasks Uninstall to ,The time consists of two parts: transmission time and execution time; and The transmission time between It is expressed as: ; Subtasks Execution time It is expressed as: ; in, For vehicles Assign to subtask computing resources; Therefore, uninstall to Total time for: ; Transmission energy consumption It is expressed as: ; at the same time, Computational energy consumption generated by auxiliary computing offloading It is expressed as: ; Step 2-1-4, unloading the formulaic representation; (1) Time delay formulation; By analyzing the calculation model, the vehicle Generated subtasks Total execution time for: ; Subtasks Task upload time for: ; (2) Energy consumption formula; Subtasks The energy consumption of vehicles is divided into two parts: Energy consumption for: ; vehicle Execution energy consumption generated by auxiliary calculation for: ; In step 2-2, a cost model is established, including the cost calculation of leasing service programs from RSU and the cost calculation of leasing computing resources between vehicles, including: Step 2-2-1, establish a cost calculation model for the RSU leasing service program; For vehicles, if you want to process tasks, you need to cache the service program for the corresponding task, and each vehicle needs to pay the rental service program fee. for: ; in, Indicates the vehicles in the current time slot Whether the service program is newly cached , represents the unit price of the service program for leasing from RSU, Indicates the cache size of the service program; Step 2-2-2, establish a rental computing resource cost model; According to the above cost calculation model, when the vehicle Deciding to offload a subtask When using the resource provider, you need to pay the resource provider for the computing resource usage fee; for the vehicle that provides the resource , and its computing resource unit price is set as : ; Unit price and vehicle Willing to do subtask Allocated computing resources Proportional, is the correlation coefficient; For vehicles with mission requirements :like Allocate relatively more computing resources to subtasks , then for the subtask , Resource providers The service fee paid is expressed as : ; In step 2-3, we construct and solve the problem of minimizing the average execution cost of the task initiator in the system, which includes: Step 2-3-1, define the task initiator cost and optimization problem, and clarify the system optimization goals and constraints; Costs to the task initiator Defined as: ; in, , and represent the relative weights of latency, energy consumption, and service cost, respectively, and , represents the exit subtask of task i, that is, the completion time of the last subtask; Total energy consumption for: ; Finally, the optimization problem is expressed as: ; in, , , represents the task offloading symbol factor; is the offloading decision of all tasks, where ; Indicates the computing resources allocated to all subtasks in the system. Indicates the computing resources allocated to each subtask; Indicates the transmission power allocated to all subtasks in the system, Indicates the transmission power allocated to each subtask; Constraints: ; ; ; ; ; ; In the formula, , Represent the total computing resources of RSU and vehicle respectively; Step 2-3-2, using a step-by-step solution strategy to solve the optimization problem, integrating multi-agent reinforcement learning and particle swarm algorithm to optimize task offloading and resource allocation, specifically including: The first step is to use multi-agent reinforcement learning to solve the subtask offloading decision and determine the optimal offloading location of the subtask through agent interaction and learning; In the second step, the particle swarm algorithm is used to solve the resource allocation decision, with the goal of minimizing the system cost and optimizing the allocation of transmission power and computing resources; The third step is to update information based on the current execution results, including system status information, buffer information, and agent learning buffer data, so that the agent can update the neural network based on the latest information and continuously optimize the decision-making process.

7. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 6, characterized in that: Step 3 specifically includes: The subtask scheduling order is calculated in step 3-1, including: Step 3-1-1, according to Internal dependencies , get the subtask The set of predecessor subtasks and the set of subsequent subtasks ; Step 3-1-2, calculate the subtask execution delay, specifically including: Step 3-1-2-1, calculate the local execution delay: Substitute the execution time formula in step 2-1-1 to obtain the subtask The time for local processing, where For local vehicles The maximum calculation frequency provided; Step 3-1-2-2, calculate Execution delay: Substitute the execution time formula of step 2-1-2 into the formula, Substitute the formula in step 1-4-1 and call the transmission time formula in step 2-1-2 to obtain the subtask exist The transmission and computation delays of for The maximum computing resources provided, so the total execution delay is : ; Step 3-1-2-3, calculate the execution delay on other vehicles: Substitute the execution time formula of step 2-1-3 into the formula, Substitute the transmission rate formula in step 1-4-2 and call the transmission time formula in step 2-1-3 to obtain the subtask Unloading to vehicle The transmission and calculation delays on the , so the total execution delay of unloading to other vehicles is obtained by the total time formula in step 2-1-3; Step 3-1-3, calculate the subtask execution cost; The execution delays obtained by respectively combining the execution time formula in step 2-1-1, the total execution delay formula in step 3-1-2-2, and the total time formula in step 2-1-3 , and Substitute the following formula: ; Get subtask Execution cost ,in Indicates the number of vehicles in the system; Step 3-1-4, execution priority; Subtasks calculated according to step 3-1-3 Execution cost of , and obtain its execution priority; In the task described by DAG, the last subtask Defined as exit subtasks, each subtask has at least one path leading to the exit subtask, and the priority of the subtask is calculated using the reverse recursive method: ; in, Represented as a subtask The set of successor subtasks of For subtask The priority of For subtask Priority; Step 3-1-5, build a subtask scheduling queue; According to the subtask priority obtained in step 3-1-4, all subtasks are sorted to form a scheduling queue, which is expressed as ; The completion time of the subtask is calculated in step 3-2, including: definition Represents a subtask The task completion time, Represents a subtask The time point when the task upload is completed; Subtasks The task completion time is expressed as: ; in, It is a subtask The set of predecessor subtasks, when hour, ; Subtasks The time when the task upload is completed It is expressed as: ; Among them, when hour, .

8. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 7, characterized in that: The vehicle-assisted edge computing network task offloading optimization algorithm based on multi-agent reinforcement learning described in step 3-3 specifically includes: (1) State space; For Agents in time slot The state space The definition is as follows: ; in, is the vehicle status information, expressed as: ; In the formula, , , , are the i-th vehicle The initial driving speed, maximum transmission power, maximum calculation frequency and maximum cache capacity; is the RSU status information, expressed as: ; In the formula, for The location coordinates, For coverage, is the total computing resources of the configured MEC servers; is the task and subtask status information, expressed as: ; In the formula, Indicates the number of subtasks contained in the task. Indicates the scheduling order of subtasks, calculated by step 3-1-5. Represents the entire task The type of task; Representation Task The submission time point; Representation Task The maximum tolerance time; It is the service program status information, expressed as: ; in, Indicates the status information of all vehicle cache services in the current time slot. Indicates vehicle Whether the service program k is cached in Represents relevant information of the service program; (2) Action space; No. Agents in time slot The action space is: ; In the formula, They represent the three processing modes of each task generated by agent n in time slot t: local execution, offloading to RSU, and offloading to other vehicles; (3) Reward function; For vehicles , if the current state satisfies the first three formulas of the constraints in step 2-3-1, then Agents in time slot The reward function It is expressed as: ; in is a negative parameter; Expand the above formula to a multi-agent environment, and the reward value of the system in the current state It is expressed as: 。 9. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 8, characterized in that: Step 4 specifically includes: Step 4-1, allocate resources, including: Step 4-1-1, particle encoding: Assume the number of particles is , particle number , number of tasks ,Task The number of subtasks , using the matrix and Represents the moving position of the particle, where the matrix Describe the transmit power resource allocation decision, the matrix Describes computing resource allocation decisions; matrix The element value of Represents a subtask Transmit power resource allocation decision, indicating the amount of transmit power allocation, matrix The element value of Describe the subtask Computing resource allocation decision, which indicates the number of computing resources allocated, in GHz; In the resource allocation matrix and In the above example, particles are encoded with real numbers; and Respectively represent the emission power allocation decision and computing resource allocation decision of the particle; the matrix Value Represents a subtask The motion trend of allocating transmission power, the matrix Value Represents a subtask The movement trend of allocating computing resources; Step 4-1-2, construct the fitness function, which is the total cost of the system : ; Step 4-2, designing a particle swarm algorithm, including: Step 4-2-1, parameter update: If the particle The speed of the iteration is , and its update formula is: ; in, represents the number of update iterations, is a random function, is the optimal solution currently searched by the particle, is the optimal solution currently searched by all particles, is the inertia factor, For the The speed of iterations; , is the learning factor; The inertia factor Dynamic update, the update formula is: ; in, and are the current number of iterations and the maximum number of iterations respectively, is the number of particles, , are coefficients, The learning factor , The update formulas are: ; ; In the formula, Respectively Iteration, The learning factor of the iteration , Respectively Iteration, The learning factor of the iteration , , and The particles are The fitness value at the time of the first update, the fitness value of the particle's current optimal solution, and the fitness value of the global optimal solution, Indicates the number of updates. , All are coefficients, used to adjust the incremental ratio; Step 4-2-2, location update; The particle position update formula is: ; In the formula, Respectively Iteration, The position of the particle at the iteration; Step 4-2-3, the specific process of the algorithm includes: Step 4-2-3-1, initialization, random generation in the solution space A particle swarm of particles, including the position matrix , and the velocity matrix , ; For speed , set the minimum speed limit to , the maximum speed limit is ; For speed , set the minimum speed limit to , the maximum speed limit is ; For location , set the minimum position limit to , the maximum position limit is ; For location , set the minimum position limit to , the maximum position limit is ; Step 4-2-3-2, calculate the fitness value of each particle according to the fitness function of step 4-1-2; Step 4-2-3-3, update individual historical optimal position and the global optimal particle position ; Step 4-2-3-4, update the particle swarm position matrix and , so that the transmit power resource allocation matrix The out-of-bounds element value in is located at and between The out-of-bounds element value is located at and between; Step 4-2-3-5, update the particle swarm velocity matrix and , so that the velocity matrix The out-of-bounds element value is located at and between, The same applies to out-of-bounds element values; Step 4-2-3-6, let ; Step 4-2-3-7, if the fitness difference between two iterations is less than the preset value, or the number of iterations When the preset value is reached, the algorithm ends and outputs the global optimal particle position , otherwise return to step 4-2-3-2 to continue iterating.

10. The method for optimizing task offloading and resource allocation of a vehicle-assisted edge computing network according to claim 1, characterized in that: Step 5 specifically includes: Step 5-1, update the vehicle service program, including: Step 5-1-1, according to the subtask unloading situation, i.e., the action space, during the unloading process, check the subtask service type unloaded to the current vehicle; Step 5-1-2: If the current vehicle already has a service program corresponding to this subtask service type, there is no need to rent the current service from the RSU; if the current vehicle does not have a corresponding service program, execute step 5-1-3; Step 5-1-3 determines whether the size of the current service program pulled exceeds the cache capacity limit of the current vehicle. If not, it is directly cached in the vehicle; if it exceeds the cache capacity of the current vehicle, the service program in the vehicle cache space is replaced by FIFO; Step 5-2, update the multi-agent reinforcement learning algorithm, including: Step 5-2-1, initialize global policy network parameters , target network parameters And the corresponding target network parameters and ; Step 5-2-2, initialize the global experience playback memory buffer ; Step 5-2-3, set the number of iterations , the initial value is 0; Step 5-2-4, initialize the observations of all agents ; Step 5-2-5, initialize the time slot is 0; Step 5-2-6, each agent generates an independent random number , ; If the random number Less than the greed factor , with probability Choose to explore, i.e. randomly choose an action , otherwise, with probability According to the strategy Select Action ; Step 5-2-7, each agent performs its own actions , and observe the current reward value and the next state ; Step 5-2-8, the sample Upload to buffer ; Step 5-2-9, update the current state to the next state: ; Step 5-2-10, if When the number of samples in the experience pool reaches the preset size, a random sample is drawn from it. Samples ; Step 5-2-11, calculate the target value ; Step 5-2-12, minimize the loss function and update the value network parameters ; Step 5-2-14, update the policy network parameters by deterministic policy gradient method ; Step 5-2-15, using soft update to update the target network parameters to a smoothed version of the policy network parameters; Step 5-3, start a new round of training until the preset end condition is reached.

Citation Information

Patent Citations

  • Edge computing unloading and resource allocation method based on multi-agent reinforcement learning

    CN116321293A