Multi-queue task distributed unloading method and device in edge cloud Internet of Vehicles and medium
By dividing edge servers to edge domains and deploying agents in an edge cloud-to-vehicle network environment, and dynamically adjusting task offloading strategies using deep reinforcement learning models, the efficiency of large-scale distributed resources and task management is solved, and the task offloading effect with low latency and low energy consumption is achieved.
Patent Information
- Application Number
- CN202510130026.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-23
AI Technical Summary
With the sharp increase in the number of edge servers, how to efficiently manage large-scale distributed resources and tasks has become a key issue that needs to be solved urgently.
By dividing multiple edge servers into at least one edge domain and deploying an agent within each edge domain, a task offload strategy is dynamically determined using a dual-latency depth deterministic policy gradient model, including task offload ratio, bandwidth allocation ratio, and Liyapunov queue parameters, to minimize task delay ratio and energy consumption ratio.
It realizes efficient, low latency and low energy consumption task offloading, improving the efficiency and performance of distributed computing resources and task management in Internet of Vehicles and mobile edge computing environments.
Smart Images

Figure CN120034908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a method, device and storage medium for distributed unloading of multi-queue tasks in an edge cloud vehicle network. Background Art
[0002] The Internet of Vehicles technology connects vehicle terminals to the network, enabling communication between vehicle terminals and interaction between vehicle terminals and infrastructure. The development of this technology not only improves the efficiency of traffic management, but also provides strong support for intelligent transportation systems. The application scenarios of the Internet of Vehicles include autonomous driving, traffic flow optimization, and vehicle terminal safety warnings. These applications have increasingly higher requirements for computing resources and low latency.
[0003] MEC technology significantly reduces the delay of data transmission and improves the response speed of the system by offloading computing tasks from the cloud to edge servers. Edge servers are deployed close to vehicle terminals, such as roadside units (RSUs) or base stations, and can quickly process tasks generated by vehicle terminals, reducing the computing burden of the vehicle terminals themselves.
[0004] With the popularization of the Internet of Vehicles, the number of edge servers has increased dramatically, forming a large-scale distributed computing environment. MEC effectively solves the problems of insufficient computing resources for vehicle terminal devices and large cloud latency by offloading computing tasks to edge servers. However, with the rapid increase in the number of edge servers, how to efficiently manage large-scale distributed resources and tasks has become a key issue that needs to be solved urgently. Summary of the invention
[0005] The present invention provides a multi-queue task distributed unloading method in an edge cloud vehicle network, which is used to solve the problem in the prior art that it is difficult to efficiently manage large-scale distributed resources and tasks as the number of edge servers increases sharply.
[0006] The present invention provides a method for distributed unloading of multi-queue tasks in an edge cloud vehicle network, comprising the following steps: Divide a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; Obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; Based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; The next time slot observation state is used as the current time slot observation state, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute, and when the number of iterations reaches a preset number, all agents are aggregated for federated learning, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute until the model converges.
[0007] According to a method for distributed unloading of multi-queue tasks in an edge cloud vehicle network provided by the present invention, the method further includes: Based on the queue length of the edge server in the edge domain, a Lyapunov function is defined, and based on the Lyapunov function, a Lyapunov drift is defined; Based on the task delay, task energy consumption and task processing waiting delay in the edge domain, a Lyapunov penalty is defined; Taking minimizing the Lyapunov drift and the Lyapunov penalty in the current time slot as the optimization goal, the Lyapunov queue parameters of the edge domain in the current time slot are determined; the Lyapunov queue parameters include the control parameters between the Lyapunov drift and the Lyapunov penalty, the weight coefficient of the task delay, the weight coefficient of the task energy consumption and the weight coefficient of the task processing waiting delay.
[0008] According to a multi-queue task distributed unloading method in an edge cloud vehicle network provided by the present invention, the current time slot reward is obtained in the following manner: Based on the task delay ratio and task energy consumption ratio in the edge domain in the current time slot, the comprehensive utility of task offloading in the current time slot is determined; Determine the maximum value of the comprehensive utility under the preset constraints as the current time slot reward; The preset constraint condition includes at least one of the following: The offloading ratio of each task in the edge domain at the current time slot is between 0 and 1; The bandwidth allocation ratio of each edge server in the edge domain at the current time slot is between 0 and 1; The sum of bandwidth allocation ratios of all edge servers in the edge domain in the current time slot is 1; The task delay in the edge domain in the current time slot is not greater than the maximum task delay; The task energy consumption of the vehicle-mounted terminal in the edge domain in the current time slot is not greater than the maximum task energy consumption.
[0009] According to a task offloading method provided by the present invention, the task delay ratio in the edge domain in the current time slot is obtained by: Determine the local task processing delay in the edge domain at the current time slot based on the amount of tasks to be unloaded in the edge domain at the current time slot, the vehicle computing frequency, and the number of CPU cycles required for each unit task; Determine the cloud-edge collaborative task processing delay in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the vehicle computing frequency, the number of CPU cycles required for each unit task, the edge server computing frequency, the communication channel transmission rate, and the task offloading mode; Based on the local task processing delay in the edge domain and the cloud-edge collaborative task processing delay in the current time slot, determine the task delay ratio in the edge domain in the current time slot.
[0010] According to a multi-queue task distributed unloading method in an edge cloud vehicle network provided by the present invention, the task energy consumption ratio in the edge domain at the current time slot is obtained by the following method: Determine the energy consumption of edge server task processing in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, and the edge server task processing power consumption in the edge domain; Based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, the vehicle computing frequency, the task offloading mode and the edge server task processing power consumption in the edge domain, determine the cloud-edge collaborative task processing energy consumption in the edge domain at the current time slot; Based on the edge server task processing energy consumption and the cloud-edge collaborative task processing energy consumption in the edge domain in the current time slot, determine the task energy consumption ratio in the edge domain in the current time slot.
[0011] According to a method for distributed unloading of multi-queue tasks in an edge cloud vehicle network provided by the present invention, the method further includes: The gated recurrent unit layers are fused into the actor network and critic network of the dual-delayed deep deterministic policy gradient model respectively to obtain the target dual-delayed deep deterministic policy gradient model. Based on the current time slot observation state, the target double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent.
[0012] According to a multi-queue task distributed unloading method in an edge cloud vehicle network provided by the present invention, the federated learning aggregation of all intelligent agents includes: For each agent, determine the agent's reward change rate under the current federated learning aggregation round, and determine the agent's reward change weight based on the agent's reward change rate; Determine the aggregate weight of the agent according to the reward change rate and the initial weight of the agent, and obtain the global model parameter based on the local model parameter of each agent and the aggregate weight; For each agent, the model parameters of the agent are updated based on the cosine similarity between the local model parameters of the agent and the global model parameters.
[0013] According to a method for distributed unloading of multi-queue tasks in an edge cloud vehicle network provided by the present invention, the method further includes: For each agent, fit a linear regression model between the agent's reward and the update round, determine the slope of the linear regression model, and determine a confidence interval for the slope of the linear regression model; Adjust the aggregation frequency based on the slope and confidence interval of the linear regression model for each agent.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for distributed unloading of multi-queue tasks in an edge cloud vehicle network as described in any one of the above-mentioned methods is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the distributed unloading method of multiple queue tasks in the edge cloud vehicle network as described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for distributed unloading of multi-queue tasks in an edge cloud vehicle network.
[0017] The method for distributed offloading of multi-queue tasks in edge cloud vehicle networking provided by the present invention divides multiple edge servers into at least one edge domain, deploys an intelligent agent in each edge domain, and uses a double-delay deep deterministic policy gradient (TD3) model to dynamically determine the task offloading strategy. The intelligent agent determines the task offloading ratio, bandwidth allocation ratio and Lyapunov queue parameters based on the current time slot observation status, including the position of the vehicle terminal, the amount of tasks to be offloaded, the communication status, the task processing waiting delay, the remaining computing resources and location of the edge server, so as to minimize the task delay ratio and energy consumption ratio. The problem of heterogeneity of the distributed environment is solved by iteratively updating the model parameters of the intelligent agent and performing federated learning aggregation after reaching the preset number of iterations. Thereby, efficient, low-latency and low-energy task offloading is achieved, and the efficiency and performance of distributed computing resources and task management in the vehicle networking and mobile edge computing environments are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 It is a flow chart of a multi-queue task distributed unloading method in an edge cloud vehicle network provided by the present invention; Figure 2 It is a structural schematic diagram of a multi-queue task distributed unloading device in an edge cloud vehicle network provided by the present invention; Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Figure 1 : is a flow chart of a multi-queue task distributed unloading method in an edge cloud vehicle network provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 110, dividing a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; In this embodiment, a four-layer edge cloud Internet of Vehicles collaboration system is established, which includes a vehicle terminal layer, an edge layer, an edge domain layer, and a central cloud layer. The vehicle terminal layer includes vehicle terminals traveling on the road, and its own resources can be used to meet the service needs of the vehicle terminals. The edge layer includes edge servers deployed beside the road, and various service resources provided by service providers are cached in the edge servers to meet the service needs of vehicle terminals. The edge domain layer is an edge domain formed by several nearby edge servers. Each edge domain is deployed with a domain server. The domain server has an intelligent agent that can make task offloading and resource allocation, and provide offloading decisions for the edge servers in this domain. The central cloud layer mainly manages the entire system, allowing different edge domains to share information.
[0022] In one example, a vehicle terminal is driving on the road. In each area, there are N edge servers deployed at a distance that meets the preset distance requirements, forming an edge domain. There are M edge domains in the entire system. The central cloud layer controls all edge domains. In each edge domain, each vehicle terminal in the domain can connect to the edge server, and in each time slot t, each edge server can accept tasks that arrive within its communication coverage. According to the connection status of each vehicle terminal, the task arrival time, and the task size, each edge server can choose to offload part of the task or all of the tasks to reduce the computing burden of the vehicle terminal.
[0023] Step 120, obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; At the beginning of each time slot t, the agent in each edge domain obtains the observed state of the current time slot. This state information includes the location of the vehicle terminal in the edge domain, the amount of tasks to be unloaded (the amount of tasks generated by each vehicle terminal in the current time slot), the communication status (the communication status between each vehicle terminal and each edge server in the edge domain), the task processing waiting delay (the waiting time for processing the task of each vehicle terminal in the edge domain), the remaining computing resources of the edge server (the remaining computing resources of each edge server in the edge domain in the current time slot), and the edge server location (the location of each edge server in the edge domain).
[0024] In one example, the current time slot observation state at time slot t It can be expressed by the following formula: ; in, represents the remaining computing resources of each edge server in the edge domain, is the location information of each edge server in the edge domain, is the location information of the jth vehicle terminal in the edge domain, is the amount of tasks to be unloaded by the jth vehicle terminal in the edge domain, is the communication status between the jth vehicle terminal in the edge domain and each edge server in the domain, is the task processing waiting delay of the jth vehicle terminal in the edge domain.
[0025] It should be noted that at the beginning of each time slot, each vehicle terminal generates at most one task request, and the tasks generated in each time slot can be completed in the current time slot. In this embodiment, it is assumed that the arrival of tasks is completely random, follows Poisson distribution, and is stable.
[0026] Here, the task processing waiting delay is defined as the time interval from when the vehicle terminal sends a service request to when the request is processed. Specifically, the task processing waiting delay is the time interval between the time when the task starts to be processed and the time when the task arrives.
[0027] Step 130, based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Here, the Twin Delayed Deep Deterministic Policy Gradient (TD3) model is a deep reinforcement learning algorithm for continuous action space. TD3 uses two neural networks, including a main network and a target network. For each type of neural network, there is an actor network and two critic networks.
[0028] The model parameters of the actor network and the critic network in the main network are denoted as φ, θ1, θ2. The parameters of the actor network and the critic network in the target network are denoted as φ′, θ1′, θ2′. The actor network uses gradient ascent as the goal to maximize the cumulative expected return, and any critic network can be selected to calculate the Q value; while the critic network is updated by minimizing the error between the current Q value and the target Q value.
[0029] First, the agent in the edge domain inputs the current time slot observation state s(t) into the actor network of the main network, and then obtains the current time slot action from the output layer of the actor network , where the use of strategy Get the current time slot action ,Right now ; in, is the noise explored by the agent, which follows a normal distribution.
[0030] In this embodiment, these actions include the task offloading ratio (the ratio of tasks of each vehicle terminal offloaded to the edge server in the current time slot), the bandwidth allocation ratio of the edge server (the ratio of bandwidth allocated to each edge server in the current time slot), and the Lyapunov queue parameter (a parameter used to control the queue length of tasks offloaded to the edge server).
[0031] In one example, the current time slot action at the current time slot t It can be expressed by the following formula: ; ; in, represents the proportion of tasks of the jth vehicle terminal in the edge domain offloaded to the edge server at the current time slot t; Indicates the bandwidth ratio of the nth edge server in the edge domain in the current time slot; are all Lyapunov queue parameters.
[0032] Execute the current time slot action After that, the intelligent body will obtain the observation state of the next time slot , and based on the current time slot action , Current time slot observation status , the agent can obtain the current time slot observation state Execute the current time slot action Rewards .
[0033] In this embodiment, the reward It is determined based on the task delay ratio and task energy consumption ratio in the edge domain at the current time slot t.
[0034] Step 140, based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; Specifically, the four-tuple ( , , , ) is stored in the experience replay buffer. The model parameters of the critic network are updated using the method that minimizes the Q-value error. The model parameters of the actor network are updated using the deterministic policy gradient method.
[0035] Step 150, taking the next time slot observation state as the current time slot observation state, returning to continue iteratively executing the step of determining the current time slot action performed by the agent based on the current time slot observation state using a double-delay deep deterministic policy gradient model, and when the number of iterations reaches a preset number, performing federated learning aggregation on all agents, and returning to continue iteratively executing the step of determining the current time slot action performed by the agent based on the current time slot observation state using a double-delay deep deterministic policy gradient model until the model converges.
[0036] After each time slot ends, the observation state of the next time slot is used as the observation state of the current time slot, and steps 130 and 140 are continued to be performed for iterative training. When the number of iterations reaches the preset number, federated learning aggregation is performed, and the updated model parameters are distributed to each agent based on the federated learning aggregation, and then steps 130, 140 and 150 are continued until the model converges.
[0037] The method for distributed offloading of multi-queue tasks in edge cloud vehicle networking provided by the present invention divides multiple edge servers into at least one edge domain, deploys an agent in each edge domain, and uses a double-delay deep deterministic policy gradient (TD3) model to dynamically determine the task offloading strategy. The agent determines the task offloading ratio, bandwidth allocation ratio and Lyapunov queue parameters based on the current time slot observation status, including the position of the vehicle terminal, the amount of tasks to be offloaded, the communication status, the task processing waiting delay, the remaining computing resources and location of the edge server, so as to minimize the task delay and energy consumption. The problem of heterogeneity of the distributed environment is solved by iteratively updating the model parameters of the agent and performing federated learning aggregation after reaching the preset number of iterations. Thereby, efficient, low-latency and low-energy task offloading is achieved, and the efficiency and performance of distributed computing resources and task management in the vehicle networking and mobile edge computing environments are improved.
[0038] In some embodiments, the current time slot reward is obtained by: Based on the task delay ratio and task energy consumption ratio in the edge domain in the current time slot, the comprehensive utility of task offloading in the current time slot is determined; Determine the maximum value of the comprehensive utility under the preset constraints as the current time slot reward; The preset constraint condition includes at least one of the following: The offloading ratio of each task in the edge domain at the current time slot is between 0 and 1; The bandwidth allocation ratio of each edge server in the edge domain at the current time slot is between 0 and 1; The sum of bandwidth allocation ratios of all edge servers in the edge domain in the current time slot is 1; The task delay in the edge domain in the current time slot is not greater than the maximum task delay; The task energy consumption of the vehicle-mounted terminal in the edge domain in the current time slot is not greater than the maximum task energy consumption.
[0039] Specifically, the reward function in the edge domain under the current time slot is It can be defined as: ;in, The first I The comprehensive utility of the tasks, I The overall utility of each task is based on I The task delay ratio and task energy consumption ratio of each task are weighted and calculated.
[0040] The preset constraints are: ; ; ; ; ; in, Indicates the current time slot t I The task offloading ratio of each task; represents the bandwidth allocation ratio of the nth edge server in the current time slot t; Indicates the current time slot t I The task delay of each task; t is the maximum task delay, which is the length of a time slot; represents the task energy consumption of the vehicle terminal, is the maximum task energy consumption.
[0041] In one example, the task delay ratio in the lower edge domain of the current time slot is obtained by: Determine the local task processing delay in the edge domain at the current time slot based on the amount of tasks to be unloaded in the edge domain at the current time slot, the vehicle computing frequency, and the number of CPU cycles required for each unit task; Determine the cloud-edge collaborative task processing delay in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the vehicle computing frequency, the number of CPU cycles required for each unit task, the edge server computing frequency, the communication channel transmission rate, and the task offloading mode; Based on the local task processing delay in the edge domain and the cloud-edge collaborative task processing delay in the current time slot, determine the task delay ratio in the edge domain in the current time slot.
[0042] In this embodiment, the current time slot I Here, the task delay ratio of the current time slot is I The task latency ratio that can be saved after the cloud-edge-end collaborative processing of tasks is ;in, The current time slot I The local task processing delay of each task, The current time slot I The cloud-edge collaborative task processing latency of a task.
[0043] Specifically, the first I The local task processing delay of a task can be obtained by the following formula: ;in, The first I The amount of tasks to be offloaded, Calculate the frequency for the vehicle. is the number of bits per unit (bit) I The number of CPU cycles required for each task.
[0044] Specifically, the first I The cloud-edge collaborative task processing latency of a task can be obtained by the following formula: ; in, The first I The amount of tasks to be offloaded, Calculate the frequency for the vehicle. is the number of bits per unit (bit) I The number of CPU cycles required for each task, Calculate frequency for edge servers, is the communication channel transmission rate, This is the task offloading mode.
[0045] It should be understood that the calculation frequency in this embodiment is a specific quantitative indicator of the processor speed, that is, the number of clock cycles that can be performed per second, which will not be described in detail here.
[0046] In one example, the energy consumption ratio of tasks in the edge domain in the current time slot is obtained by: Determine the energy consumption of edge server task processing in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, and the edge server task processing power consumption in the edge domain; Based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, the vehicle computing frequency, the task offloading mode and the edge server task processing power consumption in the edge domain, determine the cloud-edge collaborative task processing energy consumption in the edge domain at the current time slot; Based on the edge server task processing energy consumption and the cloud-edge collaborative task processing energy consumption in the edge domain in the current time slot, determine the task energy consumption ratio in the edge domain in the current time slot.
[0047] In this embodiment, the current time slot I Here, the energy consumption of the task in the current time slot is I The energy consumption ratio of tasks that can be saved by collaborative processing of cloud-edge-end is ;in, The current time slot I The edge server task processing energy consumption of each task is The current time slot I The energy consumption of cloud-edge collaborative task processing for a task.
[0048] Specifically, the first I The energy consumption of edge server task processing for each task can be obtained by the following formula: ;in, The first I The amount of tasks to be offloaded, Calculate frequency for edge servers, is the number of bits per unit (bit) I The number of CPU cycles required for each task; is the communication channel transmission rate; Processing power consumption for edge server tasks.
[0049] Specifically, the first IThe energy consumption of cloud-edge collaborative task processing for a task can be obtained by the following formula: ; in, The first I The amount of tasks to be offloaded, Calculate the frequency for the vehicle. is the number of bits per unit (bit) I The number of CPU cycles required for each task, Calculate frequency for edge servers, is the communication channel transmission rate, It is task offloading mode; Processing power consumption for edge server tasks.
[0050] In this embodiment, the power consumption of edge server task processing is determined according to the edge server calculation frequency. In one example, , is the effective energy coefficient related to the chip architecture.
[0051] In addition, it should be noted that in the network transmission model, the signal will be affected by Gaussian white noise. Therefore, in this embodiment, a network communication model including Gaussian white noise is constructed to reflect the communication channel. Assuming a fixed transmission power and a standard path loss propagation exponent . Then the wireless link channel gain is expressed as: ;in, Represents the Euclidean distance between the vehicle terminal and the edge server.
[0052] Therefore, the communication channel transmission rate It can be expressed as: ; in, represents the power of Gaussian white noise, B represents the channel bandwidth, It is a signal resistance indicator.
[0053] It should be understood that in the vehicle network environment, due to the limited computing resources of the vehicle-mounted equipment, some tasks will be offloaded to the nearby edge server for processing, while the remaining tasks are executed locally. The edge server receives the offloaded tasks and adds them to the waiting queue to execute these computing-intensive tasks, and then sends feedback to the vehicle after completion. Therefore, the task offloading mode in this embodiment includes three modes: local computing, partial offloading, and complete offloading.
[0054] Specifically, in the entire system, the task offloading mode of each vehicle is as follows: ; Local offloading: Tasks that can be completed using the available computing resources of the vehicle terminal can be executed directly on the vehicle device. Therefore, in the local offloading scenario, the local execution delay of the vehicle terminal is expressed as: .
[0055] Partial offloading: Due to the limited computing power of the vehicle terminal, some computationally intensive tasks can be offloaded to the edge server for processing, while the rest of the tasks are calculated locally. Therefore, the partial offloading delay can be expressed as: ; in, represents the data transmission delay of the task partially offloaded to the edge server, represents the edge computing latency of tasks that are partially offloaded to edge servers, It represents the local unloading delay of the task unloaded to the vehicle terminal. For details, refer to the following formula: ; ; .
[0056] Full offloading: For computationally intensive tasks with strict latency requirements, the computing power of the vehicle terminal is insufficient to meet their needs. Therefore, these tasks are fully offloaded to the edge server for execution. The full offloading latency can be expressed as: ; ; in, represents the data transmission latency of the task completely offloaded to the edge server, Represents the edge computing latency of tasks that are completely offloaded to edge servers.
[0057] It should be understood that most tasks requested by vehicle terminals contain many subtasks with related relationships, and these subtasks cooperate with each other to complete a complex task. Therefore, if the related subtasks are unloaded to different locations, data interaction delay will occur between these related subtasks. Therefore, in this embodiment, the interaction delay of the related subtasks is also added to improve the task unloading delay. Assume represents the data communication volume between associated subtasks, represents the number of associated subtask pairs mined, Indicates the number of associated subtask pairs executed locally.
[0058] Therefore, for the total queue delay of the edge server, assuming that there are queues, each queue Total delay The sum of the delays of the associated tasks generated by the cumulative sum of the delays of all tasks in the queue: + ; in, I , indicating that it belongs to the queue All tasks I .
[0059] Therefore, the total delay of the system Indicates the delay of the queue with the largest total delay among all queues.
[0060] In some embodiments, the method further comprises: Based on the queue length of the edge server in the edge domain, a Lyapunov function is defined, and based on the Lyapunov function, a Lyapunov drift is defined; Based on the task delay, task energy consumption and task processing waiting delay in the edge domain, a Lyapunov penalty is defined; Taking minimizing the Lyapunov drift and the Lyapunov penalty in the current time slot as the optimization goal, the Lyapunov queue parameters of the edge domain in the current time slot are determined; the Lyapunov queue parameters include the control parameters between the Lyapunov drift and the Lyapunov penalty, the weight coefficient of the task delay, the weight coefficient of the task energy consumption and the weight coefficient of the task processing waiting delay.
[0061] When a task requires auxiliary computing of an edge server, the task needs to be placed on a suitable edge server. In this embodiment, Lyapunov-based task allocation is adopted to achieve optimal placement.
[0062] The Lyapunov function is used to measure the state of the system at time t, and is usually defined as a quadratic function of the queue length. Specifically, the Lyapunov function of the queue length of the edge server in the edge domain is defined as: ; in, Represents a queue of edge servers within an edge domain Length.
[0063] Substitute the Lyapunov function into the Lyapunov drift definition and expand it: ; Since only the length of the queue for adding tasks will change, in this embodiment, it is assumed that the queue index that changes is , then: ; At the current time slot : Before making a decision, a queue at an edge server There are already tasks. At the next time slot : When assigning a new task to a queue Afterwards, the queue There is tasks. Therefore, we can get , from which we can infer: ; because , so the upper bound of Lyapunov drift is: ; By minimizing the Lyapunov drift, the queue length can be controlled, thus ensuring the long-term stability of the system.
[0064] In addition, the Lyapunov penalty is defined as: ; Among them, E is the task energy consumption, D is the task delay, and W is the task processing waiting delay. is the weight coefficient of task delay, is the weight coefficient of task energy consumption and The weight coefficient for task processing waiting delay.
[0065] In order to ensure system stability, the optimization objective can be transformed into: ;in, is the control parameter between the Lyapunov drift upper bound and the Lyapunov penalty.
[0066] because There is an upper bound, so the above formula is equivalent to minimizing: .
[0067] In this embodiment, the Lyapunov queue parameters of the edge servers in the edge domain under this optimization target are obtained. After that, the edge servers in the edge domain are determined to have the Lyapunov queue parameters The Lyapunov drift and Lyapunov penalty are calculated, and the task is assigned to the edge server with the smallest sum of Lyapunov drift and Lyapunov penalty.
[0068] In this embodiment, not only is it ensured that each task is allocated and processed as quickly as possible, but also the load balancing between edge servers in the edge domain and the consumption of energy are taken into consideration, thereby avoiding the situation where a certain edge server is overloaded while other edge servers are idle.
[0069] In some embodiments, the method further comprises: The gated recurrent unit layers are fused into the actor network and critic network of the dual-delayed deep deterministic policy gradient model respectively to obtain the target dual-delayed deep deterministic policy gradient model. Based on the current time slot observation state, the target double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent.
[0070] In this embodiment, the TD3 model is improved. Specifically, the GRU (Gated Recurrent Unit) layer is introduced into the executor network and critic network of the TD3 model. As a variant of the recurrent neural network, GRU can capture and utilize long-term dependencies more effectively than the traditional fully connected layer, and GRU has fewer parameters, which is also in line with the environmental characteristics of edge computing.
[0071] In this embodiment, the GRU layer is integrated into the actor network and the critic network, which not only enhances the model's ability to process temporal information, but also uses GRU as an efficient feature extractor to extract more abstract and meaningful representations from input states and actions, thereby improving the performance and generalization ability of the model in complex environments.
[0072] It should be noted that although the intelligent agent based on deep reinforcement learning can dynamically and efficiently obtain the best data cache and calculate the offloading strategy, the centralized training is usually adopted, which leads to a slow convergence process of training. In addition, in the environment of edge computing, when facing a complex environment, it is easy to cause the state space and action space dimensions to be too high, resulting in extended training time and even making the network difficult to train.
[0073] Therefore, in this embodiment, in the face of such problems, a method combining federated learning aggregation and deep reinforcement learning is adopted, and the intelligent agent is trained in a distributed manner, thereby accelerating the convergence speed and increasing the performance of the model.
[0074] Specifically, in this embodiment, all agents are aggregated for federated learning, including: For each agent, determine the agent's reward change rate under the current federated learning aggregation round, and determine the agent's reward change weight based on the agent's reward change rate; Determine the aggregated weight of the agent according to the reward change rate and the initial weight of the agent, and obtain the global model parameters based on the local model parameters of each agent and the aggregated weight; For each agent, update the model parameters of the agent based on the cosine similarity between the local model parameters of the agent and the global model parameters.
[0075] In this embodiment, in the aggregation stage, an algorithm for dynamic weight aggregation is designed. For each agent, the weight of each model aggregation is determined according to the change amount of the reward. Before aggregation, calculate the reward change rate of each agent. If the reward change rate of the agent increases before aggregation, it means that it has learned more useful strategies during this time period. Therefore, increase its weight during aggregation to facilitate global convergence. Specifically, record the rewards obtained by the agents in each edge domain at the beginning and end of the aggregation round T, calculate the reward change rate, and then calculate the weight according to the change rate.
[0076] In one example, the reward change rate of agent m in the aggregation round T refers to the following formula: ; where, is the initial reward obtained by agent m at the beginning of the aggregation round T; is the end reward obtained by agent m at the end of the aggregation round T; represents the reward change rate of agent m in the aggregation round T. If is larger, it means that agent m has learned more useful strategies during this time period.
[0077] Next, according to the reward change rate of agent m in the aggregation round T, use the following formula to calculate the reward change weight : ; where, is usually used to represent a very small number to prevent the denominator from being zero, and m is the number of agents.
[0078] This embodiment includes m agents, and each agent sets an initial weight , indicating that each agent has the same initial weight during federated learning aggregation. In this embodiment, the initial weight and the reward change weight are fused to obtain the final aggregated weight .
[0079] It should be understood that each agent trains the model in its local environment to obtain local model parameters. In this embodiment, after obtaining the aggregate weight of each agent, a weighted sum is performed based on the local model parameters and aggregate weight of each agent to obtain a global model parameter.
[0080] In the update phase, considering that there are differences between various environments in edge computing environments, in traditional federated averaging, directly using the global model to update the local model will cause the characteristics of each node to be overwritten. Therefore, a local update method is adopted in this embodiment.
[0081] For each agent, we first compare the global model parameters with the local model parameters. We use the cosine similarity method to compare. Specifically, we first calculate the cosine distance , here, Represents cosine similarity. The larger it is, the greater the difference between the two models.
[0082] In a neural network, there are usually multiple layers of network structure. In this embodiment, the local parameters and global parameters of each layer are first flattened, and then the difference between the parameters of each layer is obtained by calculating their cosine distance. , this difference can be expressed as: For each layer of difference , taking the formula Make an appropriate adjustment and turn it into the weight when updating. Then the model parameter of each agent m when updating is ,in, are local model parameters, are global model parameters.
[0083] In one example, the aggregation frequency of the federated learning of the agent is also adaptively adjusted. Specifically, it also includes: For each agent, fit a linear regression model between the agent's reward and the update round, determine the slope of the linear regression model, and determine a confidence interval for the slope of the linear regression model; Adjust the aggregation frequency based on the slope and confidence interval of the linear regression model for each agent.
[0084] In this embodiment, for each agent, after participating in the federated learning aggregation, the reward of each update round when the agent locally iterates and updates the model parameters is recorded, and the linear regression model is fitted using the collected rewards of each update round, such as , For the l The rewards for each update round are: For the corresponding lUpdate rounds, is the intercept, is the slope.
[0085] After fitting the linear regression model, we will get an estimate of the slope. This estimate is the best guess at the true slope based on the currently available data. Therefore, the results of the linear regression analysis are used in this embodiment to calculate the standard error of the slope. The standard error reflects the extent to which the slope estimate may fluctuate around its true value and is a measure of the uncertainty of the slope estimate. The standard error of the slope is then multiplied by the corresponding critical value of the t distribution to calculate the error margin, and finally the confidence interval of the slope can be calculated based on the error margin. The width of the confidence interval reflects the uncertainty of the estimate; the greater the width, the higher the uncertainty.
[0086] It should be understood that analyzing the slope of the linear regression model can determine the trend of the agent's reward over time (i.e., update rounds). If the slope is positive and the confidence interval does not contain zero, it means that as the update rounds increase, the reward increases statistically significantly.
[0087] In this embodiment, the significance of the slope is determined by checking whether the confidence interval contains zero. Specifically, if the confidence interval does not contain zero, the representation slope is statistically significant, indicating that there is a real trend. If the confidence interval contains zero, the representation slope is not statistically significant, and there may be no actual trend. Then, according to the significance and direction of the slope, the aggregation frequency is adjusted. If the slope is significantly positive, the aggregation frequency is increased (i.e., the time interval between two aggregations is reduced), global synchronization is accelerated, model performance is improved, and the convergence of the global model is promoted. If the slope is significantly negative, the aggregation frequency is reduced (i.e., the time interval between two aggregations is increased), so that the model can be trained more locally, and more useful strategies can be learned locally before aggregation. If the slope is not significant, the current aggregation frequency is maintained without adjustment.
[0088] Furthermore, since there are multiple edge domains in the entire system, there are multiple agents participating in federated learning. Therefore, in this embodiment, when adjusting the aggregation frequency of the agents, two dimensional factors can also be considered: robustness (requiring that the aggregation frequency of at least 50% of the agents be adjusted before adjustment to prevent frequent changes in aggregation frequency due to abnormal trends of individual agents) and representativeness (ensuring that the adjustment decision is based on the adjustment trend of the aggregation frequency of the majority of agents to improve the reliability of the decision).
[0089] Specifically, the evaluation results of all agents are counted, and the number of agents that recommend adjusting the aggregation frequency (increase or decrease) is calculated. If the number of agents that recommend adjusting the aggregation frequency (increase or decrease) is greater than 50% of the number of all agents participating in federated learning, it is considered necessary to adjust the aggregation frequency, otherwise, no adjustment will be made.
[0090] For all agents that recommend adjustments, calculate the adjustment ratios (such as the percentage of increase or decrease) of the aggregation frequency they recommend, and find the average of these adjustment ratios. Finally, determine the aggregation frequency adjustment strategy of the agent based on the average. For example, if the average is a positive value, it means that most agents recommend increasing the aggregation frequency, so the aggregation frequency is increased accordingly. If the average is a negative value, it means that most agents recommend decreasing the aggregation frequency, so the aggregation frequency is decreased accordingly.
[0091] The following is a description of the multi-queue task distributed unloading device in the edge cloud vehicle network provided by the present invention. The multi-queue task distributed unloading device in the edge cloud vehicle network described below and the multi-queue task distributed unloading method in the edge cloud vehicle network described above can be referred to each other. Figure 2 In this embodiment, the multi-queue task distributed unloading device in the edge cloud vehicle network includes: The first task offloading module 210 is used to divide the plurality of edge servers into at least one edge domain; wherein the deployment distance between the edge servers in each edge domain meets the preset distance requirement, and an intelligent agent is deployed in each edge domain; The second task offloading module 220 is used to obtain the current time slot observation state of the intelligent agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be offloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; The third task offloading module 230 is used to determine the current time slot action performed by the agent based on the current time slot observation state using a double-delay deep deterministic policy gradient model, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay and task energy consumption in the edge domain at the current time slot; A fourth task offloading module 240 is used to update the model parameters of the agent in the edge domain based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain; The fifth task offloading module 250 is used to take the next time slot observation state as the current time slot observation state, return to continue iteratively executing the step of determining the current time slot action performed by the agent based on the current time slot observation state using a double-delay deep deterministic policy gradient model, and when the number of iterations reaches a preset number, perform federated learning aggregation on all agents, and return to continue iteratively executing the step of determining the current time slot action performed by the agent based on the current time slot observation state using a double-delay deep deterministic policy gradient model until the model converges.
[0092] The multi-queue task distributed unloading device in the edge cloud vehicle network provided by the present invention divides multiple edge servers into at least one edge domain, and deploys an intelligent agent in each edge domain, and uses the double-delay deep deterministic policy gradient (TD3) model to dynamically determine the task unloading strategy. The intelligent agent determines the task unloading ratio, bandwidth allocation ratio and Lyapunov queue parameters based on the current time slot observation status, including the position of the vehicle terminal, the amount of tasks to be unloaded, the communication status, the task processing waiting delay, the remaining computing resources and location of the edge server, so as to minimize the task delay and energy consumption. By iteratively updating the model parameters of the intelligent agent and performing federated learning aggregation after reaching the preset number of iterations, the problem of heterogeneity in the distributed environment is solved. Thereby, efficient, low-latency and low-energy task unloading is achieved, and the efficiency and performance of distributed computing resources and task management in the vehicle network and mobile edge computing environment are improved.
[0093] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communications interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the multi-queue task distributed unloading method in the edge cloud vehicle network, and the method includes: Divide a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; Obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; Based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; The next time slot observation state is used as the current time slot observation state, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute, and when the number of iterations reaches a preset number, all agents are aggregated for federated learning, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute until the model converges.
[0094] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0095] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the multi-queue task distributed unloading method in the edge cloud vehicle network provided by the above methods, the method comprising: Divide a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; Obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; Based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; The next time slot observation state is used as the current time slot observation state, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute, and when the number of iterations reaches a preset number, all agents are aggregated for federated learning, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute until the model converges.
[0096] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the multi-queue task distributed unloading method in the edge cloud vehicle network provided by the above methods, the method comprising: Divide a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; Obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; Based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; The next time slot observation state is used as the current time slot observation state, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute, and when the number of iterations reaches a preset number, all agents are aggregated for federated learning, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute until the model converges.
[0097] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0098] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed offloading method for multi-queue tasks in edge cloud vehicle networking, characterized in that: The method comprises: Divide a plurality of edge servers into at least one edge domain; wherein the deployment distance between edge servers in each edge domain meets a preset distance requirement, and an intelligent agent is deployed in each edge domain; Obtaining the current time slot observation state of the agent in the edge domain; the current time slot observation state includes the position of the vehicle terminal in the edge domain at the current time slot, the amount of tasks to be unloaded, the communication state between the vehicle terminal and the edge server, the task processing waiting delay, the remaining computing resources of the edge server and the position of the edge server; Based on the current time slot observation state, a double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent, and determine the current time slot reward for performing the current time slot action; the current time slot action includes the task offloading ratio in the edge domain at the current time slot, the bandwidth allocation ratio of the edge server and the Lyapunov queue parameter; the current time slot reward is determined based on the task delay ratio and the task energy consumption ratio in the edge domain at the current time slot; Based on the current time slot observation state, the current time slot action, the current time slot reward and the next time slot observation state of the agent in the edge domain, update the model parameters of the agent in the edge domain; The next time slot observation state is used as the current time slot observation state, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute, and when the number of iterations reaches a preset number, all agents are aggregated for federated learning, and the step of determining the current time slot action performed by the agent based on the current time slot observation state by using a double-delay deep deterministic policy gradient model is returned to iteratively execute until the model converges.
2. The distributed offloading method for multi-queue tasks in edge cloud vehicle networking according to claim 1 is characterized in that: The method further comprises: Based on the queue length of the edge server in the edge domain, a Lyapunov function is defined, and based on the Lyapunov function, a Lyapunov drift is defined; Based on the task delay, task energy consumption and task processing waiting delay in the edge domain, a Lyapunov penalty is defined; Taking minimizing the Lyapunov drift and the Lyapunov penalty in the current time slot as the optimization goal, the Lyapunov queue parameters of the edge domain in the current time slot are determined; the Lyapunov queue parameters include the control parameters between the Lyapunov drift and the Lyapunov penalty, the weight coefficient of the task delay, the weight coefficient of the task energy consumption and the weight coefficient of the task processing waiting delay.
3. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 1 is characterized in that: The current time slot reward is obtained in the following way: Based on the task delay ratio and task energy consumption ratio in the edge domain in the current time slot, the comprehensive utility of task offloading in the current time slot is determined; Determine the maximum value of the comprehensive utility under the preset constraints as the current time slot reward; The preset constraint condition includes at least one of the following: The offloading ratio of each task in the edge domain at the current time slot is between 0 and 1; The bandwidth allocation ratio of each edge server in the edge domain at the current time slot is between 0 and 1; The sum of bandwidth allocation ratios of all edge servers in the edge domain in the current time slot is 1; The task delay in the edge domain in the current time slot is not greater than the maximum task delay; The task energy consumption of the vehicle-mounted terminal in the edge domain in the current time slot is not greater than the maximum task energy consumption.
4. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 1 is characterized in that: The task delay ratio in the edge domain at the current time slot is obtained in the following manner: Determine the local task processing delay in the edge domain at the current time slot based on the amount of tasks to be unloaded in the edge domain at the current time slot, the vehicle computing frequency, and the number of CPU cycles required for each unit task; Determine the cloud-edge collaborative task processing delay in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the vehicle computing frequency, the number of CPU cycles required for each unit task, the edge server computing frequency, the communication channel transmission rate, and the task offloading mode; Based on the local task processing delay in the edge domain and the cloud-edge collaborative task processing delay in the current time slot, determine the task delay ratio in the edge domain in the current time slot.
5. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 1 is characterized in that: The energy consumption ratio of the tasks in the edge domain at the current time slot is obtained by: Determine the energy consumption of edge server task processing in the edge domain at the current time slot based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, and the edge server task processing power consumption in the edge domain; Based on the amount of tasks to be offloaded in the edge domain at the current time slot, the number of CPU cycles required for each unit task, the communication channel transmission rate, the edge server computing frequency, the vehicle computing frequency, the task offloading mode and the edge server task processing power consumption in the edge domain, determine the cloud-edge collaborative task processing energy consumption in the edge domain at the current time slot; Based on the edge server task processing energy consumption and the cloud-edge collaborative task processing energy consumption in the edge domain in the current time slot, determine the task energy consumption ratio in the edge domain in the current time slot.
6. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 1 is characterized in that: The method further comprises: The gated recurrent unit layers are fused into the actor network and critic network of the dual-delayed deep deterministic policy gradient model respectively to obtain the target dual-delayed deep deterministic policy gradient model. Based on the current time slot observation state, the target double-delay deep deterministic policy gradient model is used to determine the current time slot action performed by the agent.
7. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 1 is characterized in that: The federated learning aggregation of all agents includes: For each agent, determine the agent's reward change rate under the current federated learning aggregation round, and determine the agent's reward change weight based on the agent's reward change rate; Determine the aggregate weight of the agent according to the reward change rate and the initial weight of the agent, and obtain the global model parameter based on the local model parameter of each agent and the aggregate weight; For each agent, the model parameters of the agent are updated based on the cosine similarity between the local model parameters of the agent and the global model parameters.
8. The multi-queue task distributed unloading method in the edge cloud vehicle network according to claim 7 is characterized in that: The method further comprises: For each agent, fit a linear regression model between the agent's reward and the update round, determine the slope of the linear regression model, and determine a confidence interval for the slope of the linear regression model; Adjust the aggregation frequency based on the slope and confidence interval of the linear regression model for each agent.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer, it implements the multi-queue task distributed unloading method in the edge cloud vehicle network as described in any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the multi-queue task distributed unloading method in the edge cloud vehicle network as described in any one of claims 1 to 8.
Citation Information
Cited By
Multi-task edge computing method for Internet of Things
CN122064430A
A task offloading terminal decision, a server decision method and a task offloading system
CN122513835A
A task offloading terminal decision, a server decision method and a task offloading system
CN122513835B