Vehicle task unloading system and method based on deep Q network
By adopting a vehicle task offload system and method based on deep Q network in the Internet of Vehicles, the calculation delay and edge node resource competition problems of vehicle computing tasks are solved, and the task processing efficiency and resource utilization are improved, reducing calculation delay and communication costs.
Patent Information
- Application Number
- CN202510074061.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
AI Technical Summary
In the Internet of Vehicles environment, the calculation delay and edge node resource competition of vehicle computing tasks are serious, resulting in increased task processing delays and task failure.
The vehicle task unloading system and method based on the deep Q network are adopted to make task unloading decisions through the deep Q network model, dynamically adjust the task unloading strategy, optimize the task unloading decisions, and on the premise of ensuring user data privacy, the initial model is distributed training and aggregation optimization.
In complex environments, dynamically adjust task offloading strategies, improve task processing efficiency and system resource utilization, reduce calculation delays and communication costs, and realize efficient processing of complex tasks in vehicle networks.
Smart Images

Figure CN119938277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile edge computing technology for Internet of Vehicles, and in particular to a vehicle task offloading system and method based on a deep Q network. Background Art
[0002] With the development of intelligent driving and Internet of Vehicles technology, vehicles are no longer just means of transportation. They have gradually evolved into mobile intelligent nodes with powerful computing capabilities. The types and number of sensors installed on vehicles are increasing year by year, including lidar, radar, cameras, etc. The amount of data generated by these sensors is huge, especially computationally intensive tasks (such as real-time image processing, object recognition, etc.), which require a lot of computing resources to process. However, the computing resources of the vehicle itself are usually limited. Processing these computing tasks directly on the on-board equipment will increase computing delays, affecting the real-time performance and response speed of the system.
[0003] Task offloading has become a hot topic in recent years, especially task offloading based on edge computing. Vehicles can offload part of the computing tasks to nearby edge nodes (such as roadside units or service stations) for processing, thereby effectively reducing computing delays. However, as multiple vehicles offload tasks to edge nodes at the same time, the problem of resource competition becomes increasingly serious, which leads to increased task processing delays and even task failures. In this case, how to efficiently make task offloading decisions to balance the computing delay of tasks and the competition for edge node resources has become an urgent problem to be solved.
[0004] Therefore, in view of the privacy issues existing in cloud computing of visual data in the vehicle Internet of Things environment and the low training efficiency of traditional federated learning solutions, there is an urgent need for an efficient task offloading method to achieve it in complex environments. Summary of the invention
[0005] The purpose of the present invention is to provide a vehicle task offloading system and method based on a deep Q network, which can provide different initial models according to the characteristics of different road sections, and perform distributed training aggregation optimization on the initial model while ensuring the privacy of user data.
[0006] To achieve the above object, the vehicle task offloading method based on deep Q network provided by the present invention comprises the following steps: S1. Each client vehicle initializes the initial parameters and target network of the deep Q network model and receives the locally generated computing tasks; S2. The client vehicle obtains the basic parameters of the task according to the current hardware configuration and task requirements, including the task data volume, computing requirements and task start time; S3, the client vehicle uses the deep Q network model to make task offloading decisions based on environmental information and task parameters; S4. After the task is unloaded to the edge node, the edge node adds the task to the computing queue and performs concurrent task scheduling by dynamically adjusting the allocation of computing resources according to the computing resource allocation of existing tasks; S5. After the edge node completes the task calculation, it returns the calculation result to the client vehicle. The client vehicle updates the weight of the deep Q network model according to the task completion status, and trains and updates the global model.
[0007] Preferably, S3 includes: S31, the client vehicle reads the current task status and resource information of the edge node, including the computing requirements of the task and the computing resources of the edge node; S32, based on the current task status, computing requirements and resource competition of edge nodes, the delay and communication cost of local execution tasks and offloaded tasks are obtained through the deep Q network algorithm, and the offload decision Q value is generated; S33. Select the optimal offloading strategy according to the Q value. When the offloading benefit is higher than the local processing, choose to offload the task to the edge node.
[0008] Preferably, the delay of executing a task locally includes the computational delay when the task is processed locally, and is expressed as: ; In the formula, Indicates the task The computation delay of Indicates the task The amount of floating point operations required, Indicates the computing power of the vehicle.
[0009] Preferably, the delay of task offloading, including the communication delay and calculation delay of task offloading to the edge node, is as follows: The expression of communication delay is: ; In the formula, Indicates communication delay, Indicates the task The data size, represents the communication rate between the vehicle and the edge node; The expression for calculating the delay is: ; In the formula, Indicates the computing power of the edge node.
[0010] Preferably, in S4, computing resources are allocated as follows: ; In the formula, Indicates the task assigned of computing resources, Indicates the number of concurrent tasks, Indicates the task With existing tasks The concurrency overhead between .
[0011] Preferably, in S5, updating the deep Q network model includes updating the Q value by Bellman equation as follows: ; in, is the learning rate, is the discount factor, , Represents the current state and action respectively. represents the actual reward for task execution under the current decision, , Respectively represent the state and action of the next step.
[0012] Preferably, in S5, the training of the global model includes defining an optimization problem for task offloading decision based on the total reward of the task, as follows: ; ; In the formula, represents the reward function of the new task, represents the service delay of the new task, Indicates the remaining computation delay of the current task, and are all weight parameters.
[0013] The vehicle task offloading system based on deep Q network is characterized by comprising: The client is used to receive local tasks and determine the task offloading strategy based on task requirements; The region head detects vehicles entering or leaving the region and distributes the concurrent overhead model to edge nodes; The edge node performs concurrent task scheduling based on the computing resource allocation of the client offloaded tasks; The blockchain ledger records the task status and resource information of edge nodes.
[0014] A computer device comprises: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, it implements any one of the vehicle task unloading methods based on a deep Q network.
[0015] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, any one of the vehicle task offloading methods based on a deep Q network is implemented.
[0016] Therefore, the present invention adopts the above-mentioned vehicle task offloading system and method based on deep Q network, which has the following technical effects: (1) In complex environments, the deep Q network can be used to dynamically adjust the task offloading strategy according to the actual situation, gradually improve the task processing efficiency and the utilization of system resources, reduce computing delays and communication costs, and thus achieve efficient processing of complex tasks in vehicle networks.
[0017] (2) Through intelligent task offloading strategies, efficient management of vehicles with different computing requirements and task characteristics can be achieved, thereby maximizing the system's resource utilization, reducing the latency of vehicle computing tasks, and improving the task processing performance of the overall system.
[0018] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of a vehicle task offloading method based on federated learning; Figure 2 The vehicle task offloading system and method embodiments are different based on the deep Q network. The loss rate change diagram of vehicle task offloading training using DQN decision under different parameter values, Figure 2 (a) is The loss rate changes in different training rounds, Figure 2 (b) is The loss rate changes in different training rounds, Figure 2 (c) The change of loss rate in different training rounds; Figure 3 The vehicle task offloading system and method embodiments are different based on the deep Q network. Under the parameter value, the reward value change diagram of vehicle task offloading training using DQN decision, Figure 3 (a) is The reward value changes in different training rounds, Figure 3 (b) is The reward value changes in different training rounds, Figure 3 (c) The reward value changes in different training rounds. DETAILED DESCRIPTION
[0020] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.
[0021] Deep Q-Network (DQN) is an algorithm based on deep reinforcement learning that can make intelligent decisions in complex dynamic environments. Through the mechanism of experience replay and target network, DQN can dynamically adjust the offloading strategy and optimize the task offloading decision according to the current state of the vehicle, task requirements and edge node resources in the face of complex task offloading environments. In addition, DQN can gradually improve the efficiency of offloading decisions by continuously learning the historical offloading decision data of the vehicle, thereby obtaining the optimal offloading solution in an uncertain and complex network environment. In the Internet of Vehicles environment, the task offloading system based on the deep Q network can not only effectively reduce the delay of vehicle computing tasks, but also improve the resource utilization efficiency of edge nodes and reduce the task processing failure rate by learning and optimizing task offloading strategies.
[0022] like Figure 1 As shown, the present invention provides a vehicle task unloading method based on a deep Q network, comprising the following steps: S1. Each client vehicle initializes the initial parameters of the deep Q network (DQN) model and the target network. The client vehicle receives locally generated computing tasks, which may include tasks with high computing requirements such as image processing, path planning, sensor data analysis, etc.
[0023] S2. The client vehicle obtains the basic parameters of the task based on the current hardware configuration and task requirements, including the task data volume, computing requirements, and task start time. These parameters determine the communication and computing resources required for task processing.
[0024] S3: The client vehicle uses the deep Q network model to make task offloading decisions based on environmental information and task parameters. The specific steps are as follows: S31. The client vehicle reads the current task status and resource information of the edge node, including the computing requirements of the task and the computing resources of the edge node.
[0025] Each task Basic parameter definition, including the data size of the task , the floating point operations required for the task , the start time of the task , the computing power of the vehicle , computing power of edge nodes , and the communication rate between the vehicle and the edge node .
[0026] Task The latency when processing locally or on an edge node is as follows: When tasks are offloaded to edge nodes, communication delays The size of the task data can be and communication rate Said, the formula is: .
[0027] Computational latency when tasks are processed locally The computing power of the vehicle Related, the formula is: .
[0028] When tasks are offloaded to edge nodes, their computational latency Computing power of edge nodes Related, the formula is: .
[0029] S32. Based on the current task status, computing requirements, and resource competition of edge nodes, the client vehicle uses the DQN algorithm to evaluate the delay and communication cost of local execution and offloading tasks, and generates an offloading decision Q value.
[0030] S33. The system selects the optimal offloading strategy based on the Q value. If the offloading benefit is higher than the local processing, the task is offloaded to the edge node.
[0031] S4. When the task is unloaded to the edge node, the edge node adds the task to the computing queue and performs concurrent task scheduling based on the computing resource allocation of existing tasks. In concurrent task scheduling, the edge node will dynamically adjust the allocation of computing resources to maximize resource utilization and reduce task processing latency.
[0032] Since concurrent execution of multiple tasks on edge nodes will lead to competition for computing resources, in order to reasonably allocate computing resources for concurrent tasks, edge nodes need to consider the priorities of all current tasks and the computing requirements of tasks. , and available computing resources .
[0033] New Mission With existing tasks The concurrency overhead between , through which the resource competition between new tasks and existing tasks can be described, and factors such as the size of the task and the load of the edge nodes are taken into account. The computing resource allocation can be expressed as: .in, Assign to task of computing resources, is the number of concurrent tasks, To describe the task Resource competition with other concurrent tasks.
[0034] S5. After the edge node completes the task calculation, it returns the calculation result to the client vehicle. After receiving the task processing result, the client vehicle will update the weight of the DQN model according to the task completion status, making the next task offloading decision more intelligent. Through multiple rounds of task offloading and model updates, the deep Q network model of the client vehicle can gradually learn the optimal task offloading strategy, thereby maximizing the system's resource utilization efficiency in complex network and computing environments, reducing latency and improving the overall task processing speed.
[0035] The update process of the DQN model follows the following process: By comparing the actual task execution delay Delayed from expectation , calculate the actual reward , and update the Q value through the Bellman equation as follows: ; in, is the learning rate; is the discount factor, , are the current state and action respectively. is the actual reward for executing the task under the current decision, , They are the state and action of the next step respectively. Contains basic task parameters, network environment and edge node load information, action Including local processing Or offload to edge nodes .
[0036] To balance the processing delay of new tasks and the remaining computational delay of the currently executing task , the reward function is defined as follows: .in, Delayed service for new tasks, The remaining calculation delay for the current task, and are weight parameters, which are used to balance the delay of new tasks and the impact of existing tasks. The system optimizes the task offloading strategy by maximizing the total reward of all tasks. Therefore, the optimization problem of task offloading decision is defined as: By solving the optimization problem of task offloading decision, the system can determine the optimal task offloading strategy, that is, dynamically choose whether to process tasks locally or offload them to edge nodes for processing.
[0037] A system for a vehicle task offloading method based on a deep Q network, comprising: The client is used to receive local tasks and determine the task offloading strategy based on task requirements; The region head detects vehicles entering or leaving the region and distributes the concurrent overhead model to edge nodes; The edge node performs concurrent task scheduling based on the computing resource allocation of the client offloaded tasks; The blockchain ledger records the task status and resource information of edge nodes.
[0038] In this embodiment, different Parameter values (0.3, 0.5, and 0.7), the vehicle task offloading is trained by DQN decision, and the loss rate changes as follows Figure 2 As shown. The system loss rate gradually stabilizes under the value of = 0.3, the system has a higher initial loss rate but a faster convergence speed; = 0.7, the system shows better stability. Combined with the dynamic adjustment strategy of DQN, the system can optimize the model according to the real-time status and environmental conditions of task offloading, so that the loss rate in different offloading environments shows a gradual downward trend, indicating that our task offloading framework can better balance resource utilization and latency in different environments.
[0039] Based on different parameter values The changing trend of reward values of DQN decisions with (0.3, 0.5 and 0.7) during task offloading, such as Figure 3 As shown. It can be seen that = 0.7, the model showed the most significant increase in reward value, which indicates that when the task is unloaded, the higher The value helps the system to improve resource utilization efficiency and overall performance of task execution by dynamically adjusting the offloading strategy. =0.7 increases rapidly, indicating that in complex vehicle task unloading scenarios, the method of this embodiment can effectively learn and optimize the unloading strategy, thereby significantly improving the accuracy and timeliness of task unloading.
[0040] In addition, this implementation also compared the task completion of the pre-trained model in the real environment and the simulated environment. The performance of the two was very close, especially when the number of services was small, with the difference not exceeding 2.3%. This shows that the pre-trained model has good adaptability in the real environment.
[0041] In another embodiment, a computer device is provided, including: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, it implements any one of the vehicle task offloading methods based on the deep Q network.
[0042] In yet another embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the vehicle task offloading methods based on a deep Q network is implemented.
[0043] Therefore, the present invention adopts the above-mentioned vehicle task offloading system and method based on deep Q network, which can dynamically adjust the offloading strategy and optimize the task offloading decision according to the current state of the vehicle, task requirements and edge node resource conditions when facing a complex task offloading environment.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A vehicle task offloading method based on a deep Q network, characterized in that: The following steps are involved: S1. Each client vehicle initializes the initial parameters and target network of the deep Q network model and receives the locally generated computing tasks; S2. The client vehicle obtains the basic parameters of the task according to the current hardware configuration and task requirements, including the task data volume, computing requirements and task start time; S3, the client vehicle uses the deep Q network model to make task offloading decisions based on environmental information and task parameters; S4. After the task is unloaded to the edge node, the edge node adds the task to the computing queue and performs concurrent task scheduling by dynamically adjusting the allocation of computing resources according to the computing resource allocation of existing tasks; S5. After the edge node completes the task calculation, it returns the calculation result to the client vehicle. The client vehicle updates the weight of the deep Q network model according to the task completion status, and trains and updates the global model.
2. The vehicle task offloading method based on deep Q network according to claim 1, characterized in that S3 include: S31, the client vehicle reads the current task status and resource information of the edge node, including the computing requirements of the task and the computing resources of the edge node; S32, based on the current task status, computing requirements and resource competition of edge nodes, the delay and communication cost of local execution tasks and offloaded tasks are obtained through the deep Q network algorithm, and the offload decision Q value is generated; S33. Select the optimal offloading strategy according to the Q value. When the offloading benefit is higher than the local processing, choose to offload the task to the edge node.
3. The vehicle task offloading method based on deep Q network according to claim 2 is characterized in that: The latency of executing a task locally includes the computational latency of the task when it is processed locally. The expression is: ; In the formula, Indicates the task The computation delay of Indicates the task The amount of floating point operations required, Indicates the computing power of the vehicle.
4. The vehicle task offloading method based on deep Q network according to claim 2 is characterized in that: The delay of task offloading, including the communication delay and calculation delay from task offloading to edge nodes, is as follows: The expression of communication delay is: ; In the formula, Indicates communication delay, Indicates the task The data size, represents the communication rate between the vehicle and the edge node; The expression for calculating the delay is: ; In the formula, Indicates the computing power of the edge node.
5. The vehicle task offloading method based on deep Q network according to claim 1 is characterized in that: In S4, computing resources are allocated as follows: ; In the formula, Indicates the assignment to a task of computing resources, Indicates the number of concurrent tasks, Indicates the task With existing tasks The concurrency overhead between .
6. The vehicle task offloading method based on deep Q network according to claim 1 is characterized in that: In S5, the update of the deep Q network model includes updating the Q value through the Bellman equation as follows: ; in, is the learning rate, is the discount factor, , Represents the current state and action respectively. represents the actual reward for task execution under the current decision, , Respectively represent the state and action of the next step.
7. The vehicle task offloading method based on deep Q network according to claim 1 is characterized in that: In S5, the training of the global model involves defining the optimization problem of task offloading decision based on the total reward of the task, as follows: ; ; In the formula, represents the reward function of the new task, represents the service delay of the new task, Indicates the remaining computation delay of the current task, and are all weight parameters.
8. The vehicle task offloading system based on deep Q network is characterized by: include: The client is used to receive local tasks and determine the task offloading strategy based on task requirements; The region head detects vehicles entering or leaving the region and distributes the concurrent overhead model to edge nodes; The edge node performs concurrent task scheduling based on the computing resource allocation of the client offloaded tasks; The blockchain ledger records the task status and resource information of edge nodes.
9. A computer device comprising: Memory and processor; The memory stores a computer program, characterized in that when the processor executes the computer program, the vehicle task offloading method based on the deep Q network described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vehicle task offloading method based on a deep Q network described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-agent visual data unloading method based on information perception and reinforcement learning
CN119917183A
Multi-Agent Visualization Data Offloading Method Based on Information Sensing and Reinforcement Learning
CN119917183B
Multi-access point edge computing service state synchronization method and system based on deep Q network
CN121462596A
A multi-access point edge computing service state synchronization method and system based on a deep Q network
CN121462596B