Internet of vehicles task unloading and caching method based on time delay-energy consumption collaborative optimization

By constructing a latency-energy consumption collaborative optimization model and a meta-reinforcement learning algorithm, the problem of comprehensive optimization of energy consumption and latency in the Internet of Vehicles task offloading method is solved, enabling efficient decision-making and rapid adaptation of vehicles in different scenarios, and improving the performance of the Internet of Vehicles system.

CN120602999AActive Publication Date: 2025-09-05NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202511109389.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-05
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing Internet of Vehicles task offloading methods fail to effectively and comprehensively consider the vehicle's energy consumption and latency issues, and are difficult to adapt to the needs of intelligent connected vehicles. In addition, traditional algorithms have shortcomings in high-efficiency autonomous decision-making and dynamic performance optimization.

Method used

A vehicle network task offloading and caching method based on latency-energy consumption co-optimization is adopted. By constructing a latency-energy consumption co-optimization model for task units, and using the Markov decision model and meta-reinforcement learning algorithm (META-A3C) for online decision-making, task offloading and caching decisions are optimized.

Benefits of technology

It realizes dynamic optimization of vehicle energy consumption and service latency in different scenarios, has rapid learning and generalization capabilities, can quickly adapt to different latency-energy consumption optimization scenarios, and improve the efficiency of task offloading and caching decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602999A_ABST
    Figure CN120602999A_ABST
Patent Text Reader

Abstract

The invention discloses an Internet of Vehicles task unloading and caching method based on time delay-energy consumption collaborative optimization, and belongs to the technical field of Internet of Vehicles edge computing. Comprises: obtaining a task unit generated by a vehicle; according to the task unit and a pre-constructed communication model of the vehicle and the road side unit, calculating processing time delay and energy consumption when the task unit is executed on the vehicle, transmission time delay when the task unit is unloaded to the road side unit, and processing time delay and energy consumption when the task unit is executed on the road side unit; calculating an energy consumption cost function and a time cost function of the vehicle to the task unit; constructing a time delay energy consumption collaborative optimization model of the task unit; and converting the time delay energy consumption collaborative optimization model into a Markov decision model of the road side unit, and performing online decision making on the Markov decision model to obtain an unloading cache result of the task unit. According to the method, the two performance indexes of the vehicle energy consumption and the service time delay can be jointly optimized, and the service performance of the vehicle in different scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet of Vehicles edge computing technology, and in particular to an Internet of Vehicles task offloading and caching method based on latency-energy consumption collaborative optimization. Background Art

[0002] With the rapid development of connected vehicle (IoV) technology, there is a growing demand for improving the computational efficiency of IoV tasks and optimizing resource utilization. Traditional IoV systems typically leverage cloud computing technology, using remote cloud systems with ample computing power to provide vehicles with powerful computing and storage services. However, due to the significant physical distance between vehicles and cloud computing centers, data transmission is subject to significant transmission delays and losses, making it unable to meet the requirements of tasks such as emergency braking decisions and real-time communication. Furthermore, as vehicles continue to become more intelligent, the computing tasks generated by vehicles are placing an increasing demand on computing resources. Mobile Edge Computing (MEC) technology provides an effective solution to this problem. It effectively reduces the vehicle's computational load by offloading tasks to the collaborative processing of MEC servers on roadside units (RSUs).

[0003] However, current MEC-assisted vehicle task offloading approaches have several challenges. These solutions focus solely on the performance of a single vehicle task, such as service latency or energy consumption. With the rise of intelligent connected vehicles, comprehensive consideration of both energy consumption and latency has become crucial. Furthermore, traditional vehicle task offloading algorithms struggle to adapt to high-efficiency autonomous decision-making, continuous decision-making, and dynamic performance targets, requiring further optimization. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a vehicle network task offloading and caching method based on delay-energy consumption collaborative optimization, which can jointly optimize the two performance indicators of vehicle energy consumption and service delay, and improve the vehicle's business performance in different scenarios by optimizing task offloading decisions and task caching decisions.

[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0006] The present invention provides a method for offloading and caching tasks in an Internet of Vehicles (IoV) based on latency-energy consumption coordinated optimization, comprising:

[0007] Get the task units generated by the vehicle;

[0008] Based on the task unit and the pre-built communication model between the vehicle and the roadside unit, the processing delay and energy consumption of the task unit when executing on the vehicle, the transmission delay when the task unit is offloaded to the roadside unit, and the processing delay and energy consumption when the task unit is executed on the roadside unit are calculated;

[0009] Calculate the energy cost function and time cost function of the vehicle for the task unit based on the processing delay and energy consumption of the task unit when executed by the vehicle, the transmission delay of the task unit unloading to the roadside unit, and the processing delay and energy consumption of the task unit when executed by the roadside unit.

[0010] Based on the energy consumption cost function and time cost function of the vehicle to the task unit, a delay energy consumption collaborative optimization model of the task unit is constructed;

[0011] The delay-energy-consumption collaborative optimization model is converted into a Markov decision model of a roadside unit, and an online decision is performed on the Markov decision model to obtain an unloading cache result of the task unit.

[0012] Optionally, the pre-built vehicle-to-roadside unit communication model is expressed as:

[0013] ;

[0014] in, represents the pre-built downlink transmission rate between vehicle v and roadside unit r; Indicates the channel bandwidth; denote the channel gain and transmission power between vehicle v and roadside unit r at time t, respectively; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t.

[0015] Optionally, the calculation of the processing delay and energy consumption of the task unit when executed on the vehicle includes:

[0016] ;

[0017] ;

[0018] in, represents the processing delay of the i-th task unit when it is executed on vehicle v; Indicates the amount of computation required for the i-th task unit; represents the local computing capability of vehicle v; represents the energy consumption of the i-th task unit when executed on vehicle v; Coefficients representing the vehicle chip structure;

[0019] The calculation of the transmission delay of the task unit unloaded to the roadside unit and the processing delay and energy consumption when the task unit is executed on the roadside unit includes:

[0020] ;

[0021] ;

[0022] ;

[0023] in, represents the transmission delay of the i-th task unit from vehicle v to roadside unit r; Indicates the task data size of the i-th task unit; represents the pre-built downlink transmission rate between vehicle v and roadside unit r; represents the processing delay of the i-th task unit when it is executed on the roadside unit; Indicates the amount of data processed by the i-th task unit; represents the task cache decision when the i-th task unit is executed on the roadside unit r at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the transmission power between vehicle v and roadside unit r at time t.

[0024] Optionally, the calculation of the energy consumption cost function of the vehicle for the task unit includes:

[0025] ;

[0026] ;

[0027] ;

[0028] in, represents the task offloading decision of vehicle v at time t Total energy consumption under the conditions; Indicates the number of task units; represents the energy consumption of the i-th task unit when executed on vehicle v; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the energy consumption when all task units are executed on vehicle v; represents the energy consumption cost function of the vehicle for the task unit; Indicates the number of vehicles;

[0029] The calculation of the time cost function of the vehicle-to-task unit includes:

[0030] ;

[0031] ;

[0032] ;

[0033] in, Represents the task offloading decision of the i-th task unit at time t and the task cache decision at time t Total processing delay under the conditions; represents the processing delay of the i-th task unit when it is executed on the roadside unit; represents the processing delay when all task units are executed on vehicle v; represents the processing delay of the i-th task unit when it is executed on vehicle v; Represents the time cost function of the vehicle to task unit.

[0034] Optionally, the delay and energy consumption collaborative optimization model of the task unit is expressed as:

[0035] ;

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, Represents the optimization problem of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; represents the time cost function of the vehicle to the task unit; Represents constraints; Represent the first constraint, second constraint, third constraint, and fourth constraint respectively; Indicates the number of task units; Indicates the number of vehicles; represents the task offloading decision at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; Indicates the upper limit of the computing capacity of the roadside unit r; represents the task cache decision at time t; Indicates the amount of data processed by the i-th task unit; Indicates the upper limit of the storage capacity of the roadside unit r.

[0041] Optionally, the Markov decision model of the roadside unit is expressed as:

[0042] ;

[0043] ;

[0044] ;

[0045] in, Indicates the system status of the roadside unit at time t; Indicates the system action of the roadside unit at time t; represents the reward function of the roadside unit at time t; represents the channel gain between vehicle v and roadside unit r at time t; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t; They represent the task offloading decision at time t and the task caching decision at time t respectively; represents the transmission power decision of the vehicle at time t; Represents the cost function of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; Represents the time cost function of the vehicle to task unit.

[0046] Optionally, a META-A3C algorithm is used to perform online decision making on the Markov decision model to obtain an offloading cache result of the task unit, including:

[0047] Determine the delay energy consumption correlation coefficient at time t based on the reward function of the roadside unit , determine the optimized scenario type at time t according to the delay-energy consumption correlation coefficient at time t , generate decision tasks based on the optimized scenario type at time t ;in, Indicates the optimization scenario type for obtaining the delay-energy consumption correlation coefficient at time t; Indicates the nth optimization scenario type;

[0048] From the historical task dataset Take out K sample data and calculate the decision task based on K sample data The loss function , according to the decision task The loss function updates the policy network Parameters, value network Parameters of the updated strategy network , Update the value network ;in, 、 Represents the policy network Parameters, value network Parameters; Represent the updated strategy network Parameters, update value network Parameters; Indicates the system status of the roadside unit; Indicates the system action of the roadside unit; Represents network parameters;

[0049] According to the update strategy network , Update the value network , computational decision-making tasks Update loss function , according to the decision task Update loss function Calculating the meta-loss function ;

[0050] According to the meta-loss function , to update the strategy network Parameters, update value network The parameters of the global policy network are updated globally to obtain the global policy network , global value network ;in, Represent the global policy network Parameters, global value network Parameters;

[0051] According to the global policy network Parameters, global value network , the system status of the roadside unit at time t , execute the system action of the roadside unit at time t , obtain the system reward of the roadside unit at time t and the system status of the roadside unit at time t+1 ; Generate task data ( , , , ), executes the task data and stores the task data in the decision task Update the historical task dataset in the experience cache ; Among them, the system action of the roadside unit at time t Cache results for task unit offloading;

[0052] Repeat the process until all task units have their cached results.

[0053] Optionally, the decision-making task The loss function Expressed as:

[0054] ;

[0055] in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Representing the use value network Estimated system status Execute system actions The action value of Representing the use value network Estimated system status Execute system actions The action value of

[0056] Update policy network Parameters , Update the value network Parameters Respectively expressed as:

[0057] ;

[0058] ;

[0059] in, Represents the policy network Learning rate, value network The learning rate; Indicates about parameters Gradient operation.

[0060] Optional, decision-making tasks Update loss function Expressed as:

[0061] ;

[0062] in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the parameter update process;

[0063] The meta-loss function Expressed as:

[0064] ;

[0065] in, Represents a collection of optimized scene types.

[0066] Optional, global policy network Parameters , global value network Parameters Respectively expressed as:

[0067] ;

[0068] ;

[0069] in, represents the meta-learning rate; Indicates about parameters Gradient operation of ; represents the meta-loss function.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] The present invention realizes dynamic optimization of vehicle energy consumption and service delay in different scenarios by jointly controlling task offloading decisions and data caching decisions. It can achieve rapid learning ability for different optimization target tasks only through online learning of decision-making experience in different scenarios, and obtain rapid generalization ability for dynamic optimization target tasks only through a small sample training process, so as to achieve rapid adaptation to different delay-energy consumption optimization scenarios and enable rapid convergence of task offloading and caching decision problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 FIG2 is a flow chart of an embodiment of a method for offloading and caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to the present invention;

[0073] Figure 2 The figure shows the relationship between the algorithm convergence performance and the training cycle when the optimization parameters remain unchanged in one embodiment of the present invention;

[0074] Figure 3 The figure shows the relationship between the algorithm convergence performance and the training cycle under the optimization parameter mutation in one embodiment of the present invention;

[0075] Figure 4 Shown is a line graph of total energy consumption of tasks under different numbers of task units in one embodiment of the present invention;

[0076] Figure 5 Shown is a line graph of total task delay under different numbers of task units in one embodiment of the present invention. DETAILED DESCRIPTION

[0077] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0078] The term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " generally indicates an "or" relationship between the related objects.

[0079] Example 1

[0080] like Figure 1 As shown, this embodiment introduces a method for caching task offloading in an IoV system based on latency-energy consumption coordinated optimization. This method considers the dynamic latency and energy consumption requirements of IoV scenarios, the computing power of edge computing servers, and the storage capacity conditions, and addresses the joint design of task offloading and data caching. The method includes:

[0081] Step 1: Establish a communication model between the vehicle and the roadside unit (RSU), expressed as:

[0082] ;

[0083] in, represents the pre-built downlink transmission rate between vehicle v and roadside unit r; Indicates the channel bandwidth; denote the channel gain and transmission power between vehicle v and roadside unit r at time t, respectively; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t.

[0084] Step 2: Establish a calculation model for the task units generated by the vehicle, specifically:

[0085] The vehicle continuously generates task units, which can be completed by the vehicle itself or offloaded to the RSU for execution. After the task is processed, the result is transmitted to the vehicle via the wireless link of the communication model.

[0086] When the i-th task unit generated by vehicle v is calculated locally on the vehicle, the processing delay of the i-th task unit when it is executed on vehicle v is Expressed as:

[0087] ;

[0088] The energy consumption of the vehicle only includes the computing energy consumption of the i-th task unit in the vehicle, and the energy consumption of the i-th task unit when it is executed on the vehicle v. Expressed as:

[0089] ;

[0090] in, Indicates the amount of computation required for the i-th task unit, in cycles; Indicates the local computing power of vehicle v, that is, the CPU frequency, in cycles / s; Coefficient representing the vehicle chip structure.

[0091] When the i-th task unit generated by vehicle v is unloaded to the RSU for execution, the transmission delay of the i-th task unit from vehicle v to the roadside unit r is Expressed as:

[0092] ;

[0093] After the task is transmitted to the RSU, the RSU first queries the local cache pool to see if there is cached data of the processing result of the i-th task unit. If there is cached data of the task, the cached result is directly transmitted to the vehicle v. If not, the task is processed by the Mobile Edge Computing (MEC) server in the RSU. At this time, the processing delay of the i-th task unit when it is executed on the roadside unit is Expressed as:

[0094] ;

[0095] The energy consumption of the vehicle includes the transmission energy consumption of the i-th task unit and its processing results, the energy consumption of the i-th task unit when it is executed on the roadside unit r Expressed as:

[0096] ;

[0097] in, Indicates the task data size of the i-th task unit; Indicates the amount of data processed by the i-th task unit; represents the task cache decision when the i-th task unit is executed on the roadside unit r at time t, =1 indicates that the RSU caches the processing data of the i-th task unit, otherwise ; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t, in cycles / s; represents the transmission power between vehicle v and roadside unit r at time t.

[0098] Step 3: Calculate the energy consumption cost function and time cost function of the vehicle for the task unit, specifically:

[0099] The energy cost function of a vehicle is defined as the ratio of the energy consumed by the vehicle under a certain task unloading decision state to the energy consumed when there is no task unloading. The vehicle v makes the task unloading decision at time t. Total energy consumption under Expressed as:

[0100] ;

[0101] When all tasks of vehicle v are executed locally, that is, no tasks are offloaded to RSU, the energy consumption when all task units are executed on vehicle v is Expressed as:

[0102] ;

[0103] Then the energy consumption cost function of the vehicle for the task unit is Expressed as:

[0104] ;

[0105] in, Indicates the number of vehicles.

[0106] The time cost function of vehicle v for a task unit is defined as the ratio of the time required for the vehicle to obtain the task processing result to the time required for local processing of the task under certain task offloading decision and task caching decision conditions. The task offloading decision of the i-th task unit at time t is: and the task cache decision at time t Total processing delay under the conditions Expressed as:

[0107] ;

[0108] When all the task units of vehicle v are calculated by the local processor, the processing delay when all the task units are executed on vehicle v Expressed as:

[0109] ;

[0110] Then the time cost function of the vehicle to task unit is expressed as:

[0111] .

[0112] Step 4: Construct a collaborative optimization model for the latency and energy consumption of the task unit, specifically:

[0113] In the Internet of Vehicles system, vehicles continuously generate task units to be processed, and RSUs achieve the goal of alleviating vehicle load and reducing service latency by offloading the task calculation process and caching the task processing results. This process involves the optimization of task offloading decisions and caching decisions, that is, by controlling the task offloading decision at time t under the constraints of RSU computing power and storage capacity. and the task cache decision at time t , to achieve the joint optimization of the mission service delay and vehicle energy consumption at time t, that is:

[0114] Constructing the cost function of the delay-energy collaborative optimization model of task units , expressed as:

[0115] ;

[0116] in, and They are the energy consumption weight coefficient and the delay weight coefficient, satisfying ;

[0117] Construct the task offloading decision constraint condition. The total amount of tasks that can be offloaded by RSU at any time must meet its computing capacity constraint, that is:

[0118] ;

[0119] Construct the task cache decision constraint. The total amount of task data that the RSU can store at any time must meet its storage space limit, that is:

[0120] ;

[0121] By optimizing task offloading decisions and task caching decisions, the vehicle's energy consumption and service delay are reduced. The delay-energy collaborative optimization model of the task unit is expressed as:

[0122] ;

[0123] ;

[0124] ;

[0125] ;

[0126] ;

[0127] in, Represents the optimization problem of the delay-energy collaborative optimization model of the task unit; Represents constraints; Represent the first constraint, second constraint, third constraint, and fourth constraint respectively; Indicates the number of task units; represents the task offloading decision at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t, in cycles / s; Indicates the upper limit of the computing capacity of the roadside unit r; represents the task cache decision at time t; Indicates the upper limit of the storage capacity of the roadside unit r.

[0128] The above optimization problem P1 can be transformed into a Markov decision process (MDP) problem based on state-action values. Due to the dynamic optimization target caused by the dynamic parameter changes in this problem, the traditional reinforcement learning (RL)-based method has low training efficiency and lacks a clear training direction. Therefore, this embodiment designs a decision method based on meta-reinforcement learning to achieve online decision-making for the optimization problem P1.

[0129] Step 5: Convert the optimization problem of the delay-energy collaborative optimization model into a Markov decision model for the roadside unit, specifically:

[0130] The channel conditions, noise and interference environment of RSU are changing dynamically. Therefore, the system status of RSU is As time goes by, the system status of the roadside unit at time t is Expressed as:

[0131] ;

[0132] RSU needs to be based on the system status of the roadside unit at time t Make dynamic decisions, the system action of the roadside unit at time t Expressed as:

[0133] ;

[0134] The reward function is used to calculate RSU in the system state Execute system actions The system reward obtained, the MDP process usually takes maximizing the reward as the decision-making goal. The goal of the optimization problem is to simultaneously reduce the energy consumption of the vehicle and the service delay. Therefore, the cost function of the delay-energy collaborative optimization model of the task unit can be taken as a negative value to construct the reward function, the reward function of the roadside unit Expressed as:

[0135] ;

[0136] in, represents the transmission power decision of the vehicle.

[0137] Step 6: Make an online decision on the Markov decision model to obtain the offloading cache result of the task unit, specifically:

[0138] A Meta-Reinforcement Learning with Asynchronous Advantage Actor-Critic (META-A3C) algorithm is designed to perform online decision-making on multi-objective tasks. The delay-energy consumption correlation coefficient at time t is determined based on the reward function of the roadside unit. , according to the delay energy consumption correlation coefficient at time t Determine the optimization scenario type at time t ,Right now:

[0139] ;

[0140] Generate decision tasks based on the optimized scenario type at time t ;in, Indicates the optimization scenario type for obtaining the delay-energy consumption correlation coefficient at time t; Indicates the nth optimization scenario type.

[0141] From the historical task dataset Take out K sample data , Calculate the decision task based on K sample data The loss function , decision-making task The loss function Expressed as:

[0142] ;

[0143] in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Representing the use value network Estimated system status Execute system actions The action value of Representing the use value network Estimated system status Execute system actions The action value of 、 Represents the policy network Parameters, value network Parameters; They represent the system state of the first sample at time t, the system state of the second sample at time t, ..., the system state of the K-th sample at time t, respectively; They represent the system action of the first sample at time t, the system action of the second sample at time t, ..., the system action of the Kth sample at time t, respectively; They represent the system reward for the first sample at time t, the system reward for the second sample at time t, ..., the system reward for the K-th sample at time t, respectively; They represent the system state of the 1st sample at time t+1, the system state of the 2nd sample at time t+1, ..., the system state of the Kth sample at time t+1, respectively.

[0144] In this embodiment, the global policy network , global value network It is updated in real time. During the current network update round, the policy network Parameters , Value Network Parameters The parameters of the global policy network and the global value network obtained in the previous network update round.

[0145] According to the decision task The loss function updates the policy network Parameters, value network Parameters of the updated strategy network , Update the value network , Indicates the system status of the roadside unit; Represents the system action of the roadside unit; updates the strategy network Parameters , Update the value network Parameters Respectively expressed as:

[0146] ;

[0147] ;

[0148] in, Represents the policy network Learning rate, value network The learning rate; Indicates about parameters Gradient operation.

[0149] According to the update strategy network , Update the value network , computational decision-making tasks Update loss function , decision-making task Update loss function Expressed as:

[0150] ;

[0151] in, Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the use of updated value network Estimated system status Execute system actions The action value of Indicates the parameter update process.

[0152] According to the decision task Update loss function Calculating the meta-loss function , meta-loss function Expressed as:

[0153] ;

[0154] in, Represents a collection of optimized scene types.

[0155] Meta-loss function Perform gradient descent, according to the meta-loss function , to update the strategy network Parameters, update value network The parameters of the global policy network are updated globally to obtain the global policy network , global value network , global policy network Parameters , global value network Parameters Respectively expressed as:

[0156] ;

[0157] ;

[0158] in, represents the meta-learning rate.

[0159] According to the global policy network Parameters, global value network , the system status of the roadside unit at time t , execute the system action of the roadside unit at time t , obtain the system reward of the roadside unit at time t and the system status of the roadside unit at time t+1 ; Generate task data ( , , , ), executes the task data and stores the task data in the decision task Update the historical task dataset in the experience cache ; Among them, the system action of the roadside unit at time t Cache results for offloading of task units.

[0160] This process is continuously executed, during which task offloading and caching decisions are updated and optimized until the offloading caching results of all task units are obtained.

[0161] Example 2

[0162] This embodiment introduces an experimental example of a method for offloading and caching tasks in an Internet of Vehicles (IoV) based on latency-energy consumption coordinated optimization:

[0163] Figure 2This graph shows the relationship between algorithm convergence performance and training cycle under fixed optimization parameters. It compares the convergence performance of the Meta-Reinforcement Learning with Asynchronous Advantage Actor-Critic (META-A3C) algorithm in this embodiment with other algorithms, such as the Asynchronous Advantage Actor-Critic (A3C) algorithm and the Deep Deterministic Policy Gradient (DDPG) algorithm, under fixed optimization parameters. The horizontal axis represents training rounds, and the vertical axis represents cumulative reward. The results show that the Meta-A3C algorithm reaches stable convergence within 100 rounds, while the A3C algorithm and the DDPG algorithm require 150 and 200 rounds, respectively. The Meta-A3C algorithm's rapid convergence is due to meta-learning's pre-training optimization of global parameters, which enables efficient utilization of historical experience and reduces exploration costs. Furthermore, the Meta-A3C algorithm's final cumulative reward is significantly higher than that of other algorithms, validating the effectiveness of its joint optimization of energy consumption and latency.

[0164] Figure 3 This is a graph showing the relationship between the algorithm convergence performance and the training cycle in a dynamic scenario (optimization parameter changes) of this embodiment, showing the convergence performance of each algorithm when the "delay-energy consumption" correlation parameter is dynamically adjusted. When training to 250 rounds, the energy consumption weight coefficient Increase the delay weight coefficient from 0.3 to 0.7 The Meta-A3C algorithm reconverged after only a few dozen rounds, while the A3C and DDPG algorithms required over 100 additional rounds. This result demonstrates that the Meta-A3C algorithm's meta-learning mechanism can quickly adapt to changes in the "latency-energy" optimization ratio. By dynamically adjusting the policy network parameters, it can quickly respond to new optimization directions, demonstrating the algorithm's robustness in dynamic environments.

[0165] Figure 4 The total energy consumption curve of tasks under different numbers of task units in this embodiment analyzes the impact of the number of task units on total energy consumption. = =0.5, the figure shows the changing trend of the total energy consumption of each algorithm task as the number of task units increases. Figure 4 As can be seen from the table, the total task energy consumption of the algorithm in this embodiment is consistently lower than that of other algorithms, and its energy efficiency advantage becomes increasingly apparent as the number of task units increases. This demonstrates that when processing a large number of tasks, the method in this embodiment can better optimize task offloading and storage decisions, balance energy consumption and latency, and achieve efficient resource utilization.

[0166] Figure 5 The total task delay curve for this embodiment under different numbers of task units shows the impact of increasing the number of task units on the total delay. = = 0.5, this figure shows the changes in the total task delay of each algorithm when the number of task units is different. Figure 5 It can be seen that the total task latency of the method in this embodiment is at a low level under each number of task units, and as the number of task units increases, its latency advantage becomes more and more obvious, further proving that the algorithm can effectively reduce latency when processing tasks of different scales, ensure the rapid processing of Internet of Vehicles tasks, and improve user experience.

[0167] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0168] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0169] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0171] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which are all protected by the present invention.

Claims

1. A method for caching and offloading tasks in an Internet of Vehicles (IoV) based on coordinated optimization of latency and energy consumption, characterized in that: include: Get the task units generated by the vehicle; Based on the task unit and the pre-built communication model between the vehicle and the roadside unit, the processing delay and energy consumption of the task unit when executing on the vehicle, the transmission delay when the task unit is offloaded to the roadside unit, and the processing delay and energy consumption when the task unit is executed on the roadside unit are calculated; Calculate the energy cost function and time cost function of the vehicle for the task unit based on the processing delay and energy consumption of the task unit when executed by the vehicle, the transmission delay of the task unit unloading to the roadside unit, and the processing delay and energy consumption of the task unit when executed by the roadside unit. Based on the energy consumption cost function and time cost function of the vehicle to the task unit, a delay energy consumption collaborative optimization model of the task unit is constructed; The delay-energy-consumption collaborative optimization model is converted into a Markov decision model of a roadside unit, and an online decision is performed on the Markov decision model to obtain an unloading cache result of the task unit.

2. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1 is characterized in that: The pre-built communication model between the vehicle and the roadside unit is expressed as: ; in, represents the pre-built downlink transmission rate between vehicle v and roadside unit r; Indicates the channel bandwidth; denote the channel gain and transmission power between vehicle v and roadside unit r at time t, respectively; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t.

3. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1, characterized in that: The calculation of the processing delay and energy consumption of the task unit when executed on the vehicle includes: ; ; in, represents the processing delay of the i-th task unit when it is executed on vehicle v; Indicates the amount of computation required for the i-th task unit; represents the local computing capability of vehicle v; represents the energy consumption of the i-th task unit when executed on vehicle v; Coefficients representing the vehicle chip structure; The calculation of the transmission delay of the task unit unloaded to the roadside unit and the processing delay and energy consumption when the task unit is executed on the roadside unit includes: ; ; ; in, represents the transmission delay of the i-th task unit from vehicle v to roadside unit r; Indicates the task data size of the i-th task unit; represents the pre-built downlink transmission rate between vehicle v and roadside unit r; represents the processing delay of the i-th task unit when it is executed on the roadside unit; Indicates the amount of data processed by the i-th task unit; represents the task cache decision when the i-th task unit is executed on the roadside unit r at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the transmission power between vehicle v and roadside unit r at time t.

4. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1, characterized in that: The calculation of the energy consumption cost function of the vehicle to task unit includes: ; ; ; in, represents the task offloading decision of vehicle v at time t Total energy consumption under the conditions; Indicates the number of task units; represents the energy consumption of the i-th task unit when executed on vehicle v; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the energy consumption when all task units are executed on vehicle v; represents the energy consumption cost function of the vehicle for the task unit; Indicates the number of vehicles; The calculation of the time cost function of the vehicle-to-task unit includes: ; ; ; in, Represents the task offloading decision of the i-th task unit at time t and the task cache decision at time t Total processing delay under the conditions; represents the processing delay of the i-th task unit when it is executed on the roadside unit; represents the processing delay when all task units are executed on vehicle v; represents the processing delay of the i-th task unit when it is executed on vehicle v; Represents the time cost function of the vehicle to task unit.

5. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1, characterized in that: The delay and energy consumption collaborative optimization model of the task unit is expressed as: ; ; ; ; ; in, Represents the optimization problem of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; represents the time cost function of the vehicle to the task unit; Represents constraints; Represent the first constraint, second constraint, third constraint, and fourth constraint respectively; Indicates the number of task units; Indicates the number of vehicles; represents the task offloading decision at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; Indicates the upper limit of the computing capacity of the roadside unit r; represents the task cache decision at time t; Indicates the amount of data processed by the i-th task unit; Indicates the upper limit of the storage capacity of the roadside unit r.

6. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1, characterized in that: The Markov decision model of the roadside unit is expressed as: ; ; ; in, Indicates the system status of the roadside unit at time t; Indicates the system action of the roadside unit at time t; represents the reward function of the roadside unit at time t; represents the channel gain between vehicle v and roadside unit r at time t; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t; They represent the task offloading decision at time t and the task caching decision at time t respectively; represents the transmission power decision of the vehicle at time t; Represents the cost function of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; Represents the time cost function of the vehicle to task unit.

7. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 6 is characterized in that: The META-A3C algorithm is used to make online decisions on the Markov decision model to obtain the offloading cache results of the task unit, including: Determine the delay energy consumption correlation coefficient at time t based on the reward function of the roadside unit , determine the optimized scenario type at time t according to the delay-energy consumption correlation coefficient at time t , generate decision tasks based on the optimized scenario type at time t ;in, Indicates the optimization scenario type for obtaining the delay-energy consumption correlation coefficient at time t; Indicates the nth optimization scenario type; From the historical task dataset Take out K sample data and calculate the decision task based on K sample data The loss function , according to the decision task The loss function updates the policy network Parameters, value network Parameters of the updated strategy network , Update the value network ;in, 、 Represents the policy network Parameters, value network Parameters; Represent the updated policy network Parameters, update value network Parameters; Indicates the system status of the roadside unit; Indicates the system action of the roadside unit; Represents network parameters; According to the update strategy network , Update the value network , computational decision-making tasks Update loss function , according to the decision task Update loss function Calculating the meta-loss function ; According to the meta-loss function , to update the strategy network Parameters, update value network The parameters of the global policy network are updated globally to obtain the global policy network , global value network ;in, Represent the global policy network Parameters, global value network Parameters; According to the global policy network Parameters, global value network , the system status of the roadside unit at time t , execute the system action of the roadside unit at time t , obtain the system reward of the roadside unit at time t and the system status of the roadside unit at time t+1 ; Generate task data ( , , , ), executes the task data and stores the task data in the decision task Update the historical task dataset in the experience cache ; Among them, the system action of the roadside unit at time t Cache results for task unit offloading; Repeat the process until all task units have their cached results.

8. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 7 is characterized in that: The decision-making task The loss function Expressed as: ; in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Representing the use value network Estimated system status Execute system actions The action value of Representing the use value network Estimated system status Execute system actions The action value of Update policy network Parameters , Update the value network Parameters Respectively expressed as: ; ; in, Represents the policy network Learning rate, value network The learning rate; Indicates about parameters Gradient operation.

9. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 7, characterized in that: Decision-making tasks Update loss function Expressed as: ; in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the parameter update process; The meta-loss function Expressed as: ; in, Represents a collection of optimized scene types.

10. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 7, characterized in that: Global Policy Network Parameters , global value network Parameters Respectively expressed as: ; ; in, represents the meta-learning rate; Indicates about parameters Gradient operation of ; represents the meta-loss function.

Citation Information

Patent Citations

  • Internet of vehicles task unloading scheduling method and system

    CN114268923A

  • Energy-saving automatic interconnected vehicle service unloading method based on deep reinforcement learning

    CN114528042A

Cited By

  • Remote area Internet of Vehicles task unloading method and device based on high-altitude platform assistance

    CN121455699A

  • Vehicle-road cooperation task scheduling method and system based on maximum entropy deep reinforcement learning

    CN121481194A