A task offloading and caching method based on delay-energy collaborative optimization for internet of vehicles

By constructing a latency-energy collaborative optimization model and a meta-reinforcement learning algorithm, the problem of comprehensive optimization of energy consumption and latency in the Internet of Vehicles task offloading method is solved, and rapid adaptation and efficient task offloading and caching decisions are achieved, thereby improving the business performance of the Internet of Vehicles system.

CN120602999BActive Publication Date: 2025-10-17NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511109389.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-17
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing Internet of Vehicles task offloading methods fail to effectively and comprehensively consider the vehicle's energy consumption and latency issues, and are difficult to adapt to the rapid decision-making needs of intelligent connected vehicles. In addition, traditional algorithms are inefficient in dynamic performance optimization.

Method used

A vehicle network task offloading and caching method based on latency-energy consumption collaborative optimization is adopted. By constructing a latency-energy consumption collaborative optimization model and using the Markov decision model and meta-reinforcement learning algorithm (META-A3C), online optimization of task offloading and caching decisions is performed to achieve joint control of energy consumption and latency.

Benefits of technology

It realizes dynamic optimization of vehicle energy consumption and service latency in different scenarios, has rapid learning and generalization capabilities, can quickly adapt to different latency-energy consumption optimization scenarios, and improve the efficiency and accuracy of task offloading and caching decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602999B_ABST
    Figure CN120602999B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on time delay-energy consumption collaborative optimization's vehicle networking task unloading cache method, belong to vehicle networking edge computing technical field. Including: obtaining the task unit generated by vehicle;According to task unit, pre-constructed communication model of vehicle and roadside unit, the processing time delay and energy consumption when task unit is executed on vehicle, the transmission time delay when task unit is unloaded to roadside unit and the processing time delay and energy consumption when task unit is executed on roadside unit;Calculate the energy consumption cost function and time cost function of vehicle to task unit;Construct the time delay-energy consumption collaborative optimization model of task unit;The Markov decision model of roadside unit is converted into time delay-energy consumption collaborative optimization model, and the unloading cache result of task unit is obtained by carrying out online decision to Markov decision model.The application can jointly optimize two performance indexes of vehicle energy consumption and service time delay, improve the business performance of vehicle in different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle networking edge computing, and in particular to a vehicle networking task offloading and caching method based on delay-energy consumption collaborative optimization. BACKGROUND

[0002] With the rapid development of vehicle networking technology, the demand for improving the computing efficiency of vehicle networking tasks and optimizing resource utilization is increasing in today's society. Traditional vehicle networking systems usually use cloud computing technology to use remote cloud systems with sufficient computing power to provide powerful computing and storage services for vehicles. However, due to the long physical distance between vehicles and cloud computing centers, there is a large transmission delay and loss in data transmission, which cannot meet the needs of tasks such as emergency braking decision-making, real-time communication, etc. In addition, with the continuous intelligent development of vehicles, the demand for computing resources of the computing tasks generated by vehicles is also increasing. Mobile edge computing (MEC) technology provides an effective solution to the above problems, which effectively reduces the vehicle computing load by offloading tasks to the MEC server of the roadside unit (RSU) for collaborative processing.

[0003] However, the current MEC-assisted vehicle task offloading method has some problems. Related solutions only focus on the single business performance of vehicle tasks, such as service delay or energy consumption. With the rise of intelligent networked vehicles, it is essential to consider the energy consumption and delay of vehicles. At the same time, traditional vehicle task offloading algorithms are difficult to adapt to efficient autonomous decision-making, continuous decision-making, and dynamic performance-oriented targets, and need further optimization. SUMMARY

[0004] The present application aims to overcome the shortcomings of the prior art and provide a vehicle networking task offloading and caching method based on delay-energy consumption collaborative optimization, which can jointly optimize the two performance indicators of vehicle energy consumption and service delay, and improve the business performance of vehicles in different scenarios by optimizing task offloading decisions and task caching decisions.

[0005] To achieve the above-mentioned purpose, the present application is implemented by using the following technical solutions:

[0006] The present application provides a vehicle networking task offloading and caching method based on delay-energy consumption collaborative optimization, comprising:

[0007] Obtaining a task unit generated by a vehicle;

[0008] According to the task unit, the pre-constructed vehicle and the road side unit communication model, the processing delay and energy consumption of the task unit when executing on the vehicle, the transmission delay of the task unit unloading to the road side unit, and the processing delay and energy consumption of the task unit when executing on the road side unit are calculated;

[0009] According to the processing delay and energy consumption of the task unit when executing on the vehicle, the transmission delay of the task unit unloading to the road side unit, and the processing delay and energy consumption of the task unit when executing on the road side unit, the energy consumption cost function and the time cost function of the vehicle to the task unit are calculated;

[0010] According to the energy consumption cost function and the time cost function of the vehicle to the task unit, a delay-energy consumption collaborative optimization model of the task unit is constructed;

[0011] The delay-energy consumption collaborative optimization model is converted into a Markov decision model of the road side unit, and online decision is made on the Markov decision model to obtain an unloading cache result of the task unit.

[0012] Optionally, the pre-constructed vehicle and the road side unit communication model is represented as:

[0013] ;

[0014] Wherein, represents the downlink transmission rate of the pre-constructed vehicle v and the road side unit r; represents the channel bandwidth; respectively represents the channel gain and transmission power between the vehicle v and the road side unit r at the t moment; represents the additive white Gaussian noise power; represents the interference power between the vehicle v and the road side unit r at the t moment.

[0015] Optionally, the calculation of the processing delay and energy consumption of the task unit when executing on the vehicle includes:

[0016] ;

[0017] ;

[0018] Wherein, represents the processing delay of the i th task unit when executing on the vehicle v; represents the calculation amount required by the i th task unit; represents the local computing capacity of the vehicle v; represents the energy consumption of the i th task unit when executing on the vehicle v; represents the coefficient of the vehicle chip structure;

[0019] The transmission delay of the task unit offloaded to the road side unit and the calculation of the processing delay and energy consumption of the task unit when executing on the road side unit include:

[0020] ;

[0021] ;

[0022] ;

[0023] wherein, denotes the transmission delay of the ith task unit from the vehicle v to the road side unit r; denotes the task data size of the ith task unit; denotes the pre-constructed downlink transmission rate of the vehicle v and the road side unit r; denotes the processing delay of the ith task unit when executing on the road side unit; denotes the data size after processing of the ith task unit; denotes the task cache decision of the ith task unit when executing on the road side unit r at the t time; denotes the computing capacity allocated by the road side unit r to the ith task unit at the t time; denotes the energy consumption of the ith task unit when executing on the road side unit r; denotes the transmission power between the vehicle v and the road side unit r at the t time.

[0024] Optionally, the calculation of the energy consumption cost function of the vehicle to the task unit includes:

[0025] ;

[0026] ;

[0027] ;

[0028] wherein, denotes the total energy consumption of the vehicle v at the t time under the task offloading decision ; denotes the number of task units; denotes the energy consumption of the ith task unit when executing on the vehicle v; denotes the energy consumption of the ith task unit when executing on the road side unit r; denotes the energy consumption when all task units execute on the vehicle v; denotes the energy consumption cost function of the vehicle to the task unit; denotes the number of vehicles;

[0029] The calculation of the time cost function of the vehicle-to-task unit includes:

[0030] ;

[0031] ;

[0032] ;

[0033] in, Represents the task offloading decision of the i-th task unit at time t and the task cache decision at time t Total processing delay under the conditions; represents the processing delay of the i-th task unit when it is executed on the roadside unit; represents the processing delay when all task units are executed on vehicle v; represents the processing delay of the i-th task unit when it is executed on vehicle v; Represents the time cost function of the vehicle to task unit.

[0034] Optionally, the delay and energy consumption collaborative optimization model of the task unit is expressed as:

[0035] ;

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, Represents the optimization problem of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; represents the time cost function of the vehicle to the task unit; Represents constraints; Represent the first constraint, second constraint, third constraint, and fourth constraint respectively; Indicates the number of task units; Indicates the number of vehicles; represents the task offloading decision at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; Indicates the upper limit of the computing capacity of the roadside unit r; denotes the task cache decision at the t-th moment; denotes the data size processed by the i-th task unit; denotes the upper limit of the storage capacity of the road side unit r.

[0041] Optionally, the Markov decision model of the road side unit is represented as:

[0042] ;

[0043] ;

[0044] ;

[0045] wherein, denotes the system state of the road side unit at the t-th moment; denotes the system action of the road side unit at the t-th moment; denotes the reward function of the road side unit at the t-th moment; denotes the channel gain between the vehicle v and the road side unit r at the t-th moment; denotes the additive white Gaussian noise power; denotes the interference power between the vehicle v and the road side unit r at the t-th moment; denote the task offloading decision at the t-th moment and the task cache decision at the t-th moment, respectively; denotes the transmission power decision of the vehicle at the t-th moment; denotes the cost function of the time delay and energy consumption collaborative optimization model of the task unit; denote the energy consumption weight coefficient and the time delay weight coefficient, respectively; denotes the energy consumption cost function of the vehicle to the task unit; denotes the time cost function of the vehicle to the task unit.

[0046] Optionally, the META-A3C algorithm is adopted to make online decisions on the Markov decision model, so as to obtain the offloading and caching results of the task unit, including:

[0047] determining the time delay and energy consumption correlation coefficient at the t-th moment according to the reward function of the road side unit determining the optimization scene type at the t-th moment according to the time delay and energy consumption correlation coefficient at the t-th moment generating a decision task according to the optimization scene type at the t-th moment ; wherein, denotes the optimization scene type constituted by the time delay and energy consumption correlation coefficient at the t-th moment; denotes the n-th optimization scene type;

[0048] from a historical task data set Take out K sample data and calculate the decision task based on K sample data The loss function , according to the decision task The loss function updates the policy network Parameters, value network Parameters of the updated strategy network , Update the value network ;in, 、 Represents the policy network Parameters, value network Parameters; Represent the updated policy network Parameters, update value network Parameters; Indicates the system status of the roadside unit; Indicates the system action of the roadside unit; Represents network parameters;

[0049] According to the update strategy network , Update the value network , computational decision-making tasks Update loss function , according to the decision task Update loss function Calculating the meta-loss function ;

[0050] According to the meta-loss function , to update the strategy network Parameters, update value network The parameters of the global policy network are updated globally to obtain the global policy network , global value network ;in, Represent the global policy network Parameters, global value network Parameters;

[0051] According to the global policy network Parameters, global value network , the system status of the roadside unit at time t , execute the system action of the roadside unit at time t , obtain the system reward of the roadside unit at time t and the system status of the roadside unit at time t+1 ; Generate task data ( , , , ), executes the task data and stores the task data in the decision task Update the historical task dataset in the experience cache ; Among them, the system action of the roadside unit at time t Cache results for task unit offloading;

[0052] Repeat the process until all task units have their cached results.

[0053] Optionally, the decision-making task The loss function Expressed as:

[0054] ;

[0055] in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Representing the use value network Estimated system status Execute system actions The action value of Representing the use value network Estimated system status Execute system actions The action value of

[0056] Update policy network Parameters , Update the value network Parameters Respectively expressed as:

[0057] ;

[0058] ;

[0059] in, Represents the policy network Learning rate, value network The learning rate; Indicates about parameters Gradient operation.

[0060] Optional, decision-making tasks Update loss function Expressed as:

[0061] ;

[0062] wherein, denotes the mathematical expectation at the t-th moment; denotes the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample, respectively; denotes the superposition factor; denotes using the updated value network to estimate the action value of performing a system action under a system state ; denotes using the updated value network to estimate the action value of performing a system action under a system state ; denotes a parameter updating process;

[0063] The meta-loss function is represented as:

[0064] ;

[0065] wherein, denotes an optimized set of scene types.

[0066] Optionally, the parameters of the global policy network , the parameters of the global value network are respectively represented as: ;

[0067] ;

[0068] ;

[0069] wherein, denotes a meta-learning rate; denotes gradient operation on the parameters ; denotes the meta-loss function.

[0070] Compared with the prior art, the present application has the beneficial effects that:

[0071] The present application realizes dynamic optimization of two indexes of vehicle energy consumption and service time delay under different scenes by jointly controlling the task offloading decision and the data caching decision, can realize rapid learning ability for different optimization target tasks by only learning decision experience under different scenes online, can obtain rapid generalization ability for dynamic optimization target tasks by only a small sample training process, can quickly adapt to different time delay-energy consumption optimization scenes, and can make the task offloading and caching decision problem quickly converge. BRIEF DESCRIPTION OF DRAWINGS

[0072] ​Figure 1 Fig. 1 shows a flowchart of the method for task offloading and caching based on latency-energy collaborative optimization in a vehicle network according to an embodiment of the present application;

[0073] Figure 2 Fig. 4 shows a graph of the relationship between the convergence performance of the algorithm and the training period under the condition of a mutation of the optimization parameter according to an embodiment of the present application;

[0074] Figure 3 Fig. 4 shows a graph of the relationship between the convergence performance of the algorithm and the training period under the condition of a mutation of the optimization parameter according to an embodiment of the present application;

[0075] Figure 4 Fig. 5 shows a line graph of the total energy consumption of tasks under different numbers of task units according to an embodiment of the present application;

[0076] Figure 5 Fig. 5 shows a line graph of the total energy consumption of tasks under different numbers of task units according to an embodiment of the present application; DETAILED DESCRIPTION

[0077] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, and are not limitations of the technical solutions of the present application. In the case of no conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0078] The term "and / or", only describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " generally represents that the associated objects before and after it are in an "or" relationship.

[0079] Embodiment 1

[0080] As Figure 1 shown, the present embodiment introduces a task offloading and caching method based on latency-energy collaborative optimization in a vehicle network, which considers the dynamic latency, energy consumption requirements of the vehicle network scene, and the joint design problem of task offloading and data caching under the condition of the computing power and storage capacity of the edge computing server. The method comprises:

[0081] Step 1: Establish a communication model between the vehicle and the road side unit (RSU), which is represented as:

[0082] ;

[0083] Among them, represents the pre-constructed downlink transmission rate of the vehicle v and the road side unit r; represents the channel bandwidth; respectively represent the channel gain, transmission power between vehicle v and road side unit r at time t; represents the additive white Gaussian noise power; represents the interference power between vehicle v and road side unit r at time t.

[0084] Step two: establish a computational model of the task unit generated by the vehicle, specifically:

[0085] The vehicle continuously generates task units, which can be calculated by the vehicle itself or offloaded to the RSU for execution. After the task is processed, the result is transmitted to the vehicle through the communication model wireless link.

[0086] When the ith task unit generated by the vehicle v is calculated locally on the vehicle, the processing delay of the ith task unit when executed on the vehicle v is is represented as:

[0087] ;

[0088] The energy consumption of the vehicle only includes the calculation energy consumption of the ith task unit locally on the vehicle, and the energy consumption of the ith task unit when executed on the vehicle v is is represented as:

[0089] ;

[0090] wherein, represents the calculation amount required by the ith task unit, in cycles; represents the local computing capacity of the vehicle v, i.e. CPU frequency, in cycles / s; represents the coefficient of the vehicle chip structure.

[0091] When the ith task unit generated by the vehicle v is offloaded to the RSU for execution, the transmission delay of the ith task unit from the vehicle v to the road side unit r is is represented as:

[0092] ;

[0093] After the task is transmitted to the RSU, the RSU first queries whether there is cached data of the processing result of the ith task unit in the local cache pool. If there is cached data of the task, the cached result is directly transmitted to the vehicle v. If there is no cached data, the task is processed by the Mobile Edge Computing (MEC) server in the RSU. At this time, the processing delay of the ith task unit when executed on the road side unit is is represented as:

[0094] ;

[0095] The energy consumption of the vehicle includes the transmission energy consumption of the ith task unit and its processing result, and the energy consumption of the ith task unit when executed on the road side unit r is expressed as:

[0096] ;

[0097] wherein, denotes the task data size of the ith task unit; denotes the data size processed by the ith task unit; denotes the task cache decision of the ith task unit when executed on the road side unit r at the tth moment, =1 indicates that the RSU caches the processing data of the ith task unit, otherwise ; denotes the computing capacity allocated by the road side unit r to the ith task unit at the tth moment, in cycles / s; denotes the transmission power between the vehicle v and the road side unit r at the tth moment.

[0098] Step three: calculate the energy consumption cost function and time cost function of the vehicle to the task unit, specifically:

[0099] The energy consumption cost function of the vehicle is defined as the ratio of the energy consumption of the vehicle under a certain task offloading decision state to the energy consumption without task offloading, and the total energy consumption of the vehicle v at the tth moment under the task offloading decision . is expressed as:

[0100] ;

[0101] When all the tasks of the vehicle v are executed locally, i.e. no task offloading to the RSU, the energy consumption of the task unit when executed on the vehicle v is expressed as:

[0102] ;

[0103] The energy consumption cost function of the vehicle to the task unit is expressed as:

[0104] ;

[0105] wherein, denotes the number of vehicles.

[0106] The time cost function of the vehicle v to the task unit is defined as the ratio of the time required for the vehicle to obtain the task processing result under certain task offloading decision and task caching decision and the time required for local processing of the task, and the total processing delay of the ith task unit under the task offloading decision of the ith task unit at the t time and the task caching decision of the ith task unit at the t time is expressed as:

[0107] ;

[0108] When all the task units of the vehicle v are calculated by the local processor, the processing delay of the task units when all the task units are executed on the vehicle v is expressed as:

[0109] ;

[0110] The time cost function of the vehicle to the task unit is expressed as:

[0111] .

[0112] Step four: constructing a time delay and energy consumption collaborative optimization model of the task unit, specifically:

[0113] In the vehicle networking system, vehicles continuously generate task units to be processed, and RSUs achieve the goal of relieving vehicle load and reducing service delay by offloading the task computing process and caching the task processing result. This process involves optimization of task offloading decision and caching decision, that is, under the constraints of RSU computing ability and storage ability, by controlling the task offloading decision of the ith task unit at the t time and the task caching decision of the ith task unit at the t time , the joint optimization of the task service delay at the t time and the vehicle energy consumption is achieved, that is:

[0114] The cost function of the time delay and energy consumption collaborative optimization model of the task unit is constructed , which is expressed as:

[0115] ;

[0116] wherein, and are the energy consumption weight coefficient and the time delay weight coefficient, respectively, satisfying ;

[0117] The task offloading decision constraint condition is constructed. The total amount of tasks that can be offloaded by the RSU at any time needs to satisfy the computing ability constraint, that is:

[0118] ;

[0119] ​The task cache decision constraint condition is constructed. The total amount of task data that the RSU can store at any time needs to meet the storage space limit, that is, the following condition is met:

[0120]

[0121] By optimizing the task offloading decision and the task cache decision, the energy consumption and the service delay of the vehicle are reduced. The delay-energy consumption collaborative optimization model of the task unit is represented as follows:

[0122]

[0123]

[0124]

[0125]

[0126]

[0127] The optimization problem of the delay-energy consumption collaborative optimization model of the task unit is represented as follows: The constraint is represented as follows: The first constraint, the second constraint, the third constraint, and the fourth constraint are represented as follows: The number of task units is represented as follows: The task offloading decision at the t-th moment is represented as follows: The computing power of the roadside unit r allocated to the i-th task unit at the t-th moment is represented as follows, and the unit is cycles / s: The upper limit of the computing power of the roadside unit r is represented as follows: The task cache decision at the t-th moment is represented as follows: The upper limit of the storage capacity of the roadside unit r is represented as follows.

[0128] The optimization problem P1 can be converted into a Markov decision process (MDP) problem based on state-action value. Due to the dynamic optimization target caused by the dynamic parameters in the problem, the training efficiency of the traditional reinforcement learning (RL) based method is low, and there is no clear training direction. Therefore, a decision method based on meta-reinforcement learning is designed to realize online decision of the optimization problem P1.

[0129] Step five: convert the optimization problem of the delay-energy consumption collaborative optimization model into a Markov decision model of the roadside unit, and the specific conversion is as follows:

[0130] The channel condition, noise, and interference environment of the RSU are dynamically changing. Therefore, the system state of the roadside unit is represented as follows: ​​​​​​​As time goes by, the system status of the roadside unit at time t is Expressed as:

[0131] ;

[0132] RSU needs to be based on the system status of the roadside unit at time t Make dynamic decisions, the system action of the roadside unit at time t Expressed as:

[0133] ;

[0134] The reward function is used to calculate RSU in the system state Execute system actions The system reward obtained, the MDP process usually takes maximizing the reward as the decision-making goal. The goal of the optimization problem is to simultaneously reduce the energy consumption of the vehicle and the service delay. Therefore, the cost function of the delay-energy collaborative optimization model of the task unit can be taken as a negative value to construct the reward function, the reward function of the roadside unit Expressed as:

[0135] ;

[0136] in, represents the transmission power decision of the vehicle.

[0137] Step 6: Make an online decision on the Markov decision model to obtain the offloading cache result of the task unit, specifically:

[0138] A Meta-Reinforcement Learning with Asynchronous Advantage Actor-Critic (META-A3C) algorithm is designed to perform online decision-making on multi-objective tasks. The delay-energy consumption correlation coefficient at time t is determined based on the reward function of the roadside unit. , according to the delay energy consumption correlation coefficient at time t Determine the optimization scenario type at time t ,Right now:

[0139] ;

[0140] Generate decision tasks based on the optimized scenario type at time t ;in, Indicates the optimization scenario type for obtaining the delay-energy consumption correlation coefficient at time t; Indicates the nth optimization scenario type.

[0141] From the historical task dataset K sample data are taken out , a loss function of the decision task is calculated according to the K sample data , the loss function of the decision task is expressed as:

[0142] ;

[0143] wherein, denotes a mathematical expectation at the t th moment; denote the t th moment of the 1 st sample, the t th moment of the 2 nd sample, …, the t th moment of the K th sample, respectively; denotes a superposition factor; denotes an action value estimated by using a value network to execute a system action under a system state ; denotes an action value estimated by using a value network to execute a system action under a system state ; , denote parameters of a policy network , parameters of a value network , respectively; denote a system state of the 1 st sample at the t th moment, a system state of the 2 nd sample at the t th moment, …, a system state of the K th sample at the t th moment, respectively; denote a system action of the 1 st sample at the t th moment, a system action of the 2 nd sample at the t th moment, …, a system action of the K th sample at the t th moment, respectively; denote a system reward of the 1 st sample at the t th moment, a system reward of the 2 nd sample at the t th moment, …, a system reward of the K th sample at the t th moment, respectively; denote a system state of the 1 st sample at the t+1 th moment, a system state of the 2 nd sample at the t+1 th moment, …, a system state of the K th sample at the t+1 th moment, respectively.

[0144] In this embodiment, the global policy network , the global value network are updated in real time, and at the current network update round, the parameters of the policy network , the parameters of the value network are the parameters of the global policy network, the parameters of the global value network obtained at the last network update round.

[0145] According to the decision task ​​loss function update strategy network parameters of the value network parameters of the value network , to obtain an updated strategy network , represents a system state of the road side unit; represents a system action of the road side unit; update the strategy network parameters of the strategy network , update the value network parameters of the value network are respectively represented as:

[0146] ;

[0147] ;

[0148] wherein, learning rate of the strategy network learning rate of the value network ; represents a gradient operation on the parameters .

[0149] According to the updated strategy network , update the value network , calculate an update loss function of the decision task , the update loss function of the decision task is represented as:

[0150] ;

[0151] wherein, represents an action value of performing a system action in a system state estimated using the updated value network ; represents an action value of performing a system action in a system state estimated using the updated value network ; represents a parameter update process.

[0152] According to the update loss function of the decision task , calculate a meta loss function , the meta loss function is represented as:

[0153] ;

[0154] wherein, represents an optimized scene type set.

[0155] The meta-loss function is subjected to gradient descent, and the parameters of the value network are updated according to the meta-loss function , to update the parameters of the policy network , the parameters of the value network , the parameters of the policy network , and the parameters of the value network , respectively.

[0156]

[0157]

[0158] wherein, represents a meta-learning rate.

[0159] According to the parameters of the policy network , the value network , and the system state of the roadside unit at the t th moment , the system action of the roadside unit at the t th moment is executed, the system reward of the roadside unit at the t th moment and the system state of the roadside unit at the t+1 th moment are obtained; the task data is generated, , , , the task data is executed and stored in the experience buffer of the decision task , and the historical task data set is updated; wherein the system action of the roadside unit at the t th moment is the unloading cache result of the task unit.

[0160] This process is continuously executed, and the update and optimization of the task unloading and cache decision are realized in the process until the unloading cache results of all task units are obtained.

[0161] Embodiment 2

[0162] This embodiment introduces a test example of a vehicle networking task unloading and caching method based on delay-energy consumption collaborative optimization:

[0163] Figure 2 ​​​​​For the embodiment, the algorithm convergence performance and training cycle relationship graph under the unchanged optimization parameters is compared with other algorithms, such as Meta-Reinforcement Learning with Asynchronous Advantage Actor-Critic (META-A3C) algorithm, Asynchronous Advantage Actor-Critic (A3C) algorithm, and Deep Deterministic Policy Gradient (DDPG) algorithm, under the fixed optimization parameters. The horizontal axis is the training round, and the vertical axis is the cumulative reward value. The results show that the Meta-A3C algorithm reaches stable convergence within 100 rounds, while the A3C algorithm and the DDPG algorithm need 150 rounds and 200 rounds respectively. The fast convergence of the Meta-A3C algorithm benefits from the pre-training optimization of the global parameters by meta-learning, which enables it to efficiently utilize historical experience and reduce exploration costs. In addition, the final cumulative reward of the Meta-A3C algorithm is significantly higher than that of other algorithms, verifying the effectiveness of its joint optimization of energy consumption and latency.

[0164] Figure 3 For the embodiment, the algorithm convergence performance and training cycle relationship graph under the dynamic scene (changing optimization parameters) is shown. When the training reaches 250 rounds, the energy consumption weight coefficient increases from 0.3 to 0.7, and the latency weight coefficient decreases from 0.7 to 0.3. The Meta-A3C algorithm only needs about dozens of rounds to re-converge, while the A3C algorithm and the DDPG algorithm need more than 100 additional rounds. This result highlights the meta-learning mechanism of the Meta-A3C algorithm, which can quickly adapt to changes in the "latency-energy" optimization ratio, dynamically adjust the policy network parameters, and quickly respond to new optimization directions, demonstrating the robustness of the algorithm in dynamic environments.

[0165] Figure 4 For the embodiment, the total energy consumption of tasks under different numbers of task units is analyzed. In = =0.5, the graph shows the trend of total energy consumption of tasks for each algorithm as the number of task units increases. From Figure 4 , it can be seen that the total energy consumption of tasks under the algorithm of the embodiment is always lower than that of other algorithms, and the energy consumption advantage becomes more obvious as the number of task units increases. This shows that the method of the embodiment can better optimize task offloading and storage decisions, balance energy consumption and latency, and achieve efficient use of resources when handling a large number of tasks.

[0166] Figure 5 The total task delay curve of the embodiment under different numbers of task units shows the influence of the increase of the number of task units on the total delay. Also in the condition of = 0.5, the figure shows the total task delay of each algorithm under different numbers of task units. It can be seen from the figure that the total task delay of the method of the embodiment is at a low level under each number of task units, and the advantage of the delay is more and more obvious with the increase of the number of task units, which further proves that the algorithm can effectively reduce the delay when processing tasks of different scales, guarantee the fast processing of the task of the Internet of Vehicles, and improve the user experience. Figure 5 The total task delay curve of the embodiment under different numbers of task units shows the influence of the increase of the number of task units on the total delay. Also in the condition of = 0.5, the figure shows the total task delay of each algorithm under different numbers of task units. It can be seen from the figure that the total task delay of the method of the embodiment is at a low level under each number of task units, and the advantage of the delay is more and more obvious with the increase of the number of task units, which further proves that the algorithm can effectively reduce the delay when processing tasks of different scales, guarantee the fast processing of the task of the Internet of Vehicles, and improve the user experience.

[0167] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media containing computer usable program code (including but not limited to disk storage, CD-ROM, optical storage, etc.).

[0168] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks. Figure 1 The functions specified in one or more flows and / or blocks.

[0169] These computer program instructions can also be stored in a computer readable storage medium that can guide the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks. Figure 1 The functions specified in one or more flows and / or blocks.

[0170] ​​These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are generated to realize the computer-implemented processes, and the instructions executed on the computer or other programmable devices provide a process for implementing the functions specified in the flowchart Figure 1 one flow or multiple flows and / or the functions specified in the block Figure 1 one flow or multiple flows and / or the functions specified in the block

[0171] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and those of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims, which are all within the protection of the present application.

Claims

1. A method for caching and offloading tasks in an Internet of Vehicles (IoV) based on coordinated optimization of latency and energy consumption, characterized in that: include: Get the task units generated by the vehicle; Based on the task unit and the pre-built communication model between the vehicle and the roadside unit, the processing delay and energy consumption of the task unit when executing on the vehicle, the transmission delay when the task unit is offloaded to the roadside unit, and the processing delay and energy consumption when the task unit is executed on the roadside unit are calculated; Calculate the energy cost function and time cost function of the vehicle for the task unit based on the processing delay and energy consumption of the task unit when executed by the vehicle, the transmission delay of the task unit unloading to the roadside unit, and the processing delay and energy consumption of the task unit when executed by the roadside unit. Based on the energy consumption cost function and time cost function of the vehicle to the task unit, a delay energy consumption collaborative optimization model of the task unit is constructed; Converting the delay-energy collaborative optimization model into a Markov decision model of a roadside unit, performing online decision-making on the Markov decision model, and obtaining an unloading cache result of the task unit; The delay and energy consumption collaborative optimization model of the task unit is expressed as: ; ; ; ; ; in, Represents the optimization problem of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; represents the time cost function of the vehicle to the task unit; Represents constraints; Represent the first constraint, second constraint, third constraint, and fourth constraint respectively; Indicates the number of task units; Indicates the number of vehicles; represents the task offloading decision at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; Indicates the upper limit of the computing capacity of the roadside unit r; represents the task cache decision at time t; Indicates the amount of data processed by the i-th task unit; Indicates the upper limit of the storage capacity of the roadside unit r.

2. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1 is characterized in that: The pre-built communication model between the vehicle and the roadside unit is expressed as: ; in, represents the pre-built downlink transmission rate between vehicle v and roadside unit r; Indicates the channel bandwidth; denote the channel gain and transmission power between vehicle v and roadside unit r at time t, respectively; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t.

3. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 2 is characterized in that: The calculation of the processing delay and energy consumption of the task unit when executed on the vehicle includes: ; ; in, represents the processing delay of the i-th task unit when it is executed on vehicle v; Indicates the amount of computation required for the i-th task unit; represents the local computing capability of vehicle v; represents the energy consumption of the i-th task unit when executed on vehicle v; Coefficients representing the vehicle chip structure; The calculation of the transmission delay of the task unit unloaded to the roadside unit and the processing delay and energy consumption when the task unit is executed on the roadside unit includes: ; ; ; in, represents the transmission delay of the i-th task unit from vehicle v to roadside unit r; Indicates the task data size of the i-th task unit; represents the pre-built downlink transmission rate between vehicle v and roadside unit r; represents the processing delay of the i-th task unit when it is executed on the roadside unit; Indicates the amount of data processed by the i-th task unit; represents the task cache decision when the i-th task unit is executed on the roadside unit r at time t; It represents the computing capacity allocated by the roadside unit r to the i-th task unit at time t; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the transmission power between vehicle v and roadside unit r at time t.

4. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 3 is characterized in that: The calculation of the energy consumption cost function of the vehicle to task unit includes: ; ; ; in, represents the task offloading decision of vehicle v at time t Total energy consumption under the conditions; Indicates the number of task units; represents the energy consumption of the i-th task unit when executed on vehicle v; represents the energy consumption of the i-th task unit when it is executed on the roadside unit r; represents the energy consumption when all task units are executed on vehicle v; represents the energy consumption cost function of the vehicle for the task unit; Indicates the number of vehicles; The calculation of the time cost function of the vehicle-to-task unit includes: ; ; ; in, Represents the task offloading decision of the i-th task unit at time t and the task cache decision at time t Total processing delay under the conditions; represents the processing delay of the i-th task unit when it is executed on the roadside unit; represents the processing delay when all task units are executed on vehicle v; represents the processing delay of the i-th task unit when it is executed on vehicle v; Represents the time cost function of the vehicle to the task unit.

5. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 1, characterized in that: The Markov decision model of the roadside unit is expressed as: ; ; ; in, Indicates the system status of the roadside unit at time t; Indicates the system action of the roadside unit at time t; represents the reward function of the roadside unit at time t; represents the channel gain between vehicle v and roadside unit r at time t; represents the additive white Gaussian noise power; represents the interference power between vehicle v and roadside unit r at time t; They represent the task offloading decision at time t and the task caching decision at time t respectively; represents the transmission power decision of the vehicle at time t; Represents the cost function of the delay-energy collaborative optimization model of the task unit; Respectively represent the energy consumption weight coefficient and the delay weight coefficient; represents the energy consumption cost function of the vehicle for the task unit; Represents the time cost function of the vehicle to the task unit.

6. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 5 is characterized in that: The META-A3C algorithm is used to make online decisions on the Markov decision model to obtain the offloading cache results of the task unit, including: Determine the delay energy consumption correlation coefficient at time t based on the reward function of the roadside unit , determine the optimized scenario type at time t according to the delay-energy consumption correlation coefficient at time t , generate decision tasks based on the optimized scenario type at time t ;in, Indicates the optimization scenario type for obtaining the delay-energy consumption correlation coefficient at time t; Indicates the nth optimization scenario type; From the historical task dataset Take out K sample data and calculate the decision task based on K sample data The loss function , according to the decision task The loss function updates the policy network Parameters, value network Parameters of the updated strategy network , Update the value network ;in, 、 Represents the policy network Parameters, value network Parameters; Represent the updated strategy network Parameters, update value network Parameters; Indicates the system status of the roadside unit; Indicates the system action of the roadside unit; Represents network parameters; According to the update strategy network , Update the value network , computational decision-making tasks Update loss function , according to the decision task Update loss function Calculating the meta-loss function ; According to the meta-loss function , to update the strategy network Parameters, update value network The parameters of the global policy network are updated globally to obtain the global policy network , global value network ;in, Represent the global policy network Parameters, global value network Parameters; According to the global policy network Parameters, global value network , the system status of the roadside unit at time t , execute the system action of the roadside unit at time t , obtain the system reward of the roadside unit at time t and the system status of the roadside unit at time t+1 ; Generate task data ( , , , ), executes the task data and stores the task data in the decision task Update the historical task dataset in the experience cache ; Among them, the system action of the roadside unit at time t Cache results for task unit offloading; Repeat the process until all task units have their cached results.

7. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 6 is characterized in that: The decision-making task The loss function Expressed as: ; in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Representing the use value network Estimated system status Execute system actions The action value of Representing the use value network Estimated system status Execute system actions The action value of Update policy network Parameters , Update the value network Parameters Respectively expressed as: ; ; in, Represents the policy network Learning rate, value network The learning rate; Indicates about parameters Gradient operation.

8. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 7 is characterized in that: Decision-making tasks Update loss function Expressed as: ; in, represents the mathematical expectation at time t; They represent the t-th moment of the first sample, the t-th moment of the second sample, …, the t-th moment of the K-th sample respectively; represents the superposition factor; Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the use of updated value network Estimated system status Execute system actions The action value of Represents the parameter update process; The meta-loss function Expressed as: ; in, Represents a collection of optimized scene types.

9. The method for caching Internet of Vehicles tasks based on latency-energy consumption coordinated optimization according to claim 8, characterized in that: Global Policy Network Parameters , global value network Parameters Respectively expressed as: ; ; in, represents the meta-learning rate; Indicates about parameters Gradient operation of ; represents the meta-loss function.

Citation Information

Patent Citations

  • Internet of vehicles task unloading scheduling method and system

    CN114268923A

  • Energy-saving automatic interconnected vehicle service unloading method based on deep reinforcement learning

    CN114528042A