State update method for efficient caching and task offloading in drone-assisted connected vehicles

By combining caching and drone technologies, a strategy for cache updating and resource allocation in a dynamic environment was designed. A deep deterministic strategy gradient algorithm was used to optimize vehicle state updates, solving the problems of information freshness and energy consumption in traditional vehicle-to-everything (V2X) architectures and improving the safety and timeliness of autonomous driving.

CN114626298BActive Publication Date: 2025-10-28BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210246271.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-10-28
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

Traditional vehicle-to-everything (V2X) architectures struggle to effectively guarantee the freshness and security of computation results in autonomous driving scenarios. Especially in unexpected traffic situations, existing technologies often neglect the freshness of cached data, resulting in insufficient latency metrics during the computation offloading process to ensure the timeliness and accuracy of information.

Method used

Combining caching and UAV technologies, a strategy for cache updates and resource allocation in a dynamic environment is designed. The vehicle state update is optimized through a deep deterministic policy gradient algorithm to ensure information freshness and minimize system energy consumption. The deep deterministic policy gradient algorithm is used for resource allocation decisions, and an experience cache pool is partitioned to improve training efficiency.

Benefits of technology

It ensures the timeliness of information in autonomous driving scenarios, reduces system energy consumption, improves the freshness and security of calculation results, and adapts to the vehicle-to-everything (V2X) environment characterized by high vehicle mobility, intensive tasks, and rapid changes in network topology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626298B_ABST
    Figure CN114626298B_ABST
Patent Text Reader

Abstract

This invention discloses an efficient state update method for caching and task unloading in UAV-assisted vehicle-to-everything (V2X) networks. First, to comprehensively ensure the safety of autonomous driving, the freshness of both the caching model used for computation and the freshness of the computation unloading process are considered as the information freshness for vehicle state updates. Second, combining vehicle edge computing, caching, and UAV technologies, a strategy is designed to minimize system energy consumption and ensure information timeliness by making decisions on cache updates, user associations, and resource allocation in a dynamic environment. Finally, a deep reinforcement learning algorithm based on deep deterministic policy gradients is used. The experience cache pool is divided, and experience is selected proportionally from two different cache pools to train the neural network. This accelerates the convergence speed, reduces reward oscillations after convergence, increases the reward value after convergence, and improves algorithm performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a method for efficient caching and task unloading state updates in unmanned aerial vehicle-assisted vehicle networks. Background Technology

[0002] With the continuous development of the Internet of Vehicles (IoV), various IoV applications and services have emerged, bringing a large amount of data traffic to the IoV. This has exacerbated the contradiction between computational overload and limited spectrum and computing resources in the traditional vehicle-to-base station IoV architecture. This contradiction poses a challenge to the traditional IoV architecture in supporting latency-sensitive and resource-demanding applications and services, as well as responding to sudden traffic situations. As one of the important application scenarios of IoV, autonomous driving, due to its frequent perception of the external environment and the resulting large data traffic, makes the above problems even more severe. Therefore, numerous studies have shown that introducing mobile edge computing into the IoV architecture, using vehicle edge computing technology to offload the vehicle's task computation process to the network edge, reduces the pressure on the core network, and can effectively address the above challenges.

[0003] Caching, a technique that predicts and pre-stores content on servers, reduces frequent response times and is often used to assist edge computing in optimizing performance metrics. Caching is also frequently used as an aid in designing strategies for vehicle edge computing. By introducing caching, computational latency can be effectively reduced, while system energy consumption can also be decreased.

[0004] Drone technology, with its flexibility and mobility, can be used to assist vehicle edge computing. When encountering sudden traffic situations or large surges in traffic flow, vehicle edge computing architectures often struggle to support various task requirements. In such cases, dispatching nearby drones equipped with edge computing servers to the affected areas can effectively address the problem. Therefore, in existing technologies, drone technology is frequently used as an auxiliary technology for vehicle edge computing.

[0005] In various applications and services of the Internet of Vehicles (IoV), vehicles often have certain requirements regarding the freshness of the information they receive. In autonomous driving scenarios, the task calculation results obtained by the vehicle need to reflect the external situation in a timely, accurate, and efficient manner, and need to have a small difference from the actual external situation; that is, fresh calculation results are required. In recent years, information age has been regarded as a method for measuring information freshness, defined as the time elapsed since the information was generated. Currently, a large number of studies focus on this freshness in the computational unloading process to characterize the freshness of the results obtained in the computational unloading process.

[0006] The primary goal of research on vehicle edge computing in autonomous driving scenarios is to ensure safety during the autonomous driving process. To guarantee vehicle safety, vehicles need to frequently perceive environmental information and interact with the environment extensively and frequently, while also continuously and intensively generating computational tasks. This poses a challenge to traditional vehicle-to-everything (V2X) architectures. Existing technologies often ensure the safety and user experience of autonomous driving by controlling the latency of the computation offloading process. In reality, using latency as an indicator to characterize whether the computational results obtained by autonomous vehicles can timely and accurately reflect the actual external situation is clearly insufficient. Therefore, this invention uses information age to characterize the difference between the computational results obtained by the vehicle and the actual external situation. Furthermore, previous computational offloading strategies that consider information age often only focus on the computational offloading process, neglecting the information freshness of the cached computational model used. This invention addresses this deficiency by incorporating the information freshness of cached content into the freshness of the computational results, thus comprehensively ensuring the freshness of the information results obtained by the vehicle. Summary of the Invention

[0007] This invention addresses the challenges of cache-assisted UAV-vehicle edge computing architectures and autonomous driving scenarios, which face high requirements for safety and timeliness, high vehicle mobility, intensive tasks, rapid network topology changes, and high system overhead. By combining caching and UAV technologies, this invention designs a strategy to minimize system energy consumption and ensure information timeliness by making decisions on cache updates, user associations, and resource allocation in a dynamic environment. This provides an efficient state update method for cache and task offloading in UAV-assisted vehicle networks.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A method for efficient caching and task offloading state updates in a drone-assisted vehicle network, the system model including a macro base station with an MEC server, a drone with computing capabilities, and multiple vehicles with computing capabilities, and employing the following steps:

[0010] S1. The freshness of the data used in the calculation task and the freshness of the unloading process are used as the freshness of the vehicle's status update.

[0011] S2. Minimize the total energy consumption of the system in time slot t. The total energy consumption includes the energy consumption generated by updating the cached content in time slot t and the energy consumption generated by vehicle task processing in time slot t.

[0012] S3. An algorithm based on deep deterministic policy gradients divides the experience cache pool and selects experience from two different cache pools proportionally to train the neural network.

[0013] Furthermore, the method for calculating the freshness of the cache model in step S1 is as follows:

[0014] definition and This indicates whether vehicle i and UAV j have cached content w in time slot t. as well as This indicates that there is a cache w, otherwise it is 0. The cache capacity of vehicles and UAVs is expressed as:

[0015]

[0016]

[0017] in, and Indicates the cache capacity of vehicle i and UAV j, l w This represents the amount of data in content w;

[0018] definition and These are two binary variables, representing whether vehicle i and UAV j update or replace cache w in time slot t. as well as This indicates that cache w is updated or replaced in time slot t, otherwise it is 0; for vehicle i:

[0019]

[0020] Similarly, for UAV j, we can obtain:

[0021]

[0022] and Let represent the timeliness of the cached content for vehicles and UAVs, respectively, i.e., the freshness of the cache. Assuming that the time it takes for content w to be replaced by other types of cached data is set to an infinite number I, we get:

[0023]

[0024]

[0025] Furthermore, the method for calculating the freshness of the unloading process in step S1 is as follows:

[0026] The vehicle generates a computational task and sends a resource access request to the MeNB, the request including the {e} of task w. i (t), s w , z w}, driving direction, driving speed, current location information, where e i (t) represents the task type of vehicle i in time slot t, s wThis indicates the size of the task, w, and z. w This indicates the CPU cycles required for the computation of subtask w; after collecting this information, the MeNB makes a decision on the processing of vehicle task i, using... Indicates whether the task is processed locally, using It can be expressed in the following ways:

[0027] 1) Local computation:

[0028] 2) Uninstall to UAV:

[0029] 3) Uninstall to MeNB:

[0030] 4) Not calculated for now:

[0031] Then we have:

[0032]

[0033] Furthermore, during local computation, use This represents the computing power of vehicle i, and the local computing latency of vehicle i is expressed as:

[0034]

[0035] The corresponding energy consumption is expressed as follows:

[0036]

[0037] Weighting factor μ i The energy consumption required for vehicle i's CPU to perform calculations per cycle is represented as:

[0038]

[0039] The local computation delay does not exceed one time slot τ, and we have:

[0040]

[0041] Furthermore, when unloading to a UAV, the transmission rate between vehicle i and UAV j is expressed as:

[0042]

[0043] Among them, b i,j (t) is the bandwidth when vehicle i communicates with UAV j and MeNB, p is the vehicle's mission transmit power, and β is the bandwidth. i,j (t) is the channel gain when vehicle i communicates with UAV j, σ 2 It is the power spectral density of white noise;

[0044] The channel gain between vehicle i and UAV j is expressed as:

[0045]

[0046] Where, d i,j (t) represents the distance between vehicle i and UAV j;

[0047] Establish a coordinate axis on this area, using Let x and y coordinates represent the positions of the vehicle and UAV, and let x and y coordinates represent the speeds of the vehicle and UAV. Positive and negative signs indicate the direction of travel. The distance between vehicle i and UAVj is expressed as:

[0048]

[0049] If the task of vehicle i is offloaded to UAV j at time t, then i must be within the coverage area of ​​UAV j within one time slot. That is, the connection distance between vehicle i and UAV j must be greater than the communication radius R of the UAV at the start of the next time slot. Therefore:

[0050]

[0051] Furthermore, when offloading to the MeNB, the transmission rate between vehicle i and the MeNB is as follows:

[0052]

[0053] Among them, b i,M+1 (t) is the bandwidth when vehicle i communicates with the MeNB, g i,M+1 (t) is the square of the average channel gain between vehicle i and MeNB, where:

[0054] The time it takes for vehicle i to transmit the task to UAVj or MeNB is represented as:

[0055]

[0056] The latency of the MEC server in the UAV and MeNB calculating the task of vehicle i is expressed as:

[0057]

[0058] in, This represents the computing resources allocated to vehicle i in MEC server j during time slot t;

[0059] The total uninstallation time is:

[0060]

[0061] The total unloading delay is limited to one time slot:

[0062]

[0063] Energy consumption during unloading:

[0064]

[0065] definition The freshness of the state update of vehicle i in time slot t is represented by, i.e., the age of the state update, using... The age of the state update represents the time when task w (where w is the first task in the task queue to be processed) was generated.

[0066]

[0067] The regulation stipulates that it cannot exceed threshold A. th :

[0068]

[0069] Furthermore, the energy consumption ξ(t) generated by cache updates within time slot t is expressed as:

[0070]

[0071] Wherein, θ(J / bit) is a weighting factor that converts the amount of cached data to energy, representing the energy required to cache 1 bit of data.

[0072] Energy consumption E generated by task processing within time slot t i (t) is represented as:

[0073]

[0074] The energy consumption φ(t) of the system due to task transmission in time slot t is expressed as:

[0075]

[0076] The system's total energy consumption in time slot t Represented as:

[0077]

[0078] Furthermore, step S2 minimizes the system energy consumption over T time slots. This optimization problem is expressed as:

[0079]

[0080] Furthermore, in step S3, the explored experiences are stored in different experience buffer pools according to their quality. The buffer pool division method is as follows: find the threshold for dividing better and worse experiences, and divide the experience buffer pool into two parts, a better experience pool and a worse experience pool, based on the threshold.

[0081] Furthermore, in step S3, the minimization problem of step S2 is transformed into an MDP problem, defined by a tuple {Sc, Ac, Tc, Rc}, where Sc is the set of system states, Ac is the set of system actions, and Tc is the set of system actions. c ={p(s c′ |s c a c R is the set of transition probabilities. c :S c ×A c →R c Here, π is the reward function, and policy π is the mapping from Sc to Sc. The MDP problem is defined as follows:

[0082] State space: In time slot t, the set of system states is defined as the coordinates of the UAV and vehicle, the task request type of each vehicle, the cache state of the UAV and vehicle, and the information age of the first pending task of the vehicle.

[0083] Action space: In time slot t, the system's action space consists of the association between the vehicle and each cache node, whether to cache, and whether to update the existing cache.

[0084] Reward function: The reward function is set as the sum of the system's energy consumption and the penalty function.

[0085] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0086] Previous task unloading strategies often focused only on the computational unloading process, neglecting the freshness of cached data used by the computational tasks. This invention addresses this deficiency by considering not only the freshness of information during the computational unloading process but also the freshness of cached data used during computational tasks in relation to the freshness of vehicle state updates, thus comprehensively ensuring the freshness of vehicle state updates.

[0087] Regarding system energy consumption, compared with previous autonomous driving research, this invention considers green systems while taking into account the freshness of the computational results obtained by the vehicle and the energy consumption of the system. It addresses the problems of high safety and timeliness requirements, high vehicle mobility, high task density, rapid network topology changes and high system overhead in autonomous driving scenarios by combining vehicle edge computing, caching and drone technologies. It designs a strategy to minimize system energy consumption and ensure the timeliness of vehicle state updates by making decisions on cache updates, user associations and resource allocation in a dynamic environment.

[0088] This invention addresses the problem-solving requirements by employing a deep reinforcement learning strategy, specifically the Deep Deterministic Policy Gradient (DDPG) algorithm. However, this method suffers from a large state and action space, resulting in slow convergence. Therefore, this invention improves upon the traditional DDPG algorithm by establishing a threshold for distinguishing between good and bad experiences through extensive experimentation. This threshold is used to divide the experience pool into two parts: a good experience pool and a bad experience pool. During training, experiences are selected proportionally from each pool to train the neural network. Simulations demonstrate that the improved algorithm converges faster, yields a higher reward value after convergence, and exhibits less oscillation. Attached Figure Description

[0089] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0090] Figure 1 This invention provides a system model for an efficient caching and task unloading state update method in unmanned aerial vehicle-to-everything (UAV) assisted vehicle-to-everything (V2X) network. Detailed Implementation

[0091] This invention designs an efficient state update method for caching and task unloading in UAV-assisted vehicle networking. The system model is as follows: Figure 1As shown, consider a region within a city where the roads are two-lane, two-way traffic sections. A macro base station (MeNB) equipped with an MEC server provides full signal coverage for the entire road section. M unmanned aerial vehicles (UAVs) with computing capabilities are deployed above the roads. Each vehicle (equipped with an MEC server) periodically generates different types of computing tasks; there are N vehicles and W types of tasks. The density and speed of vehicles in the region are randomized to simulate real-world urban traffic conditions. Within this region, the MeNB collects information and makes various decisions. This embodiment considers a time-slot system.

[0092] 1. Cache update strategy

[0093] In this invention, the MEC server can only process corresponding tasks if the relevant content has already been cached on it. It is assumed that the MeNB caches all content and continuously updates it to ensure the content is up-to-date. The cache capacity for UAVs and vehicles is limited. The cache update process can be completed within one time slot. Consider T time slots.

[0094] definition and This indicates whether vehicle i and UAV j have cached content w in time slot t. as well as This indicates that there is a cache w, otherwise it is 0. The cache capacity of vehicles and UAVs is expressed as:

[0095]

[0096]

[0097] in, and Indicates the cache capacity of vehicle i and UAV j, l w This indicates the amount of data in content w.

[0098] definition and These are two binary variables, representing whether vehicle i and UAV j update or replace cache w in time slot t. as well as This indicates that cached data w will be updated or replaced in time slot t; otherwise, it is 0. Since the cache capacity of vehicles and UAVs is limited, outdated cached data will be updated to a newer version or replaced with other content. Obviously, if the vehicle or UAV has not cached w before time slot t, then the cached data w cannot be updated in time slot t. Therefore, we can conclude that:

[0099]

[0100] Similarly, for UAV j, we can obtain:

[0101]

[0102] In this invention, the timeliness of the calculation results is related to the timeliness of the cache model used to calculate the vehicle task, as well as the transmission and calculation process. and Let represent the timeliness of the cached content for vehicles and UAVs, respectively, i.e., the freshness of the cache. Assume that when content w is replaced by other types of cached data, the freshness is set to an infinite number I. Therefore, we can obtain:

[0103]

[0104]

[0105] 2. Calculate the unloading strategy

[0106] Once a vehicle generates a computing task, it sends a resource access request to the MeNB, the request including the {e} of task w. i (t), s w , z w Information such as driving direction, speed, and current location. i (t) represents the task type of vehicle i in time slot t, s w This indicates the size of the task, w, and z. w This indicates the CPU cycles required for the sub-task computation; after collecting this information, the MeNB makes a decision on the processing of vehicle i-tasks, using... Indicates whether the task is processed locally, using It can be expressed in the following ways:

[0107] 1) Local computation:

[0108] 2) Uninstall to UAV:

[0109] 3) Uninstall to MeNB:

[0110] 4) Not calculated for now:

[0111] Then we have:

[0112]

[0113] 2.1 Local Calculation:

[0114] use This represents the computing power of vehicle i, and the local computing latency of vehicle i is expressed as:

[0115]

[0116] Corresponding energy consumption:

[0117]

[0118] Weighting factor μ i The energy consumption required for vehicle i's CPU to perform calculations per cycle is:

[0119]

[0120] In this study, the latency of local computation is limited to no more than one time slot τ. Therefore:

[0121]

[0122] 2.2 Task Unloading

[0123] When tasks are offloaded, the MEC server in the UAV or MeNB performs the computational tasks.

[0124] (1) Unload to UAV:

[0125] When a vehicle transmits a computing task to a UAV, other UAVs will be affected by this transmission process. Therefore, the transmission rate between vehicle i and UAV j can be expressed as:

[0126]

[0127] Among them, b i,j (t) is the bandwidth when vehicle i communicates with UAV j and MeNB, p is the transmit power of the vehicle for the mission, and β is the bandwidth of the UAV j and MeNB for the communication. i,j (t) is the channel gain when vehicle i communicates with UAV j, σ 2 It is the power spectral density of white noise.

[0128] The channel gain between vehicle i and UAV j can be expressed as:

[0129]

[0130] Where, d i,j (t) represents the distance between vehicle i and UAV j. Establish a coordinate axis on this region, using... Let x and y coordinates represent the positions of the vehicle and UAV, and let x and y coordinates represent the speeds of the vehicle and UAV. Positive and negative signs indicate the direction of travel. The distance between vehicle i and UAV j is expressed as:

[0131]

[0132] If the task of vehicle i is offloaded to UAV j at time t, then i must be within the coverage area of ​​UAV j within one time slot. That is, the connection distance between vehicle i and UAV j must be greater than the communication radius R of the UAV at the start of the next time slot. Therefore:

[0133]

[0134] (2) Uninstall to MeNB:

[0135] Transmission rate between vehicle i and MeNB:

[0136]

[0137] Among them, b i,M+1 (t) is the bandwidth when vehicle i communicates with the MeNB, p is the transmit power when vehicle i transmits the task to the MeNB, and g i,M+1 (t) is the square of the average channel gain between vehicle i and MeNB.

[0138] The time it takes for vehicle i to transmit task w to UAV j or MeNB can be expressed as:

[0139]

[0140] The latency of the MEC server in UAVs and MeNBs calculating the task of vehicle i can be expressed as:

[0141]

[0142] in, This represents the computing resources allocated to MEC server j for vehicle i in time slot t. Therefore, the total unloading latency is:

[0143]

[0144] The unloading delay is limited to one time slot:

[0145]

[0146] Energy consumption during unloading:

[0147]

[0148] definition This represents the freshness of the state update for vehicle i in time slot t, i.e., the age of the state update. Let w represent the time when task w was generated; therefore, the age of the state update can be expressed as:

[0149]

[0150] The regulation stipulates that it cannot exceed threshold A. th :

[0151]

[0152] 3. System Objectives

[0153] System energy consumption comes from two aspects: updating cached content and processing vehicle tasks. The energy consumption ξ(t) caused by cache updates within time slot t is:

[0154]

[0155] Wherein, θ(J / bit) is a weighting factor that converts the amount of cached data to energy, representing the energy required to cache 1 bit of data.

[0156] Energy consumption E generated by task processing within time slot t i (t)):

[0157]

[0158] The energy consumption φ(t) of the system due to task transmission in time slot t is expressed as:

[0159]

[0160] In summary, the system's total energy consumption in time slot t is...

[0161]

[0162] The objective of this invention is to minimize the system energy consumption over T time slots while satisfying various constraints. This optimization problem can be formulated as:

[0163]

[0164] st status update age limit: (3)(4)(5)(6)(22)(23)

[0165] Cache limitations: (1)(2)(24)

[0166] Task processing restrictions: (7)-(21)(25)(26)(27)

[0167] 4. Improved Deep Deterministic Policy Gradient Algorithm

[0168] In the system objectives of this invention, and It is a discrete variable, bandwidth-related bi,j (t) is a continuous variable. This problem is complex, falls under the category of MINLP, and is difficult to handle. To address the challenges of continuous task generation, high mobility, and rapidly changing channel conditions in the vehicle-to-everything (V2X) environment, a fast decision-making algorithm based on deep reinforcement learning is proposed.

[0169] Therefore, this invention transforms the problem that the system objective aims to solve into an MDP problem, defined by a tuple {Sc, Ac, Tc, Rc}, where Sc is the set of system states, Ac is the set of system actions, and Tc is the set of system actions. c ={p(s c′ |s c a c )} is the set of transition probabilities, and R c :S c ×A c →R c It is the reward function. The policy π is a mapping from Sc to Sc, therefore the MDP problem is defined as follows:

[0170] (1) State space: In time slot t, the set of system states is defined as the coordinates of the UAV and vehicle, the task request type of each vehicle, the cache state of the UAV and vehicle, and the information age of the first pending task of the vehicle.

[0171] (2) Action space: In time slot t, the action space of the system is the association between the vehicle and each cache node, and whether to cache and update the existing cache.

[0172] (3) Reward function: The reward function is set as the sum of the system's energy consumption and the penalty function. The penalty function is to prevent the size of the cached data from exceeding the cache capacity, and to prevent the vehicle from staying within the UAV's communication range for more than one time slot if it is determined that the vehicle's task is offloaded to the UAV.

[0173] Due to the randomness of the system, state transition probabilities are difficult to model. Therefore, we employ a model-free reinforcement learning algorithm based on deep deterministic policy gradient (DDPG) to learn and update computational resource allocation strategies. Unlike the traditional DDPG algorithm, which randomly samples training data without considering data quality, in this invention, explored experiences are stored in different experience buffers according to their quality. Then, poor and good experiences are randomly selected proportionally within a certain number of steps to eliminate correlations between data, improve the stability of neural network training, and reduce oscillations.

[0174] In the method of this invention, firstly, to comprehensively ensure the safety of autonomous driving, this invention simultaneously considers the freshness of the cache model used for calculation and the freshness of the calculation unloading process as the information freshness of the calculation results obtained by the vehicle. Compared with previous research, this consideration is more practically significant.

[0175] Secondly, in autonomous driving scenarios, vehicles need to frequently interact with the external environment to perceive it. Therefore, this invention mainly addresses the challenges of cache-assisted edge computing architectures, high requirements for safety and timeliness, high vehicle mobility, intensive tasks, rapid network topology changes, and high system overhead in autonomous driving scenarios. By combining vehicle edge computing, caching, and drone technologies, a strategy is designed to minimize system energy consumption and ensure information timeliness by making decisions on cache updates, user associations, and resource allocation in a dynamic environment.

[0176] Finally, the problem exists in the form of mixed-integer nonlinear programming (MINLP), which is a challenging problem. Therefore, this invention uses a deep reinforcement learning algorithm, employing an algorithm based on deep deterministic policy gradients, to divide the experience buffer pool and select experience from two different buffer pools proportionally to train the neural network. This accelerates the convergence speed, reduces reward oscillations after convergence, increases the reward value after convergence, and improves the algorithm performance.

[0177] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. However, these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for efficient caching and task unloading state updates in a drone-assisted vehicle network, characterized in that, The system model includes a macro base station equipped with an MEC server, a drone with computing capabilities, and multiple vehicles equipped with computing capabilities, and employs the following steps: S1. Calculate the freshness of the cache model and the freshness of the unloading process as the freshness of vehicle state updates; the calculation method for the freshness of the cache model is as follows: definition and This indicates whether vehicle i and UAV j have cached content w in time slot t. as well as This indicates that there is a cache w, otherwise it is 0. The cache capacity of vehicles and UAVs is expressed as: in, and Indicates the buffer capacity of vehicle i and UAV j, l w This indicates the amount of data in content w; definition and These are two binary variables, representing whether vehicle i and UAV j update or replace cache w in time slot t. as well as This indicates that cache w is updated or replaced in time slot t, otherwise it is 0; for vehicle i: For UAV j: and Let represent the timeliness of the cached content for vehicles and UAVs, respectively, i.e., the freshness of the cache. Assuming that the time it takes for content w to be replaced by other types of cached data is set to an infinite number I, we get: The method for calculating freshness during the unloading process is as follows: The vehicle generates a computational task and sends a resource access request to the MeNB, the request including the {e} of task w. i (t),s w ,z w }, driving direction, driving speed, current location information, where e i (t) represents the task type of vehicle i in time slot t, s w This indicates the size of the task, w, and z. w This indicates the CPU cycles required for the computation of subtask w; after collecting this information, the MeNB makes a decision on the processing of vehicle task i, using... Indicates whether the task is processed locally, using It can be expressed in the following ways: 1) Local computation: 2) Uninstall to UAV: 3) Uninstall to MeNB: 4) Not calculated for now: Then we have: S2. Taking into account both the freshness of the cache model and the freshness of the unloading process, the optimization problem is to minimize the total energy consumption of the system in time slot t. The total energy consumption includes the energy consumption generated by updating the cached content in time slot t and the energy consumption generated by vehicle task processing in time slot t. S3. The minimization problem in step S2 is transformed into an MDP problem. A model-free reinforcement learning algorithm based on deep deterministic policy gradient is used to learn and update the computational resource allocation strategy. When training data, the algorithm divides the experience cache pool according to the different experience quality and selects experience from the two different cache pools in proportion to train the neural network.

2. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, When calculating locally, use This represents the computing power of vehicle i, and the local computing latency of vehicle i is expressed as: The corresponding energy consumption is expressed as follows: Weighting factor μ i The energy consumption required for vehicle i's CPU to perform calculations per cycle is represented as: The local computation delay does not exceed one time slot τ, and we have:

3. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, When unloading to a UAV, the transmission rate between vehicle i and UAV j is expressed as: Among them, b i,j (t) is the bandwidth when vehicle i communicates with UAV j and MeNB, p is the vehicle's mission transmit power, and β is the bandwidth. i,j (t) is the channel gain when vehicle i communicates with UAV j, σ 2 It is the power spectral density of white noise; The channel gain between vehicle i and UAV j is expressed as: Where, d i,j (t) represents the distance between vehicle i and UAV j; Establish a coordinate axis on this area, using Let x and y coordinates represent the positions of the vehicle and UAV, and let x and y coordinates represent the speeds of the vehicle and UAV. Positive and negative signs indicate the direction of travel. The distance between vehicle i and UAV j is expressed as: If the task of vehicle i is offloaded to UAV j at time t, then i must be within the coverage area of ​​UAV j within one time slot. That is, the connection distance between vehicle i and UAV j must be greater than the communication radius R of the UAV at the start of the next time slot. Therefore: R represents the communication coverage radius of the UAV.

4. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, When offloading to the MeNB, the transmission rate between vehicle i and the MeNB is as follows: Among them, b i,M+1 (t) is the bandwidth when vehicle i communicates with the MeNB, g i,M+1 (t) is the square of the average channel gain between vehicle i and MeNB, g i,M+1 (t) is the channel gain between vehicle i and MeNB: The time it takes for vehicle i to transmit task w to UAV j or MeNB is represented as: The latency of the MEC server in the UAV and MeNB calculating the task of vehicle i is expressed as: in, This represents the computing resources allocated to vehicle i in MEC server j during time slot t; The total uninstallation time is: The unloading delay is limited to one time slot: Energy consumption during unloading: definition The freshness of the state update of vehicle i in time slot t is represented by, i.e., the age of the state update, using... The age of the state update represents the time when task w was generated. The regulation stipulates that it cannot exceed threshold A. th :

5. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, The energy consumption ξ(t) generated by cache updates within time slot t is expressed as: Where θ(J / bit) is a weighting factor that converts the amount of cached data to energy, representing the energy required to cache 1 bit of data; Energy consumption E generated by task processing within time slot t i (t) is represented as: The energy consumption φ(t) of the system due to task transmission in time slot t is expressed as: The system's total energy consumption in time slot t Represented as:

6. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, Step S2 minimizes the system energy consumption over T time slots. This optimization problem is expressed as:

7. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, In step S3, the explored experiences are stored in different experience buffer pools according to their quality. The buffer pool division method is as follows: find the threshold for dividing better and worse experiences, and divide the experience buffer pool into two parts, a better experience pool and a worse experience pool, based on the threshold.

8. The method for efficient caching and task unloading state update in UAV-assisted vehicle networking according to claim 1, characterized in that, In step S3, the minimization problem in step S2 is transformed into an MDP problem, defined by a tuple {Sc, Ac, Tc, Rc}, where Sc is the set of system states, Ac is the set of system actions, and Tc is the set of system actions. c ={p(s c′ |s c ,a c R is the set of transition probabilities. c :S c ×A c →R c Here, π is the reward function, and policy π is the mapping from Sc to Sc. The MDP problem is defined as follows: State space: In time slot t, the set of system states is defined as the coordinates of the UAV and vehicle, the task request type of each vehicle, the cache state of the UAV and vehicle, and the information age of the first pending task of the vehicle. Action space: In time slot t, the system's action space consists of the association between the vehicle and each cache node, whether to cache, and whether to update the existing cache. Reward function: The reward function is set as the sum of the system's energy consumption and the penalty function.

Citation Information

Patent Citations

  • Collaborative edge caching method based on deep reinforcement learning in Internet of Vehicles

    CN113012013A

  • Cache auxiliary task cooperative unloading and resource allocation method based on meta reinforcement learning

    CN113434212A