Cooperative cache decision-making method based on user preference under assistance of unmanned aerial vehicle in Internet of Vehicles
By introducing drones as edge cache nodes into the vehicle edge caching system, collaborative caching decisions based on user preferences, and optimizing content cache placement using the D3QN algorithm, the problem of limited resources in the vehicle edge caching system is solved, and the efficiency of content retrieval and service response rate are improved.
Patent Information
- Application Number
- CN202511173250.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Vehicle edge caching systems have limited resources. How to effectively manage and optimize these resources and formulate reasonable caching strategies to improve content retrieval efficiency and reduce data access latency and network load is a key challenge.
By introducing drones as edge caching nodes to assist in user-preference-based collaborative caching decisions in the Internet of Vehicles, an edge collaborative caching architecture model is constructed. The D3QN algorithm is used for deep learning to solve the problem, calculate the cache value and transmission rate, and optimize the content caching placement strategy.
It improves the response rate of vehicle user services, saves traffic consumption, reduces the load on the central server, and provides low-latency, high-availability services.
Smart Images

Figure CN120897233A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology, and specifically relates to a collaborative caching decision-making method based on user preferences in the Internet of Vehicles (IoV) with the assistance of unmanned aerial vehicles (UAVs). Background Technology
[0002] With the advent of the era of the Internet of Things (IoT), the Internet of Vehicles (IoV), as a key branch of IoT, has become an indispensable component of intelligent transportation systems. Autonomous driving, as an in-vehicle application, is highly sensitive to computation and latency, requiring high levels of data communication and computing storage. However, current in-vehicle network systems cannot meet the ever-increasing latency requirements of in-vehicle applications. In-vehicle networks based on Mobile Edge Computing (MEC) have been proposed as a feasible solution.
[0003] To improve the efficiency and Quality of Service (QoS) of connected vehicle systems, Vehicle Edge Caching (VEC) technology has emerged. By caching data on vehicles and edge nodes, data access latency and network load can be significantly reduced. The development of Mobile Edge Computing (MEC) provides mobile network operators with the opportunity to deeply deploy Content Message Networks (CDNs) into their mobile network infrastructure. MEC technology effectively allows CDN edge nodes to be deployed closer to vehicle users, enabling them to access services faster. This technology not only improves content retrieval speed but also reduces the pressure on central network nodes. However, vehicle edge caching systems have limited resources; therefore, effectively managing and optimizing these resources and developing reasonable caching strategies have become key research challenges.
[0004] By adding more cache nodes among vehicle users, drones, and roadside infrastructure, and leveraging distributed caching resources, the efficiency of content retrieval is improved. The combination of edge computing and caching not only reduces the load on central servers but also provides users with low-latency, high-availability services. Summary of the Invention
[0005] This invention proposes a collaborative caching decision-making method based on user preferences in the Internet of Vehicles (IoV) with the assistance of drones. The drones are used as edge caching nodes to assist in caching decisions for content storage based on user preferences in the IoV. This method calculates the corresponding caching value of the requested content for the surrounding caching nodes from the perspective of the vehicle user requesting the content, and finds a reasonable caching location for the requested content.
[0006] The technical solution of this invention is as follows:
[0007] A collaborative caching decision-making method based on user preferences with drone assistance in the Internet of Vehicles (IoV) includes the following steps:
[0008] Step 1: Construct an edge collaborative caching architecture model in the Internet of Vehicles (IoV) consisting of vehicle users, drones, and macro base stations;
[0009] Step 2: Calculate the cache value of the corresponding content using the vehicle user's historical request content topics and the current request content topics;
[0010] Step 3: Calculate the cache value under different conditions and the transmission rate under different types of links, and construct a latency model and energy consumption model for the content download request process;
[0011] Step 4: Identify the optimization problem and use the D3QN algorithm for deep learning to solve it, providing a request cache placement and replacement strategy that maximizes the caching benefits for the requested content.
[0012] Furthermore, in step 1, the vehicle user set is defined as follows: It caches its own preferred content locally and shares the cached content with users in nearby vehicles via vehicle-to-vehicle links. m Let m be the m-th vehicle user, and V be the total number of vehicle users; define the set of drones as... u n Let U be the nth drone, and U be the total number of drones. Each drone is equipped with a cache server to assist the macro base station in providing content caching services to vehicle users. They share and transmit content among themselves, providing service to vehicle users only during peak hours. During off-peak hours, service is transferred to the macro base station to reduce drone energy consumption. In each time slot, the set of content requests randomly issued by vehicle users is Q = {f...} 1,1 ,...,f 1,k ,...,f m,1 ,...,f m,k}, f m,k For the k-th content requested by the m-th vehicle user; define the content library as... f k Let F be the k-th content, and F be the total number of content.
[0013] Furthermore, the specific process of step 2 is as follows:
[0014] Step 2.1: Calculate the probability of a vehicle user's request based on the request sent by the vehicle user. The formula is as follows:
[0015]
[0016] Where, p f ρ is the probability of requesting content f; ρ is the Zipf exponent.
[0017] Step 2.2: Calculate the content caching value using the following formula:
[0018]
[0019] in, ω1 and ω2 are the cache values of content f in time slot t at different node levels; ω1 and ω2 are different adjustment parameters. This represents the m-th vehicle user v. m For the k-th content f of the current request k Interests and preferences;
[0020] Interests and preferences are closely related to content topics. The set of topics is defined as W = {w1, w2, ..., w...} j ,...,w J}, w j Let j represent the j-th topic, where J is the total number of topics; define a binary variable e. k (f k ,w j ) represents the k-th content f k and the j-th topic w j The correlation;
[0021] The cache value of the current request is determined based on the topics of the vehicle user's historical requests, whether it's a spur-of-the-moment decision or due to content preference; historical average request information is defined based on mean cosine theory. The formula is:
[0022]
[0023] Among them, f a f b These are two different contents from the vehicle user's historical requests; K is the number of contents in the historical requests; e(f a ,w j ) is f a with w j The correlation; e(f) b ,w j ) is f b with w j The correlation;
[0024] Determine the m-th vehicle user v based on the topic of the request content. m The requested content f is at this point. k and historical average request information The relevance is used to represent the relationship between the m-th vehicle user and the k-th content f requested in the current request. k Degree of interest preference:
[0025]
[0026] Furthermore, the specific process of step 3 is as follows:
[0027] Step 3.1: Construct formulas for calculating cache value under different conditions;
[0028] Step 3.2: Calculate the transmission rate under different types of links;
[0029] Step 3.3: Construct a latency model for the content download request process;
[0030] Step 3.4: Construct an energy consumption model for the content download request process.
[0031] Furthermore, step 3.1 includes three scenarios;
[0032] Scenario 1: When vehicle user v caches content f, the cache value of time slot t is
[0033]
[0034] Wherein, ω1 and ω2 are different adjustment parameters. In case 1, ω1 = 1 and ω2 = 0.
[0035] Scenario 2: For content f requested by drone u within the coverage area, the cache value of time slot t is
[0036]
[0037] In case 2, ω1 = 0 and ω2 = 1;
[0038] Scenario 3: In the content f request, the cache of adjacent vehicle user v' within the vehicle's communication coverage area is [value missing]. The cache value of time slot t is [value missing].
[0039]
[0040] Where α is the moderating factor for social contribution participation; SC m The social contribution rate of the m-th vehicle user.
[0041] Furthermore, in step 3.2, the vehicle user, as both the content requester and provider, will complete the request immediately if the content exists in their own cache, without needing to establish a communication link; otherwise, the vehicle user will obtain the content from four types of links based on the various content cache locations, and calculate the transmission rate under the four types of links.
[0042] Type 1: Obtain data from nearby adjacent vehicle users and calculate the transmission rate between vehicle users; the formula is:
[0043]
[0044] Where v and v' are two adjacent vehicle users; Rv,v' B represents the transmission rate of shared content between vehicle user v and its adjacent vehicle user v'. v,v' P is the transmission bandwidth between vehicle user v and its adjacent vehicle user v'; v,v' N0 is the transmit power between vehicle user v and its neighboring vehicle user v'; N0 is the noise power; g v,v' It is the channel gain between vehicle user v and adjacent vehicle user v';
[0045] Type 2: Obtain data from drones above the activity area and calculate the transmission rate from vehicle users to the drones; the formula is:
[0046]
[0047] Among them, R u,v B represents the transmission rate between vehicle user v and drone u. u,v This represents the transmission bandwidth between the drone u and the vehicle user v; It is the transmit power of the drone u to the vehicle user v in time slot t; h v,u σ represents the air-to-ground channel gain between vehicle user v and drone u. 2 It is the variance of Gaussian noise;
[0048] Type 3: Acquire data in cooperation with neighboring drones, and calculate the inter-drone transmission rate; the formula is:
[0049]
[0050] Where u and u' are two adjacent drones; R u,u' B represents the transmission rate between drone u and its neighboring drone u'. u,u' P represents the transmission bandwidth between drone u and its neighboring drone u'; u,u' The transmission power between UAV u and its neighboring UAV u'; The position of UAV u in time slot t; Let g0 be the position of the adjacent UAV u' in time slot t; g0 represents the power gain when the reference distance is 1 meter.
[0051] Type 4: Obtain from macro base stations and calculate the transmission rate between vehicle users and macro base stations; the formula is:
[0052]
[0053] Among them, R v,b B represents the transmission rate between vehicle user v and macro base station b. b,v P represents the transmission bandwidth between vehicle user v and macro base station b. v,b The transmit power between vehicle user v and macro base station b; g v,bLet v be the channel gain between vehicle user v and macro base station b.
[0054] Furthermore, in step 3.3, the working process of the content download request process latency model is as follows:
[0055] Step 3.3.1: When a vehicle user sends a download request for content f, it first checks whether the content is found in the vehicle user's own cache. If the content is found in the vehicle user's cache, the vehicle user does not need to make a cache decision and directly outputs f with zero latency. The vehicle user's own cache hit decision is then confirmed. Increase by 1; otherwise, proceed to step 3.3.2 for judgment.
[0056] Step 3.3.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to directly download content f from neighboring vehicle users; if not, proceed to step 3.3.3 for judgment; wherein, obtaining content f from neighboring vehicle users... k The required time is:
[0057]
[0058] in, The size of the k-th content requested in time slot t; For vehicle user v and adjacent vehicle user v', the k-th content f k Transmission delay;
[0059] Step 3.3.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. If the value is increased by 1, proceed to step 3.3.4 for judgment; in addition, the waiting time for the vehicle user's request task is calculated by the drone. If it does not exceed the maximum tolerable delay, it enters the buffer queue to wait for the drone to transmit to the vehicle user; otherwise, it is forwarded to the macro base station; drone u and vehicle user v are the service content f. k The transmission delay is
[0060]
[0061] Among them, s k t-1 , These represent the sizes of the kth content requested before time slot t-1 and t, respectively. The average transmission rate between drones and vehicle users;
[0062] Step 3.3.4: If the requested content is not found in the cache list of the adjacent drone u, determine whether the requested content can be found in the cache list of the adjacent drone u'. If so, a cache hit decision is made for the adjacent drone u'. Increase by 1, u' transmits the content to u via content sharing; otherwise, proceed to step 3.3.5 for calculation; for newly arrived content, drone u acts as a relay for forwarding; calculate the content transmission delay:
[0063]
[0064] in, For drone u and neighboring drone u', the service content for vehicle users f k Transmission delay; This represents the average transmission rate between drones.
[0065] Step 3.3.5: Calculate the k-th content f k Time required for data to be transmitted from the macro base station to the vehicle user:
[0066]
[0067] in, For macro base station b and vehicle user v, the k-th content f k Transmission delay; R v,b The transmission rate between vehicle user v and macro base station b.
[0068] Furthermore, in step 3.4, the calculation process of the energy consumption model for the content download request process is as follows:
[0069] Step 3.4.1: The vehicle user sends a content request f. k First, determine if the content can be found in vehicle user v. If so, then make a cache hit decision for the vehicle user itself. Increase by 1, the vehicle directly outputs f k If the energy consumption is 0, proceed to step 3.4.2 for judgment;
[0070] Step 3.4.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to directly download content f from neighboring vehicle users. k If not, proceed to step 3.4.3 for judgment;
[0071] Content f between vehicle user v and adjacent vehicle user v' k Transmission energy consumption for:
[0072]
[0073] Step 3.4.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. Increase by 1; otherwise, proceed to step 3.4.4 for judgment.
[0074] Content f between vehicle user v and drone u k The transmission energy consumption is
[0075]
[0076] Among them, P u,v The transmit power from the drone (u) to the vehicle user (v); For content f between drone u and vehicle user v k Transmission delay; P u The transmit power of the UAV u;
[0077] Step 3.4.4: Determine whether the requested content can be found in the cache list of neighboring drones u'. If so, make a cache hit decision for neighboring drones u'. Increase by 1; otherwise, proceed to step 3.4.5 of the calculation.
[0078] The transmission energy consumption between vehicle user v and adjacent drone u' is
[0079]
[0080] in, For the content f between drone u and its neighboring drone u' k Transmission energy consumption;
[0081] Steps 3, 4, and 5, Content f k The energy consumption for transmission via macro base stations is:
[0082]
[0083] in, For content f between macro base station b and vehicle user v k Transmission energy consumption.
[0084] Furthermore, in step 4, the final optimization problem P is determined to be the function of maximizing cache efficiency:
[0085]
[0086] in, Indicates time slot t, vehicle user v, and content f. k Cache status; Indicates time slot t, UAV u, and content f k The caching status; T represents a time period; C1, C2, C3, C4, and C5 represent different constraints; For the content f of the Mth node in time slot t k Cache status; s k Indicates the size of the k-th element; δ fk For content f k Maximum tolerable delay; U cache The caching efficiency function is expressed by the following formula:
[0087] U cache =(U time +U ene )p f (twenty one);
[0088] Among them, U time For time efficiency; U ene For energy consumption cost-effectiveness; the formulas are as follows:
[0089]
[0090] Among them, R v2v For the transmission rate between vehicle users; For vehicle user v, content f is received in each time slot t via four types of links. k The total time is calculated using the following formula:
[0091]
[0092] in, For the content f of adjacent UAVs in time slot t. k Cache status;
[0093] For vehicle user v, content f is received in each time slot t via four types of links. k The system energy consumption cost is given by the formula:
[0094]
[0095] in, These represent the energy consumption cost of transmitting content through different link units.
[0096] Furthermore, in step 4, deep reinforcement learning based on D3QN is used to select the most appropriate cache placement location for the requested content. In the D3QN algorithm, the agent obtains a quadruple of the action taken in the current state, the reward after taking the action, the next state, and the current state through interaction with the environment. in, Indicates the state of time slot t, a t Indicates the action taken in time slot t, r t This represents the reward received in time slot t. This represents the state of time slot t+1; these quadruplets are stored in the experience replay buffer as the basis for learning; when the experience replay buffer is full, the D3QN algorithm will delete the oldest quadruplet sample.
[0097] Each quadruple in the experience replay buffer is used to train the D3QN algorithm; the goal of the D3QN algorithm is to optimize the evaluation network and action network by calculating a target value; the formula for calculating the target value is:
[0098]
[0099] in, The target value for time slot t; Q target It is the Q-value function of the target network; It is the discount factor; 'a' represents the action; θ represents the discount factor. target This represents the parameters of the target network; the D3QN algorithm uses the target network to compute the Q-value and reduces overestimation of the Q-value.
[0100] Using gradient descent, the D3QN algorithm updates the parameters of the evaluation network based on the calculated target value; the evaluation network parameters θ eval The update formula is:
[0101]
[0102] Where ← represents the update operation; It is about finding the gradient; It evaluates the network's learning rate; Q eval It is a function for evaluating the Q-value of a network;
[0103] The D3QN algorithm updates the parameters of the action network through policy gradient ascent, aiming to maximize the agent's reward; the action network parameters θ action The update formula is:
[0104]
[0105] Where β is the learning rate of the action network;
[0106] After each training iteration, the parameters of the target network are updated based on the evaluation network; the parameters θ of the target network are... target The update formula is:
[0107] θ target ←τθ eval +(1-τ)θ target (29);
[0108] Where τ is the soft update coefficient;
[0109] The D3QN algorithm learns iteratively and updates the evaluation network and action network to obtain a high-performance value function, thereby obtaining the optimal cache placement scheme; finally, it optimizes the content caching decision based on the cache placement scheme.
[0110] The beneficial technical effects of this invention are as follows: From the perspective of reactive caching, this invention introduces drones as edge caching nodes to assist in caching the request content of mobile vehicle users in the Internet of Vehicles, finds a reasonable caching location for the request content, thereby improving the response rate of vehicle user services and saving traffic consumption. Attached Figure Description
[0111] Figure 1 This is a flowchart of the collaborative caching decision-making method based on user preferences in the Internet of Vehicles (IoV) with the assistance of unmanned aerial vehicles (UAVs) in this invention.
[0112] Figure 2 This is the joint edge caching architecture model in the Internet of Vehicles in this invention.
[0113] Figure 3 This is a flowchart of the cache replacement strategy based on user preferences in this invention. Detailed Implementation
[0114] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0115] like Figure 1The diagram shows the method block diagram of this invention, which includes the following three processes: constructing a joint edge caching architecture model in a vehicle network composed of vehicle users, drones, and macro base stations; calculating the caching value of corresponding content using the attributes of users' historical request content and regional features; and calculating a request caching placement strategy that maximizes caching benefits based on the attributes of the requested content and the corresponding user preferences, and proving that it is NP-hard by mapping it to a multiple knapsack problem, thereby using deep reinforcement learning D3QN to solve it. Specifically, it involves constructing a collaborative edge caching system composed of macro base stations, drones equipped with edge servers, mobile vehicle users, and backhaul links connected to the Internet and remote servers. Drones, macro base stations, and vehicle users serve as three caching nodes; calculating user preference features based on users' historical request information; calculating the degree of matching between the content's feature attributes and user preferences when requesting new content, thereby calculating the corresponding caching value; calculating the time efficiency function and energy efficiency function of caching content at different caching nodes to construct a caching benefit function; establishing an optimization function with the popularity of requested content; using the D3QN algorithm for simulation; iteratively training the D3QN network to find the caching placement strategy that maximizes the Q value.
[0116] The present invention specifically includes the following steps:
[0117] Step 1: Construct an edge collaborative caching architecture model in the Internet of Vehicles (IoV) consisting of vehicle users, drones, and macro base stations.
[0118] Figure 2 This is a schematic diagram of the edge collaborative caching architecture model in the UAV-assisted vehicle network of the present invention. It constructs a collaborative edge caching system consisting of a macro base station, UAVs equipped with edge servers, mobile vehicle users, and backhaul links connected to the Internet and remote servers. The UAV, macro base station, and vehicle users act as three caching nodes. User preference characteristics are calculated based on historical user request information. When new content is requested, the matching degree between the content's characteristic attributes and the user's preferences is calculated, thereby determining the corresponding caching value. Content is cached sequentially at different caching nodes. Vehicle users can find content themselves or obtain content from neighboring vehicle users within their communication range. The UAV can cache popular content within its coverage area. The macro base station is connected to a cloud server, providing services to all vehicle users within its coverage area and acting as a central controller to manage all edge servers and synchronously store information. The content in the cloud server is provided by content providers and updated in real time.
[0119] Consider a vehicle edge computing network with three different types of edge caching nodes: macro base stations (MBS), unmanned aerial vehicles (UAVs), and vehicle users. The set of vehicle users is represented as... It caches its own preferred content locally and can share the cached content with users in nearby vehicles via vehicle-to-vehicle (V2V) links. m Let m be the m-th vehicle user, and V be the total number of vehicle users. The set of drones is represented as... u n Let U be the nth drone, and U be the total number of drones. Each drone is equipped with a cache server to assist the macro base station in providing content caching services to vehicle users. They can share and transmit content among themselves, providing service to ground vehicle users only during peak hours. During off-peak hours, service is transferred to the macro base station to reduce drone energy consumption. Let u be the number of the nth drone. n The set of vehicle users within the coverage area is Let R be the r-th vehicle user within the coverage area of the n-th drone. In each time slot t, the set of content requests randomly issued by the vehicle user is Q = {f}. 1,1 ,...,f 1,k ,...,f m,1 ,...,f m,k}, f m,k This refers to the k-th content requested by the m-th vehicle user.
[0120] Content library by It means that f k Let F be the k-th content and F be the total number of content; define the vehicle request task characteristic set as P = {s} k ,δ k ,w j}, where s k δ represents the size of the k-th element. k w represents the vehicle user's tolerance for the delay of the k-th content task. j Let represent the j-th topic. Assume each drone is deployed within a sub-region, providing edge storage services to vehicle users, and that there is no overlap between sub-regions. The location of drone u in time slot t is... in This indicates the horizontal position of the UAV u projected onto the ground in time slot t. Let x and y be the x and y coordinates of drone u in time slot t, respectively; h is the altitude of the drone, assuming all drones fly at a fixed altitude. The position of vehicle user v in time slot t is represented as follows: Let x and y be the x and y coordinates of vehicle user v in time slot t, respectively. Assume that time period T is divided into several time slots. A content request issued by a vehicle user follows a Poisson distribution.
[0121] Step 2: Calculate the cache value of the corresponding content using the vehicle user's historical request topics and the current request topic. The specific process is as follows:
[0122] Step 2.1: Calculate the probability of a vehicle user's request based on the request sent by the vehicle user;
[0123] Let the nth drone u be in the current time slot. n The set of vehicle users within the coverage area is During the content acquisition phase, vehicle users' request behavior is influenced by factors such as user preferences and content popularity; vehicle users will randomly request popular content. The frequency distribution set is defined as follows: The set contains the frequency distribution of requests for each piece of content in the content library by vehicle users; where p f This represents the request probability of content f, following a Zipf distribution, and is used here as the global popularity within the region. For content libraries... For content f, the probability of a user requesting it is:
[0124]
[0125] Here, ρ is the Zipf index, indicating the skewness of popularity. The larger the ρ, the more concentrated the popular content is; conversely, the smaller the ρ, the more evenly the popular content is distributed. As a normalization factor, it ensures that the sum of the probabilities of all content is 1;
[0126] Step 2.2: Calculate the value of content caching;
[0127] Due to limited cache space, to maximize its caching efficiency, and considering both global and individual popularity, the caching value of content f in time slot t is defined as follows:
[0128]
[0129] Here, ω1 and ω2 are different adjustment parameters, representing their respective weights. They are adjusted based on feedback information from vehicle users and drones, and there are different caching decision schemes for vehicle users and drones. This represents the m-th vehicle user v. m For the k-th content f in the current request k User interests and preferences are closely related to content theme characteristics;
[0130] Define the set of topics as W = {w1, w2, ..., w...} j ,...,w J}, w j Let j represent the j-th topic, where J is the total number of topics; define a binary variable e. k (f k ,w j ) represents the k-th content f k and the j-th topic w j The correlation, e(f)k ,w j ) equal to 1 means f k and w j Relevant; a value of 0 indicates no relevance.
[0131] Based on the vehicle user's historical request information, determine the cache value of the current requested content, whether it's a spur-of-the-moment decision or due to content preference; define the historical average request information based on mean cosine theory. This represents the m-th vehicle user v. m The average topic similarity among K pieces of content from historical requests is calculated using the following formula:
[0132]
[0133] Among them, f a f b The content requested by different vehicle users; w j Represents the j-th topic; e(f a ,w j ) is f a with w j The correlation; e(f) b ,w j ) is f b with w j The correlation;
[0134] Determine the m-th vehicle user v based on the topic of the request content. m The requested content f is at this point. k and historical average request information The relevance is used to represent the relationship between the m-th vehicle user and the k-th content f requested in the current request. k Degree of interest preference:
[0135]
[0136] Step 2.3: Construct the caching model;
[0137] The caching status of each node's cache list is determined by a matrix. Represented as:
[0138]
[0139] Where M represents the total number of nodes for all vehicle users and drones; f k ' represents the k-th item in the cache list; using a binary decision variable. Indicates the Mth node pair f k The caching status of ' An equal value of 1 indicates that f k ' is cached by the Mth node; An equal value of 0 indicates that f k Not cached by the Mth node.
[0140] Assuming each piece of content is cached in only one place, and considering the cache size constraint, we have:
[0141]
[0142] Among them, C u Indicates the buffering capacity of the drone u; C v This represents the caching capacity of vehicle user v; C M This represents the maximum cache resource limit for all vehicle users and drone nodes. In each time slot, edge nodes can add newly arriving request content to their cache or replace / update older content.
[0143] The replacement principle is determined based on the value of the cache node where content f is located in the current time slot:
[0144]
[0145] in, For the Mth node in time slot t, f k The caching status of '; For time slot t to f k The cache value of f; k " represents the k'th item in the cache list; For time slot t-1, pair f k The cache value of "". The caching decision conditions are set. This represents the cache value of content f in time slot t at different node levels; μ is the attribution factor. Indicates the size of the content f requested in time slot t; according to Sort the contents of the cache in descending order.
[0146] Step 3: Cache value based on request content. The specific process is as follows:
[0147] Step 3.1: Calculate the cache value under different conditions based on the matching features between cache nodes and content; the cache value calculation is specifically divided into the following three cases:
[0148] Scenario 1: When vehicle user v caches content f, the cache value of time slot t is
[0149]
[0150] Where ω1 and ω2 are different adjustment parameters, ω1 = 1 and ω2 = 0 when the vehicle itself has buffered its contents. Sort the cached content from highest to lowest.
[0151] Scenario 2: For content f requested by drone u within the coverage area, the cache value of time slot t is
[0152]
[0153] Where ω1=0, ω2=1;
[0154] Scenario 3: In the content f request, the cache of adjacent vehicle user v' within the vehicle's communication coverage area is [value missing]. The cache value of time slot t is [value missing].
[0155]
[0156] Where α is the moderating factor for social contribution participation; SC m Let the social contribution rate of the m-th vehicle user be defined as the ratio of the content effectively contributed to its cache space usage. This represents the socio-economic benefits generated by the social cache node resources it occupies. When a cached vehicle successfully shares content with other vehicles within its communication range, its contribution rate increases accordingly, as shown in the formula:
[0157]
[0158] Where Θ represents the contribution rate of the increase; s k This indicates the size of the k-th element; For time slot t, the Mth node has the kth content f. k Cache status;
[0159] Use h v2v If we consider shared decision-making among vehicle users, then the average social contribution rate of each vehicle user can be calculated using the following formula:
[0160]
[0161] in, The average social contribution rate of the m-th vehicle user;
[0162] On one hand, when a vehicle user requests content that is cached in nearby vehicles, these cache holders will increase their corresponding social contribution value. On the other hand, when cache nodes receive requests from vehicle users simultaneously, the drone will queue and prioritize the requesters based on the contribution value of each cache holder, prioritizing services for vehicle users with higher social contribution rates. This incentivizes vehicle users to provide services to others in order to be better served. Furthermore, considering the impact of tolerable latency and packet loss rate, it is assumed that vehicle-to-vehicle content transmission will not exceed two hops.
[0163] Step 3.2: Calculate the transmission rate under different types of links; the specific process is as follows:
[0164] As content requesters and providers, vehicle users can request content immediately if it exists in their own cache, without needing to establish a communication link. Otherwise, depending on the various content cache locations in the system, vehicle users can obtain content from four types of links: (1) from nearby neighboring vehicle users, (2) from drones above their activity area, (3) from in cooperation with neighboring drones, and (4) from macro base stations.
[0165] Calculate the transmission rate under four types of links:
[0166] Type 1: Transmission rate between vehicle users;
[0167] According to Shannon's theory, the transmission rate of shared content between vehicle users is:
[0168]
[0169] Where v and v' are two adjacent vehicle users; R v,v' B represents the transmission rate of shared content between vehicle user v and its adjacent vehicle user v'. v,v' P is the transmission bandwidth between vehicle user v and its adjacent vehicle user v'; v,v' N0 is the transmit power between vehicle user v and its neighboring vehicle user v'; N0 is the noise power; g v,v' It is the channel gain between vehicle user v and its neighboring vehicle user v', and the formula is:
[0170]
[0171] Where g0 represents the power gain at a reference distance of 1 meter. This represents the distance between vehicle user v and its adjacent vehicle user v'.
[0172] use This represents the decision variable for sharing content f between vehicle user v and its neighboring vehicle user v', where γ is the signal-to-noise ratio of the communication link between vehicle user v and its neighboring vehicle user v'. v,v' Greater than the first signal-to-noise ratio threshold η veh At this time, a communication link is formed. The formula is:
[0173]
[0174] Type 2: Transmission rate from vehicle user to drone:
[0175] Based on the position coordinates between the drone u and the vehicle user in time slot t, the distance between the drone u and the vehicle user v can be represented as l. u,v:
[0176]
[0177] in, The position of UAV u in time slot t; The location of vehicle user v in time slot t;
[0178] The air-to-ground channel gain h between vehicle user v and drone u v,u This can be represented by a free-space path loss model:
[0179]
[0180] The transmission rate R between vehicle user v and drone u is then... u,v for:
[0181]
[0182] Among them, B u,v This represents the transmission bandwidth between the drone u and the vehicle user v, with the bandwidth being evenly distributed among the associated vehicle users. σ is the transmit power of the drone u to the vehicle user v in time slot t; 2 It is the variance of Gaussian noise.
[0183] Type 3: Inter-UAV transmission rate:
[0184] The transmission rate between drones is:
[0185]
[0186] Where u and u' are two adjacent drones; R u,u' B represents the transmission rate between drone u and its neighboring drone u'. u,u' P represents the transmission bandwidth between drone u and its neighboring drone u'; u,u' The transmission power between UAV u and its neighboring UAV u'; The position of UAV u in time slot t; The position of the adjacent UAV u' in time slot t;
[0187] To ensure the vehicle receives the requested content, the minimum transmission power of the drone should meet the following requirements:
[0188]
[0189] in, The minimum transmit power of the drone u is required for vehicle user v to receive the requested content in time slot t. The signal-to-noise ratio of the communication link between vehicle user v and drone u in time slot t;
[0190] use This represents the decision variable for sharing content f between drone u and its neighboring drone u', where γ is the signal-to-noise ratio of the communication link between drone u and its neighboring drone u'. u,u' Greater than the second signal-to-noise ratio threshold η UAV At this time, a communication link is formed. The formula is:
[0191]
[0192] Type 4: Transmission rate between vehicle users and macro base stations (V2B):
[0193] Macro base stations can communicate with vehicle users via V2B links, with the following transmission rates:
[0194]
[0195] Among them, R v,b B represents the transmission rate between vehicle user v and macro base station b. b,v P represents the transmission bandwidth between vehicle user v and macro base station b. v,b The transmit power between vehicle user v and macro base station b; g v,b The channel gain between vehicle user v and macro base station b;
[0196] Step 3.3: Construct a content download request process latency model; the working process of the content download request process latency model is as follows:
[0197] Step 3.3.1: When a vehicle user sends a download request for content f, it first checks whether the content is found in the vehicle user's own cache. If the content is found in the vehicle user's cache, the vehicle user does not need to make a cache decision and directly outputs f with zero latency. The vehicle user's own cache hit decision is then confirmed. Increase by 1; otherwise, proceed to step 3.3.2 for judgment.
[0198] Step 3.3.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to directly download content f from neighboring vehicle users. Considering that the channel gain changes over time, content f is obtained from neighboring vehicle users at this time. k The required time can be expressed as:
[0199]
[0200] If not, proceed to step 3.3.3 for judgment;
[0201] in, The size of the k-th content requested in time slot t; For vehicle user v and adjacent vehicle user v', the k-th content f k Transmission delay;
[0202] Step 3.3.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. If the value is increased by 1, proceed to step 3.3.4 for judgment; additionally, the waiting time for the vehicle user's request task is calculated by the drone. If it does not exceed the maximum tolerable delay, the request enters the buffer queue to wait for drone-to-vehicle user (U2V) transmission; otherwise, it is forwarded to the macro base station service. Drone u and vehicle user v represent the service content f. k The transmission delay is
[0203] Among them, s k t-1 , These represent the sizes of the kth content requested before time slot t-1 and t, respectively. The average transmission rate between drones and vehicle users;
[0204] To facilitate the representation of transmission time R v,v' (t), formula (25) is obtained by taking the average rate over that time period. To simplify the problem, the first part of the above formula represents the latency of waiting for the drone to process the content in the queue, and the second part represents the latency of the drone processing the currently requested content.
[0205] Step 3.3.4: If the requested content is not found in the cache list of the adjacent drone u, determine whether the requested content can be found in the cache list of the adjacent drone u'. If so, a cache hit decision is made for the adjacent drone u'. If 1 is added, u' transmits the content to u via content sharing; otherwise, proceed to step 3.3.5 for judgment. For newly arrived content, drone u acts as a relay for forwarding. The content transmission delay is:
[0206]
[0207] in, For drone u and neighboring drone u', the service content for vehicle users f k Transmission delay; This represents the average transmission rate between drones.
[0208] Step 3.3.5, the k-th content f kThe time required for data to be transmitted from the macro base station to the vehicle user is:
[0209]
[0210] in, For macro base station b and vehicle user v, the k-th content f k Transmission delay; R v,b The transmission rate between vehicle user v and macro base station b.
[0211] Step 3.4: Construct an energy consumption model for the content download request process; the calculation process for the energy consumption model for the content download request process is as follows:
[0212] Step 3.4.1: The vehicle user sends a content request f. k First, determine if the content can be found in vehicle user v. If so, then make a cache hit decision for the vehicle user itself. Increase by 1, the vehicle directly outputs f k If the energy consumption is 0, proceed to step 3.4.2 for judgment;
[0213] Step 3.4.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to download content f directly from nearby service vehicle users. k At this time, the energy consumption is:
[0214]
[0215] in, For the content f between vehicle user v and adjacent vehicle user v' k Transmission energy consumption; For the content f between vehicle user v and adjacent vehicle user v' k Transmission delay;
[0216] If not, proceed to step 3.4.3 for judgment;
[0217] Step 3.4.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. Increasing by 1, the energy consumption is now:
[0218]
[0219] in, For content f between vehicle user v and drone u k Transmission energy consumption; P u,vThe transmit power from the drone (u) to the vehicle user (v); For content f between drone u and vehicle user v k Transmission delay; P u The transmit power of the UAV u;
[0220] If not, proceed to step 3.4.4 for judgment;
[0221] Step 3.4.4: Determine whether the requested content can be found in the cache list of neighboring drones u'. If so, make a cache hit decision for neighboring drones u'. Increasing by 1, the energy consumption is now:
[0222]
[0223] in, The energy consumption for transmission between vehicle user v and adjacent drone u'; For the content f between drone u and its neighboring drone u' k Transmission energy consumption;
[0224] If not, proceed to step 3.4.5 of the calculation;
[0225] Steps 3, 4, and 5, Content f k The energy consumption for transmission via macro base stations is:
[0226]
[0227] in, For content f between macro base station b and vehicle user v k Transmission energy consumption; P b,v The transmit power from macro base station b to vehicle user v;
[0228] Step 4: Define the optimization problem and solve it using deep reinforcement learning based on D3QN to provide a request cache placement and replacement strategy that maximizes caching benefits for the requested content; the specific process is as follows:
[0229] Step 4.1: Determine the final optimization problem as minimizing system latency and energy consumption while maximizing system benefits.
[0230] Define the caching benefit function as U cache :
[0231] U cache =(U time +U ene )p f (32);
[0232] Among them, U time For time efficiency; U eneFor energy cost-effectiveness;
[0233] System latency and energy consumption are two key components of system cost. According to step 3.2, the content f received by vehicle user v in each time slot t via four types of links can be calculated. k The total time is
[0234]
[0235] in, Indicates time slot t, vehicle user v, and content f. k Cache status; Indicates time slot t, UAV u, and content f k Cache status, Indicates the content f of time slot t k It is cached by the drone and does not exceed the waiting time. Similarly, For the content f of adjacent UAVs in time slot t. k Cache status; Indicates the content f of time slot t k It is cached by the adjacent drone u'.
[0236] Considering transmission costs, according to step 3.4, vehicle user v receives content f through four types of links in each time slot t. k The system energy consumption cost can be expressed as
[0237]
[0238] in, These represent the energy consumption costs of transmitting content over different link units, depending on factors such as transmission distance.
[0239] Because there is a difference in magnitude between delay and energy consumption, the time benefit can be written in the following form:
[0240]
[0241] in, This indicates the time improvement in obtaining content from edge nodes compared to the time from base stations; normalization has been performed in this invention.
[0242] Similarly, the energy consumption cost-benefit function can be expressed as:
[0243]
[0244] The final optimization problem P is determined to be the function of maximizing cache efficiency:
[0245]
[0246] Among them, C1, C2, C3, C4, and C5 are different constraints; For the content f of the Mth node in time slot t k Cache status; To receive content f k Total time; For content f k Maximum tolerable delay; C1 and C2 respectively represent the time to be ... The constraints are in binary; C3 indicates that the cache size for vehicle users and drones does not exceed the maximum cache resource limit; C4 indicates that content can only be cached in one place; C5 indicates that the maximum tolerable latency does not exceed the limit.
[0247] Step 4.2: In order to solve the above optimization problem P, we map it to a multiple knapsack problem and prove that it is NP-hard (nondeterministic polynomial). Next, we use deep reinforcement learning to solve and analyze it.
[0248] First, prove that P is NP-hard:
[0249] The central idea of proving that optimization problem P is NP-hard is to find a mapping process that maps the Multiple Knapsack Problem (MKP) to optimization problem P. This type of problem is characterized by a finite quantity of each item; it's neither a problem where only one item can be chosen (as in the 0 / 1 knapsack problem) nor a problem where an unlimited number of items can be chosen (as in the unbounded knapsack problem). The Multiple Knapsack Problem is a known NP-hard problem whose goal is to select items from a set of knapsacks to maximize the total value of the items while satisfying the capacity constraints of each knapsack. Its definition is as follows:
[0250]
[0251] Where `maximize` means to maximize; `X` is the number of knapsacks, and `Y` is the number of items. `C1'`, `C2'`, and `C3'` are different constraints. Parameters and Let L represent the weight and profit of the y-th item, respectively. x Let x represent the capacity of the x-th knapsack. (Decision variable) The decision variable represents whether the y-th item should be assigned to the x-th knapsack if and only if the y-th item is assigned to the x-th knapsack. It is set to 1. Constraint C2' ensures that each item is assigned to at most one knapsack; constraint C3' indicates that the decision variable is binary.
[0252] Transform the above MKP problem instance into an optimization problem P instance. Treat each knapsack capacity in the MKP problem as the cache capacity of vehicle users and drones in the optimization problem P. Treat each item as a content item in the optimization problem P. Map the weight and profit of each item in the MKP problem to the cache size and cache value of each content item in the optimization problem P, respectively.
[0253] Given the parameters above, we construct an instance of an optimization problem P with Z request nodes and F content items. For each request node (considering vehicle user v and drone u as request node M), the content... set up: Among them, L z ω is the cache (knapsack) capacity of the z-th request node; f The space occupied by the content f (volume);
[0254] The optimization problem P can be written as P1 + question:
[0255]
[0256] in, For the z-th request node in time slot t, the content f k Cache status; For different constraints; For content f k The space occupied by the size;
[0257] Since the power consumption of knapsack movement is not involved in the multiple knapsack problem, the following proof assumes that the vehicle user, drone coordinates, and transmission power are fixed, thus for the content f k Transmission latency obtained by vehicle users It also needs to be rewritten as:
[0258]
[0259] in, For the z-th request node, obtain the content f. k Time;
[0260] P1 + The problem can now be written as question:
[0261]
[0262] in, The energy cost of transmission through different link units;
[0263] The objective function consists of a linear sum of two terms, where, in a given context... Since it can be considered a constant, the interpretation of the cache energy efficiency part is equivalent to the solution of the cache time efficiency part. The solution of cache time efficiency is arguably the key step in solving the entire problem, and the complexity of the entire problem depends on the computational complexity of this part. Therefore, if the time efficiency problem is reduced to a multiple knapsack problem, the entire problem can be described as NP-hard. The objective function of the problem becomes question:
[0264]
[0265] According to the question, P3 + The objective function of the problem can be written in the form of: Since the constant has no effect, P3 can be... + The partial extraction from the question can be written in the following form:
[0266]
[0267] For the z-th request node and content f k ,exist The question and MKP By performing equivalence, a polynomial time reduction is constructed, which maps instances of the MKP problem to instances of the optimization problem P.
[0268] We use deep reinforcement learning based on D3QN to select the most appropriate cache placement location for the requested content;
[0269] Iterative training is performed based on the D3QN algorithm. Each iteration involves the following operations: the number and size of requested content, cached resources, and external environment information are used as the state. The input is fed into the D3QN algorithm to generate a policy μ and select action a. t Whether it is cached and where it is cached, based on a t Place the request content; execute a t The environment will then return a reward r. t And it will get the new state in the next time slot. Will a t r t s t+1 Input the D3QN algorithm for training;
[0270] In the D3QN algorithm, the agent obtains a quadruple of the action taken in the current state, the reward after taking the action, the next state, and the current state through interaction with the environment: middle, Indicates the state of time slot t, a t Indicates the action taken in time slot t, r t This represents the reward received in time slot t. This represents the state of time slot t+1. These quadruplets are stored in the experience replay buffer as the basis for learning. To avoid storing too many experience samples, the D3QN algorithm deletes the oldest quadruplet sample when the experience replay buffer is full.
[0271] Each quadruple in the experience replay buffer is used to train the D3QN algorithm. The goal of the D3QN algorithm is to optimize the evaluation network and the action network by calculating a target value. The target value is calculated using the following formula:
[0272]
[0273] in, The target value for time slot t; Q target It is the Q-value function of the target network; It is the discount factor; 'a' represents the action; θ represents the discount factor. target This represents the parameters of the target network. In this way, the D3QN algorithm uses the target network to compute the Q-value and reduces overestimation of the Q-value.
[0274] Using gradient descent, the D3QN algorithm updates the parameters of the evaluation network based on the calculated target value. The evaluation network parameters θ... eval The update formula is:
[0275]
[0276] Where ← represents the update operation; It is about finding the gradient; It evaluates the network's learning rate; Q eval It is the Q-value function for evaluating the network.
[0277] The D3QN algorithm updates the parameters of the action network through policy gradient ascent, aiming to maximize the agent's reward. The action network parameters θ... action The update formula is:
[0278]
[0279] Where β is the learning rate of the action network, Q eval It evaluates the Q-value of the network output.
[0280] After each training iteration, the parameters of the target network are updated based on the evaluation network. The parameters θ of the target network... target The update formula is:
[0281] θ target ←τθ eval +(1-τ)θ target (47);
[0282] Here, τ is the soft update coefficient, which controls the weight update rate between the target network and the evaluation network. The D3QN algorithm learns iteratively and updates the evaluation network and action network to obtain a high-performance value function. Ultimately, it can optimize content caching decisions based on the system's cache placement strategy, ensuring improved cache hit rate and energy efficiency.
[0283] The optimization problem aims to maximize cache efficiency under latency and cache space constraints. The reward function is calculated using the network model of the D3QN algorithm, and consists of two parts:
[0284] R m (t)=R(t)+αSC m (t) (48);
[0285] Among them, R m (t) represents the total reward for the m-th vehicle user (one vehicle user is one agent) in time slot t; SC m R(t) represents the social contribution rate of the m-th vehicle user in time slot t; α is the factor that adjusts the social contribution participation rate. Additionally, let R(t) be the reward for time slot t, representing the buffering utility of the objective function time slot t, expressed as a piecewise function:
[0286]
[0287] Among them, D t The delay for acquiring content in time slot t; The maximum tolerable delay for the content f in time slot t; Indicates whether a cache node hit occurred; R max The maximum reward value is set to a custom value.
[0288] After a predetermined number of iterations, a high-performance value function network was finally trained, thus obtaining the optimal cache placement scheme.
[0289] Table 1 shows the pseudocode table for caching based on users' personal historical request preferences.
[0290] Table 1 Algorithm Pseudocode
[0291]
[0292] like Figure 3 As shown, the algorithm mainly consists of the following steps:
[0293] ① Macro base stations acquire user and request information;
[0294] ② Calculate the historical preferences of each user and the cache matching degree of the requested content relative to each node;
[0295] ③ Calculate the cache efficiency function over a period of time;
[0296] ④ The D3QN algorithm is used for training to obtain high-performance value function network parameters.
[0297] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A collaborative caching decision-making method based on user preferences in a vehicle-to-everything (V2X) network with unmanned aerial vehicle (UAV) assistance, characterized in that, Includes the following steps: Step 1: Construct an edge collaborative caching architecture model in the Internet of Vehicles (IoV) consisting of vehicle users, drones, and macro base stations; Step 2: Calculate the cache value of the corresponding content using the vehicle user's historical request content topics and the current request content topics; Step 3: Calculate the cache value under different conditions and the transmission rate under different types of links, and construct a latency model and energy consumption model for the content download request process; Step 4: Identify the optimization problem and use the D3QN algorithm for deep learning to solve it, providing a request cache placement and replacement strategy that maximizes the caching benefits for the requested content.
2. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with the assistance of unmanned aerial vehicles (UAVs) as described in claim 1, characterized in that, In step 1, the vehicle user set is defined as follows: It caches its own preferred content locally and shares the cached content with users in nearby vehicles via vehicle-to-vehicle links. m Let m be the m-th vehicle user, and V be the total number of vehicle users; define the set of drones as... u n Let U be the nth drone, and U be the total number of drones. Each drone is equipped with a cache server to assist the macro base station in providing content caching services to vehicle users. They share and transmit content among themselves, providing service to vehicle users only during peak hours. During off-peak hours, service is transferred to the macro base station to reduce drone energy consumption. In each time slot, the set of content requests randomly issued by vehicle users is Q = {f...} 1,1 ,...,f 1,k ,...,f m,1 ,...,f m,k }, f m,k For the k-th content requested by the m-th vehicle user; define the content library as... f k Let F be the k-th content, and F be the total number of content.
3. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with UAV assistance as described in claim 2, characterized in that, The specific process of step 2 is as follows: Step 2.1: Calculate the probability of a vehicle user's request based on the request sent by the vehicle user. The formula is as follows: Where, p f ρ is the probability of requesting content f; ρ is the Zipf exponent. Step 2.2: Calculate the content caching value using the following formula: in, ω1 and ω2 are the cache values of content f in time slot t at different node levels; ω1 and ω2 are different adjustment parameters. This represents the m-th vehicle user v. m For the k-th content f of the current request k Interests and preferences; Interests and preferences are closely related to content topics. The set of topics is defined as W = {w1, w2, ..., w...} j ,...,w J }, w j Let j represent the j-th topic, where J is the total number of topics; define a binary variable e. k (f k ,w j ) represents the k-th content f k and the j-th topic w j The correlation; The cache value of the current request is determined based on the topics of the vehicle user's historical requests, whether it's a spur-of-the-moment decision or due to content preference; historical average request information is defined based on mean cosine theory. The formula is: Among them, f a f b These are two different contents from the vehicle user's historical requests; K is the number of contents in the historical requests; e(f a ,w j ) is f a with w j The correlation; e(f) b ,w j ) is f b with w j The correlation; Determine the m-th vehicle user v based on the topic of the request content. m The requested content f is at this point. k and historical average request information The relevance is used to represent the relationship between the m-th vehicle user and the k-th content f requested in the current request. k Degree of interest preference:
4. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with UAV assistance as described in claim 3, characterized in that, The specific process of step 3 is as follows: Step 3.1: Construct formulas for calculating cache value under different conditions; Step 3.2: Calculate the transmission rate under different types of links; Step 3.3: Construct a latency model for the content download request process; Step 3.4: Construct an energy consumption model for the content download request process.
5. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with unmanned aerial vehicle assistance as described in claim 4, characterized in that, Step 3.1 includes three scenarios; Scenario 1: When vehicle user v caches content f, the cache value of time slot t is Wherein, ω1 and ω2 are different adjustment parameters. In case 1, ω1 = 1 and ω2 = 0. Scenario 2: For content f requested by drone u within the coverage area, the cache value of time slot t is In case 2, ω1 = 0 and ω2 = 1; Scenario 3: In the content f request, the cache of adjacent vehicle user v' within the vehicle's communication coverage area is [value missing]. The cache value of time slot t is [value missing]. Where α is the moderating factor for social contribution participation; SC m The social contribution rate of the m-th vehicle user.
6. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with UAV assistance as described in claim 5, characterized in that, In step 3.2, the vehicle user acts as both the content requester and provider. If the content exists in their own cache, the request will be completed immediately without establishing a communication link. Otherwise, based on the various content cache locations, the vehicle user retrieves the content from four types of links and calculates the transmission rate under each of the four types of links. Type 1: Obtain data from nearby adjacent vehicle users and calculate the transmission rate between vehicle users; the formula is: Where v and v' are two adjacent vehicle users; R v,v' B represents the transmission rate of shared content between vehicle user v and its adjacent vehicle user v'. v,v' P is the transmission bandwidth between vehicle user v and its adjacent vehicle user v'; v,v' N0 is the transmit power between vehicle user v and its neighboring vehicle user v'; N0 is the noise power; g v,v' It is the channel gain between vehicle user v and adjacent vehicle user v'; Type 2: Obtain data from drones above the activity area and calculate the transmission rate from vehicle users to the drones; the formula is: Among them, R u,v B represents the transmission rate between vehicle user v and drone u. u,v This represents the transmission bandwidth between the drone u and the vehicle user v; It is the transmit power of the drone u to the vehicle user v in time slot t; h v,u σ represents the air-to-ground channel gain between vehicle user v and drone u. 2 It is the variance of Gaussian noise; Type 3: Acquire data in cooperation with neighboring drones, and calculate the inter-drone transmission rate; the formula is: Where u and u' are two adjacent drones; R u,u' B represents the transmission rate between drone u and its neighboring drone u'. u,u' P represents the transmission bandwidth between drone u and its neighboring drone u'; u,u' The transmission power between UAV u and its neighboring UAV u'; The position of UAV u in time slot t; Let g0 be the position of the adjacent UAV u' in time slot t; g0 represents the power gain when the reference distance is 1 meter. Type 4: Obtain from macro base stations and calculate the transmission rate between vehicle users and macro base stations; the formula is: Among them, R v,b B represents the transmission rate between vehicle user v and macro base station b. b,v P represents the transmission bandwidth between vehicle user v and macro base station b. v,b The transmit power between vehicle user v and macro base station b; g v,b Let v be the channel gain between vehicle user v and macro base station b.
7. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with unmanned aerial vehicle assistance as described in claim 6, characterized in that, In step 3.3, the working process of the content download request latency model is as follows: Step 3.3.1: When a vehicle user sends a download request for content f, it first checks whether the content is found in the vehicle user's own cache. If the content is found in the vehicle user's cache, the vehicle user does not need to make a cache decision and directly outputs f with zero latency. The vehicle user's own cache hit decision is then confirmed. Increase by 1; Otherwise, proceed to step 3.3.2 for judgment; Step 3.3.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to directly download content f from neighboring vehicle users; if not, proceed to step 3.3.3 for judgment; wherein, obtaining content f from neighboring vehicle users... k The required time is: in, The size of the k-th content requested in time slot t; For vehicle user v and adjacent vehicle user v', the k-th content f k Transmission delay; Step 3.3.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. If the value is increased by 1, proceed to step 3.3.4 for judgment; in addition, the waiting time for the vehicle user's request task is calculated by the drone. If it does not exceed the maximum tolerable delay, it enters the buffer queue to wait for the drone to transmit to the vehicle user; otherwise, it is forwarded to the macro base station; drone u and vehicle user v are the service content f. k The transmission delay is Among them, s k t-1 , These represent the sizes of the kth content requested before time slot t-1 and t, respectively. The average transmission rate between drones and vehicle users; Step 3.3.4: If the requested content is not found in the cache list of the adjacent drone u, determine whether the requested content can be found in the cache list of the adjacent drone u'. If so, a cache hit decision is made for the adjacent drone u'. Increase by 1, u' transmits the content to u via content sharing; otherwise, proceed to step 3.3.5 for calculation; for newly arrived content, drone u acts as a relay for forwarding; calculate the content transmission delay: in, For drone u and neighboring drone u', the service content for vehicle users f k Transmission delay; This represents the average transmission rate between drones. Step 3.3.5: Calculate the k-th content f k Time required for data to be transmitted from the macro base station to the vehicle user: in, For macro base station b and vehicle user v, the k-th content f k Transmission delay; R v,b The transmission rate between vehicle user v and macro base station b.
8. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with unmanned aerial vehicle assistance as described in claim 7, characterized in that, In step 3.4, the calculation process of the energy consumption model for the content download request process is as follows: Step 3.4.1: The vehicle user sends a content request f. k First, determine if the content can be found in vehicle user v. If so, then make a cache hit decision for the vehicle user itself. Increase by 1, the vehicle directly outputs f k If the energy consumption is 0, proceed to step 3.4.2 for judgment; Step 3.4.2: Determine if the content can be found in the nearest neighboring vehicle user within the communication range. If so, make a cache hit decision for the nearest neighboring vehicle user. Add 1, requesting vehicle user v to directly download content f from neighboring vehicle users. k If not, proceed to step 3.4.3 for judgment; Content f between vehicle user v and adjacent vehicle user v' k Transmission energy consumption for: Step 3.4.3: Determine whether the requested content can be found in the drone u within the communication range of the requesting vehicle user v. If so, make a cache hit decision for the drone u within the communication range. Increase by 1; otherwise, proceed to step 3.4.4 for judgment. Content f between vehicle user v and drone u k The transmission energy consumption is Among them, P u,v The transmit power from the drone (u) to the vehicle user (v); For content f between drone u and vehicle user v k Transmission delay; P u The transmit power of the UAV u; Step 3.4.4: Determine whether the requested content can be found in the cache list of neighboring drones u'. If so, make a cache hit decision for neighboring drones u'. Increase by 1; otherwise, proceed to step 3.4.5 of the calculation. The transmission energy consumption between vehicle user v and adjacent drone u' is in, For the content f between drone u and its neighboring drone u' k Transmission energy consumption; Steps 3, 4, and 5, Content f k The energy consumption for transmission via macro base stations is: in, For content f between macro base station b and vehicle user v k Transmission energy consumption.
9. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with unmanned aerial vehicle assistance as described in claim 8, characterized in that, In step 4, the final optimization problem P is determined to be the function of maximizing cache efficiency: in, Indicates time slot t, vehicle user v, and content f. k Cache status; Indicates the time slot t of the UAV u to the content f k The caching status; T represents a time period; C1, C2, C3, C4, and C5 represent different constraints; For the content f of the Mth node in time slot t k Cache status; s k This indicates the size of the k-th element; For content f k Maximum tolerable delay; U cache The caching efficiency function is expressed by the following formula: IN cache =(U time +U ene )p f (21); Among them, U time For time efficiency; U ene For energy consumption cost-effectiveness; the formulas are as follows: Among them, R v2v For the transmission rate between vehicle users; For vehicle user v, content f is received in each time slot t via four types of links. k The total time is calculated using the formula: in, For the content f of adjacent UAVs in time slot t. k Cache status; For vehicle user v, content f is received in each time slot t via four types of links. k The system energy consumption cost is given by the formula: in, These represent the energy consumption cost of transmitting content through different link units.
10. The collaborative caching decision-making method based on user preferences in the Internet of Vehicles with unmanned aerial vehicle assistance as described in claim 9, characterized in that, In step 4, deep reinforcement learning based on D3QN is used to select the most appropriate cache placement location for the requested content. In the D3QN algorithm, the agent obtains a quadruple of the action taken in the current state, the reward after taking the action, the next state, and the current state through interaction with the environment. in, Indicates the state of time slot t, a t Indicates the action taken in time slot t, r t This represents the reward received in time slot t. This represents the state of time slot t+1; these quadruplets are stored in the experience replay buffer as the basis for learning; when the experience replay buffer is full, the D3QN algorithm will delete the oldest quadruplet sample. Each quadruple in the experience replay buffer is used to train the D3QN algorithm; the goal of the D3QN algorithm is to optimize the evaluation network and action network by calculating a target value; the formula for calculating the target value is: in, The target value for time slot t; Q target It is the Q-value function of the target network; It is the discount factor; 'a' represents the action; θ represents the discount factor. target This represents the parameters of the target network; the D3QN algorithm uses the target network to compute the Q-value and reduces overestimation of the Q-value. Using gradient descent, the D3QN algorithm updates the parameters of the evaluation network based on the calculated target value; the evaluation network parameters θ eval The update formula is: Where ← represents the update operation; It is about finding the gradient; It evaluates the network's learning rate; Q eval It is a function for evaluating the Q-value of a network; The D3QN algorithm updates the parameters of the action network through policy gradient ascent, aiming to maximize the agent's reward; the action network parameters θ action The update formula is: Where β is the learning rate of the action network; After each training iteration, the parameters of the target network are updated based on the evaluation network; the parameters θ of the target network are... target The update formula is: i target ←tth eval +(1-τ)θ target (29); Where τ is the soft update coefficient; The D3QN algorithm learns iteratively and updates the evaluation network and action network to obtain a high-performance value function, thereby obtaining the optimal cache placement scheme; finally, it optimizes the content caching decision based on the cache placement scheme.
Citation Information
Patent Citations
Priori knowledge driven active cache decision-making method in Internet of Vehicles
CN117460001A
Content deployment optimization method for unmanned aerial vehicle cooperative caching 6G network
CN118740929A
Method for drone-assisted caching in vehicular network on basis of geographical location
WO2024164528A1