An intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles
By adopting an intelligent cooperative caching strategy that combines asynchronous federated learning and social perception in the Internet of Vehicles, and using deep reinforcement learning to optimize the selection and update of cache content, the edge caching problem affected by vehicle mobility is solved, achieving a higher cache hit rate and lower request latency.
Patent Information
- Application Number
- CN202411380723.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-09-30
AI Technical Summary
In the Internet of Vehicles, edge cache capacity is limited, and vehicle mobility affects the performance of asynchronous federated learning, resulting in a decrease in global model accuracy. In addition, existing learning algorithms fail to effectively optimize the selection and update of cached content, resulting in request content transmission delays and resource consumption problems.
An intelligent cooperative caching strategy combining asynchronous federated learning and social perception is adopted. Global and local caching models are trained through deep reinforcement learning. Social perception and asynchronous federated learning framework between vehicles are utilized to optimize the update and selection of cache content, reduce the transmission delay of requested content, and improve the cache hit rate.
It significantly improves cache utilization, reduces the average vehicle request latency, and improves cache hit rate and edge cache performance, especially outperforming other baseline solutions in highly dynamic vehicle network environments.
Smart Images

Figure CN119277346B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and in particular relates to an intelligent cooperative caching strategy based on asynchronous federated learning and social perception in an Internet of Vehicles. Background Art
[0002] With the development of the Internet of Vehicles (IoV) and edge computing, caching technology has become a key enabler for diverse multimedia real-time applications, such as intelligent navigation, efficient pattern recognition, and rich multimedia entertainment experiences, significantly improving driving convenience and entertainment. Given the high demands of in-vehicle applications for computing, communication, and storage resources, the requirements for service response speed and network bandwidth are very stringent. Intelligent vehicles, leveraging their powerful computing and caching capabilities, enable real-time interaction with edge devices, pedestrians, and other vehicles through communication technologies such as C-V2X and LTE-V. Edge caching technology reduces duplicate data transmission, optimizes transmission efficiency, addresses the challenges of telematics data, reduces system burden and energy consumption, and meets the requirements of computationally intensive, low-latency in-vehicle applications. Furthermore, edge cache capacity is limited, making the design of an efficient collaborative caching solution a significant research topic.
[0003] In collaborative caching systems, the selection of service nodes and the updating of cached content are two key issues that need to be addressed. The rise of federated learning and certain learning algorithms has addressed some of these challenges. For example, some studies have used federated learning algorithms to select high-quality service nodes, aggregate all vehicle models, and construct a global model. This model predicts popular content and pre-caches it to cooperating vehicles or roadside units. Collaborative caching reduces access latency and improves vehicle service quality. In traditional federated learning, the global model is periodically updated by aggregating the local models of all vehicles. However, vehicles may frequently move out of the coverage area of edge servers and fail to upload all local models as expected, resulting in a decrease in the accuracy of the global model. Asynchronous federated learning eliminates the need to aggregate all vehicle local models. The more local models uploaded, the more accurate the global model. However, vehicle mobility can significantly impact the performance of asynchronous federated learning. Furthermore, the rise of intelligent transportation has also applied deep reinforcement learning methods to edge caching systems, introducing reinforcement learning caching strategies whereby each roadside unit agent independently selects caching actions, optimizing resources to maximize cumulative rewards. However, the vast state-action space in the learning algorithm has led to increasingly significant resource consumption issues. Summary of the Invention
[0004] The present invention provides an intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles. It incorporates social perception between vehicles and collaboratively updates the cache through deep reinforcement learning training under the asynchronous federated learning framework, which can reduce the transmission delay of the requested content and improve the cache hit rate.
[0005] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0006] An intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles (IoV) includes:
[0007] The user vehicle sends a cache request to the local RSU;
[0008] After receiving the cache request, the local RSU determines whether there is a request content corresponding to the cache request in its own cache content. If so, the request content is sent to the user vehicle. If not, a nearby user vehicle, a neighboring RSU or a macro base station with the request content is selected according to the preset priority to respond to the cache request of the user vehicle, and the request content is also sent to the local RSU.
[0009] Each user vehicle within the coverage of the local RSU: After receiving the request content corresponding to its own cache request, it uses the vehicle's local cache update model to update the vehicle's local cache content, and updates and trains the vehicle's local cache update model based on the vehicle's current local experience pool data;
[0010] After the local cache update model of each selected vehicle is updated, it is uploaded to the local RSU; each selected vehicle is selected by its local RSU based on the vehicle's residence time within its coverage area;
[0011] When the local RSU receives any local cache update model, it updates the global cache update model based on the corresponding vehicle weight;
[0012] The local RSU uses the current global cache update model to update the cache content of the local RSU based on the request content corresponding to the cache requests of all user vehicles in its service area.
[0013] Furthermore, the preset priorities, from high to low, are: nearby user vehicles, adjacent RSUs, and macro base stations, and the methods for sending the request content are:
[0014] Nearby user vehicles use social perception between vehicles to send the request content to the user vehicle;
[0015] The neighboring RSU first sends the request content to the local RSU, which then sends it to the user vehicle;
[0016] The macro base station receives the cache request from the user vehicle and then directly sends the request content to the user vehicle.
[0017] Furthermore, the local cache update model of the vehicle is referred to as the local model, and the global cache update model of the RSU is referred to as the global model. The initial local model of each vehicle is obtained by downloading the global model of the local RSU. The global model and the local model use a deep reinforcement learning algorithm to decide whether to update their respective cache contents.
[0018] The global model uses a deep reinforcement learning algorithm to decide whether to update the cache content and is defined as:
[0019] Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the global model of the nth RSU in time slot t, It represents the cache status of K preset cache contents corresponding to the global model of the nth RSU in time slot t, represents the cache status of the global model of the nth RSU for the kth preset cache content in time slot t, Represents cached. Represents not cached; represents the request status of the i-th user vehicle for K preset cache contents in time slot t, represents the request status of the i-th vehicle for the k-th preset cache content in time slot t, Send a request on behalf of Indicates that the request was not sent; P k Indicates the cache location of the kth preset cache content; is the set of user vehicle numbers within the coverage area of the nth RSU;
[0020] The action is: whether to replace multiple cache contents with the request content corresponding to the cache request, expressed as in, represents the action of the global model of the nth RSU in time slot t, Indicates whether the global model of the nth RSU replaces the kth preset cache content with the request content corresponding to the cache request currently received from the i-th user vehicle in time slot t, Indicates replacement, Indicates no replacement;
[0021] The reward is: minimize the cache request latency of each user's vehicle;
[0022] The local model of the i-th user vehicle uses a deep reinforcement learning algorithm to decide whether to update the cache content and is defined as:
[0023] Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the current local model at time slot t, Indicates the cache status of the current local model for K preset cache contents in time slot t, Indicates the cache status of the current local model for the kth preset cache content in time slot t, Represents cached. Represents not cached;
[0024] The action is: whether to replace multiple cache contents with the request content corresponding to the cache request, expressed as in, represents the action of the current local model in time slot t, Indicates whether the current local model replaces the k-th preset cache content with the requested content obtained from the cache request in time slot t. Indicates replacement, Indicates no replacement;
[0025] The reward is: minimize the cache request latency of the i-th user vehicle.
[0026] Furthermore, the reward functions of the global model and each local model are:
[0027]
[0028] In the reward function of the global model, (S t ,A t )Pick In the reward function of the local model, (S t ,A t )Pick D t represents the delayed reward at time t; λ1, λ2, λ3, and λ4 represent the learning rates in the reward function of nearby vehicles, local RSU, nearby RSU, and macro base station response cache requests, respectively. V,i d R,i 、 d M,i They represent the transmission delays of nearby vehicles, local RSU, nearby RSU, and macro base station in response to the cache request of the i-th vehicle.
[0029] Furthermore, the transmission delay d of the nearby vehicles in responding to the cache request of the i-th vehicle is V,i The calculation formula is:
[0030]
[0031] Where s represents the size of the request content in response to the cache request; RV,i Represents the distance between nearby vehicles V and the i-th vehicle V i The transmission rate between nearby vehicles V; α represents the contact rate between nearby vehicles V and vehicle V i Probability of establishing communication; N represents the total number of vehicles with requested content cached in time slot t within the coverage area of the local RSU. t represents the total number of vehicles within the coverage area of the local RSU in time slot t; B is the available bandwidth; p V Indicates the transmit power level used by nearby vehicles, dis(V,V i ) indicates the proximity of vehicles V and V i The distance between i (dis(V,V i )) indicates vehicle V i The channel gain with the nearby vehicle V, is the noise power;
[0032] The transmission delay d of the local RSU responding to the cache request of the i-th vehicle R,i The calculation formula is:
[0033]
[0034] Where R R,i Represents the local RSU and the user vehicle V i The transmission rate between B Indicates the transmit power level used by the local RSU; dis(S,V i ) represents the local RSU and vehicle V i The distance between i (dis(S,V i )) represents the local RSU and vehicle V i The channel gain between
[0035] Transmission delay of the nearby RSU in responding to the cache request of the i-th vehicle The calculation formula is:
[0036]
[0037] Where R R-R Indicates the transmission rate between adjacent RSUs;
[0038] The transmission delay d of the macro base station in responding to the cache request of the i-th vehicle M,i The calculation formula is:
[0039]
[0040] Where R M,i Represents the macro base station and the user vehicle Vi The transmission rate between M Indicates the transmit power level used by the macro base station; dis(M,V i ) represents the macro base station and vehicle V i The distance between i (dis(M,V i )) represents the macro base station and vehicle V i The channel gain between .
[0041] Furthermore, let x = V, M, S represent nearby vehicles, macro base stations, and local RSUs, respectively. Then the nearby vehicles, local RSUs, macro base stations, and vehicles V i The channel gain calculation formula between is:
[0042] h i (dis(x,V i ))=α i (dis(x,V i ))g i
[0043] Where, α i (dis(x,V i )) is the large-scale fading effect including path loss and shadowing, where the path loss is calculated as 128.1+37.6log 10 (dis(x,V i ), the shadow follows a log-normal distribution; g i It is a small-scale fading effect.
[0044] Furthermore, the method for the local RSU to select a vehicle based on the vehicle's residence time within its coverage area is as follows:
[0045] Calculate V for any vehicle i The duration of stay within the local RSU coverage area
[0046]
[0047] Where, L s is the coverage of the local RSU, P i For user vehicle V i The distance that the current round travels within the local RSU coverage area, S i For user vehicle V i speed;
[0048] If the vehicle V i The duration of stay within the local RSU coverage area Greater than the average training time T T and inference time T I The sum of vehicle Vi Participate in asynchronous federated training as a selected vehicle.
[0049] Furthermore, the vehicle weight used to update the global cache update model is calculated as follows:
[0050]
[0051] Where, χ i For user vehicle V i The weight of L; μ1 and μ2 are the coefficients of position weight and transmission weight respectively, μ1+μ2=1; s is the coverage of the local RSU, P i For user vehicle V i The distance that the current round travels within the local RSU coverage area, R R,i Represents the local RSU and the user vehicle V i The transmission rate between R,i' Represents the local RSU and any vehicle V within its coverage area i' The transmission rate between them, V' is the number set of selected vehicles within the coverage of the local RSU.
[0052] An intelligent cooperative caching system based on asynchronous federated learning and social perception in an Internet of Vehicles includes a macro base station, several RSUs, and several user vehicles. Each RSU has several user vehicles within its coverage area. The user vehicles send a cache request to the RSU within its coverage area. The RSU that receives the cache request is the local RSU, and the remaining RSUs are nearby RSUs. The intelligent cooperative caching system based on asynchronous federated learning and social perception in an Internet of Vehicles is used to implement any of the above-mentioned intelligent cooperative caching strategies based on asynchronous federated learning and social perception in an Internet of Vehicles.
[0053] Compared with the prior art, the significant advantages of the present invention include:
[0054] 1. By considering the mobility characteristics of vehicles such as position and speed, an asynchronous federated learning algorithm is proposed to improve the accuracy of the global model.
[0055] 2. Use the DQN reinforcement learning algorithm to learn the request content data of vehicle users in each edge device. DQN can make optimal caching decisions, reduce the average request latency of vehicles, and improve the caching performance of each edge device.
[0056] 3. Leveraging vehicle users' social networks to capture vehicle contact rates in different areas and dynamically detect content popularity based on changes in content requests. Furthermore, by leveraging opportunistic V2V communication and data sharing informed by the concept of vehicle social networks, the system integrates multiple in-vehicle communication methods, significantly improving cache utilization.
[0057] 4. The collaborative caching strategy (AFDQN) based on asynchronous federation and deep reinforcement learning proposed in this paper fully utilizes idle road resources, significantly improving edge caching performance and reducing average vehicle request latency. Experimental results show that AFDQN outperforms other baseline caching schemes in cache hit rate and average vehicle request latency in highly dynamic vehicle network environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is the technical process of the embodiment of the present application;
[0059] Figure 2 This is the network architecture of the embodiment of the present application;
[0060] Figure 3 It is the DQN algorithm under asynchronous federation in the embodiment of the present application;
[0061] Figure 4 is the cache hit rate under different local RSU cache capacities in the embodiment of the present application;
[0062] Figure 5 is the request delay under different local RSU cache capacities in the embodiment of the present application;
[0063] Figure 6 is the cache hit rate under different content amounts in the embodiment of the present application;
[0064] Figure 7 is the request delay under different content amounts in the embodiment of the present application;
[0065] Figure 8 is the average training time per round of asynchronous federated learning and federation in the embodiment of the present application; DETAILED DESCRIPTION
[0066] The following is a detailed description of an embodiment of the present invention. This embodiment is based on the technical solution of the present invention, provides a detailed implementation method and a specific operation process, and further explains the technical solution of the present invention.
[0067] This embodiment provides an intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles. The network architecture consists of a macro base station (MBS), several roadside units (RSUs), and several user vehicles. The MBS is connected to a cloud center. In the network architecture, there are several user vehicles within the coverage area of each RSU. The user vehicle sends a cache request to the RSU within its coverage area. The RSU that receives the cache request is the local RSU, and the remaining RSUs are nearby RSUs. The vehicle stores a large amount of historical vehicle user data, including personal information, requested content, and ratings. RSUs communicate with each other via wired links and are all equipped with edge servers for computing and caching. Therefore, RSUs can cache various content to meet the content service needs of vehicle users. The vehicle will search for nearby vehicles or the nearest RSU that have cached the requested content and establish communication to obtain the content. Otherwise, the nearest RSU to the vehicle will download the content from nearby RSUs that have cached the content or directly from the MBS to provide services to the vehicle.
[0068] The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles of this embodiment is referenced. Figure 1 As shown, the following process is included:
[0069] S1. The user vehicle sends a cache request to the local RSU to obtain the corresponding request content.
[0070] S2. After receiving the cache request, the local RSU determines whether there is a request content corresponding to the cache request in its own cache content: if so, the request content is sent to the user vehicle; if not, a nearby user vehicle, neighboring RSU or macro base station with the request content is selected according to the preset priority to respond to the cache request of the user vehicle, and the request content is also sent to the local RSU.
[0071] Local RSU: If the cache content of the local RSU already contains the request content corresponding to the cache request, the local RSU sends the request content back to the vehicle.
[0072] Nearby vehicles: Cached vehicles can share content with other vehicles. Due to the limited coverage of wireless signals, the premise of vehicle-to-vehicle communication is that the distance between vehicles is less than a certain value. The event in which two vehicles are very close and establish a communication relationship with each other is called vehicle contact. These areas may have different vehicle densities and driving patterns, resulting in different vehicle social attributes. Therefore, the vehicle contact rate can be used as a key indicator to characterize vehicle social relationships. The vehicle contact time also mainly depends on the contact rate α v , the contact time follows an exponential distribution. In the service area of RSU n, the contact probability between vehicles is: Socially aware caching leverages social network concepts such as node centrality and social communities to ensure that selected collaborative vehicles achieve the required data transmission ratio within a given timeframe. If the requested content is not cached in the local RSU, the vehicle directly sends the request to a nearby vehicle. If the nearby vehicle has the requested content cached, the request is sent to the vehicle.
[0073] Neighboring RSU: If the requested content is not cached in the local RSU or in a nearby vehicle, the local RSU forwards the request to the neighboring RSU. If the neighboring RSU has the requested content cached, the requested content is sent to the local RSU. The local RSU then sends the content back to the vehicle.
[0074] MBS: If the requested content is not cached in the local RSU, nearby vehicles, or adjacent RSUs, the vehicle sends a cache request to the MBS, which then directly sends the requested content back to the vehicle.
[0075] S3. Each user vehicle within the coverage of the local RSU: After receiving the request content corresponding to each cache request, the local cache update model of the vehicle is used to update the local cache content of the vehicle, and the local cache update model of the vehicle is updated and trained based on the current experience pool data of the vehicle.
[0076] The local cache update model of the vehicle and the global cache update model of the RSU in the present invention are used to make the current optimal decision based on the cache status and request status of the current cache content of the vehicle itself: whether to use the request content to update its own cache content so that the cache content can respond to subsequent cache requests more quickly.
[0077] The vehicle's local cache update model is referred to as the local model, and the RSU's global cache update model is referred to as the global model. Each vehicle's initial local model is obtained by downloading the global model of its local RSU. The global and local models use a deep reinforcement learning algorithm to decide whether to update their respective cache contents.
[0078] (1) Global model.
[0079] The global model uses a deep reinforcement learning algorithm to decide whether to update the cache content, which is defined as:
[0080] Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the global model of the nth RSU in time slot t, It represents the cache status of K preset cache contents corresponding to the global model of the nth RSU in time slot t, represents the cache status of the global model of the nth RSU for the kth preset cache content in time slot t, Represents cached. Represents not cached; represents the request status of the i-th user vehicle for K preset cache contents in time slot t, represents the request status of the i-th vehicle for the k-th preset cache content in time slot t, Send a request on behalf of Indicates that the request was not sent; P k Indicates the cache location of the kth preset cache content; is the set of user vehicle numbers within the coverage of the nth RSU.
[0081] Action: Preset whether multiple cache contents are replaced with the request content corresponding to the cache request, expressed as
[0082] in, represents the action of the global model of the nth RSU in time slot t, Indicates whether the global model of the nth RSU replaces the kth preset cache content with the request content corresponding to the cache request currently received from the i-th user vehicle in time slot t, Indicates replacement, Indicates no replacement.
[0083] The reward is: Minimize the cache request delay of each user's vehicle. Expressed as:
[0084]
[0085] In the formula, λ1+λ2+λ3+λ4=1, 0<λ1<λ2<<λ3<λ4; D t represents the delayed reward at time t; λ1, λ2, λ3, and λ4 represent the learning rates in the reward functions of nearby vehicles, local RSU, nearby RSU, and macro base station in response to cache requests, respectively; d V,i d R,i 、 d M,i They represent the transmission delays of nearby vehicles, local RSU, nearby RSU, and macro base station in response to the cache request of the i-th vehicle.
[0086] In this embodiment, vehicles use orthogonal frequency division multiplexing technology to communicate with RSUs and cache vehicles, so interference is not considered in the communication model. Communication between the local RSU and adjacent RSUs uses wired links. Each vehicle maintains the same communication model in a round and changes its communication model in different asynchronous rounds. The transmission delays of nearby vehicles, local RSUs, nearby RSUs, and macro base stations in response to the i-th vehicle cache request are calculated as follows:
[0087] ① Transmission delay d of nearby vehicles responding to the cache request of vehicle i V,i The calculation formula is:
[0088]
[0089] Where s represents the size of the request content in response to the cache request; R V,i Represents the distance between nearby vehicles V and the i-th vehicle V i The transmission rate between nearby vehicles V; α represents the contact rate between nearby vehicles V and vehicle V i Probability of establishing communication; N represents the total number of vehicles with requested content cached in time slot t within the coverage area of the local RSU. t represents the total number of vehicles within the coverage area of the local RSU in time slot t; B is the available bandwidth; p V Indicates the transmit power level used by nearby vehicles, dis(V,V i ) indicates the proximity of vehicles V and V i The distance between i (dis(V,V i )) indicates vehicle V i The channel gain with the nearby vehicle V, is the noise power.
[0090] The vehicle model in this embodiment assumes that all vehicles travel in the same direction and arrive at one RSU based on Poisson distribution, with an average arrival rate of λ v . Once a vehicle enters the coverage area of the local RSU, it will send a request message to the local RSU. We assume that no vehicle enters or leaves the coverage area of the local RSU in a round, so the number of vehicles in the local RSU in a round is fixed. Each vehicle maintains the same mobility characteristics, including position and speed, within a round, and may change its mobility characteristics at the beginning of each round. The speeds of different vehicles follow independent and identical distributions. The speed of each vehicle is generated by a truncated Gaussian distribution, which is flexible and conforms to the real dynamic vehicle environment. For any r-th round, the number of vehicles traveling in the coverage area of the local RSU is N, and the vehicle set is denoted as Where V iLet {S1, S2, ..., S i ,…,S N} is the speed of all vehicles traveling in the local RSU, where S i V i Speed. i The probability density function of is calculated as:
[0091]
[0092] S max and S min are the maximum and minimum speed thresholds for each vehicle, For S i With mean μ and variance σ 2 The Gaussian error function under . Let For user vehicle V i The distance traveled within the local RSU coverage area, T(r) is the duration of the rth round, then V i The distance traversed can be calculated as:
[0093] ② Transmission delay d of the local RSU responding to the cache request of the i-th vehicle R,i The calculation formula is:
[0094]
[0095]
[0096] Where R R,i Represents the local RSU and the user vehicle V i The transmission rate between B Indicates the transmit power level used by the local RSU; dis(S,V i ) represents the local RSU and vehicle V i The distance between i (dis(S,V i )) represents the local RSU and vehicle V i The channel gain between
[0097] ③ Transmission delay of the nearby RSU in response to the cache request of the i-th vehicle The calculation formula is:
[0098]
[0099] Where R R-R Indicates the transmission rate between adjacent RSUs;
[0100] ④ Transmission delay d of the macro base station in responding to the cache request of the i-th vehicleM,i The calculation formula is:
[0101]
[0102] Where R M,i Represents the macro base station and the user vehicle V i The transmission rate between M Indicates the transmit power level used by the macro base station; dis(M,V i ) represents the macro base station and vehicle V i The distance between i (dis(M,V i )) represents the macro base station and vehicle V i The channel gain between .
[0103] Let x = V, M, S represent nearby vehicles, macro base stations, and local RSUs, respectively. Then the relationship between nearby vehicles, local RSUs, macro base stations, and vehicles V i The channel gain calculation formula in the current round is:
[0104] h i (dis(x,V i ))=α i (dis(x,V i ))g i
[0105] Where, α i (dis(x,V i )) is the large-scale fading effect including path loss and shadowing, where the path loss is calculated as 128.1+37.6log 10 (dis(x,V i ), the shadow follows a log-normal distribution; g i is the small-scale fading effect, which is assumed to be exponentially distributed with unit mean.
[0106] In addition, the global model mentioned above needs to be pre-trained using a training set initially so that the local model can download it as the initial local model.
[0107] The global model uses a deep reinforcement learning algorithm and is pre-trained using the training set:
[0108] Estimate the action-value function Q(S) by updating the Q-learning iteration t ,A t ), as shown below:
[0109]
[0110] Where γ≥0 represents the learning rate. The current RSU takes system action A tAnd get instant reward R(S t ,A t ) and then the current state S t Update to the next state S t+1 Then, the optimal control strategy for each RSU is obtained according to the c-∈-greedy strategy, and A ※ express:
[0111]
[0112] Where ∈ represents the probability of choosing a random action. This greedy strategy can effectively avoid saddle points by trying random new actions. The size of each RSU cache is f m , store historical data through the experience pool M and update it according to the latest experience. t Through Action A t Transform into action t+1 And get a reward R(S t ,A t ), we can get a batch training set (S t ,A t ,R(S t ,A t ),S t+1 ), and store the training set in the experience playback M. During the entire system training process, at each moment RSU randomly selects a small batch of experience pools from the experience pool M And by minimizing Q(S t ,A t ;ω t ) to train QNN. Then, the prediction loss function of QNN is defined as follows:
[0113] Loss(ω t )=E[(q(S t+1 ,A t+1 )-Q(S t ,A t ;ω t )) 2 ]
[0114] Where, ω t Represents the weight parameter of each iteration of QNN training, Q(S t ,A t ;ω t ) is the output of QNN at time slot t, and its approximate value function is defined as follows:
[0115]
[0116] in, is the QNN in a given state St+1 The maximum output obtained when selecting an action is A t+1 . Then, the update is performed using the semi-gradient algorithm as follows:
[0117]
[0118] Where ρ>0 is the gradient step size, Yes t The updated gradient.
[0119] (2) Local model.
[0120] The initial local model of each vehicle is obtained by downloading the global model of the local RSU. That is, the local model of the i-th user vehicle also uses the deep reinforcement learning algorithm to decide whether to update the cache content, which is defined as:
[0121] Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the current local model at time slot t, Indicates the cache status of the current local model for K preset cache contents in time slot t, Indicates the cache status of the current local model for the kth preset cache content in time slot t, Represents cached. Represents not cached; represents the request status of the i-th user vehicle for K preset cache contents in time slot t, represents the request status of the i-th vehicle for the k-th preset cache content in time slot t, Send a request on behalf of Indicates that the request was not sent; P k Indicates the cache location of the kth preset cache content; is the set of user vehicle numbers within the coverage area of the nth RSU;
[0122] The action is: whether to replace multiple cache contents with the request content corresponding to the cache request, expressed as in, represents the action of the current local model in time slot t, Indicates whether the current local model replaces the k-th preset cache content with the requested content obtained from the cache request in time slot t. Indicates replacement, Indicates no replacement;
[0123] The reward is: Minimize the cache request delay of the i-th user vehicle. The cache request delay of the user vehicle It is consistent with the global model and is expressed as:
[0124]
[0125] After obtaining the initial local model of each vehicle by downloading the global model of the local RSU, the local model is retrained on the local dataset.
[0126] S4. After the local cache update model of each selected vehicle is updated, it is uploaded to the local RSU; each selected vehicle is selected by its local RSU based on the vehicle's stay time within its coverage area.
[0127] This invention builds, trains, and updates global and local models based on an asynchronous federation framework. For each round of asynchronous federation, after a vehicle uploads its local model, the local RSUs aggregate and update their global model. This embodiment considers vehicle mobility and proposes an asynchronous federation algorithm applicable to the Internet of Vehicles.
[0128] (1) Vehicle Selection: Vehicles that have stayed at the local RSU long enough are selected to ensure that they have the opportunity to participate in the asynchronous federation and complete the training process. Each vehicle first sends its mobility characteristics, including speed and position, to the local RSU. The local RSU then selects a vehicle based on the residence time calculated from the vehicle's mobility characteristics.
[0129] User Vehicle V i The dwell time within the local RSU coverage area is calculated as:
[0130]
[0131] Among them, L s is the coverage of the local RSU.
[0132] V i The residence time should be greater than the average training time T T and inference time T I The sum of the user's vehicle V i can complete a round of training. Therefore, if User Vehicle V i There is an opportunity to participate in asynchronous FL training. Otherwise, the user vehicle V i Ignored.
[0133] (2) Download model: Send the parameters obtained by RSU training as initial values to the selected vehicle.
[0134] (3) Local training: selected vehicle V in the local RSU iThe downloaded global model ω is set as the initial local model and the local model is updated through training iterations. Afterwards, the updated local model is uploaded to the local RSU. For each iteration k, V i Randomly extract some training data n from the original training set k The loss function under the local model is calculated as:
[0135]
[0136] Calculate n k After the loss function of all data in , the local loss function after k iterations is calculated as:
[0137]
[0138] Among them, |N k | for n k The number of data in .
[0139] Then calculate the regularized local loss function to reduce the local model ω i,k The deviation from the global model ω improves the convergence of the algorithm:
[0140]
[0141] where ρ is the regularization parameter.
[0142] set up is g(ω i,k ) is called a local gradient. In the previous round, some vehicles may fail to upload their updated local models due to training delays, thus affecting the convergence of the global model. These vehicles are referred to as stragglers, and the local gradients of these stragglers in the previous round are called delayed local gradients. The delayed local gradients are aggregated into the local gradients of the current round:
[0143]
[0144] Where β is the attenuation coefficient, is the delayed local gradient. If V i If the upload was successful in the previous round,
[0145] Then the local model for the next iteration is updated as:
[0146]
[0147] Where η is the local learning rate of the current round, which is calculated as: η = ηmax{1, log(r)}. l is the initial value of the local learning rate. r is the current round.
[0148] Then iteration k ends, V i Randomly sample some training data again and start the next iteration. When the number of iterations reaches the threshold e, V i Complete local training and update the local model ω i ,Right now Upload to the local RSU. Further get the global objective function:
[0149]
[0150] where d i V i The size of the local data in , d is the total size of the local data of the selected vehicle. The overall goal of asynchronous federated learning is to find a global model ω * To minimize the objective function, that is:
[0151] ω * =MinG(ω)
[0152] (4) Update and upload model: After completing the current round of training of the local model, the vehicle uploads the updated local model to the local RSU.
[0153] S5. When the local RSU receives any local cache update model, it updates the global cache update model based on the corresponding vehicle weight.
[0154] Considering the mobility of vehicles, we design an asynchronous aggregation of the global model based on mobility awareness:
[0155]
[0156] Among them, χ i It is V i The weight of the asynchronous aggregation.
[0157] By considering the user vehicles V within the coverage area of the local RSU i The traversing distance and the distance from the local RSU to V i The content transmission delay is calculated by asynchronous aggregation V i To improve the accuracy of the global model and reduce content transmission delay. Specifically, if V i If the traversal distance is large, it may take a long time to participate in training, so its local model should occupy a larger weight for aggregation to improve the accuracy of the global model. In addition, from the local RSU to V i The content transmission delay of V is important because when the content is cached in the local or neighboring RSU, i Eventually, the content will be downloaded from the local RSU. Therefore, if the local RSU is connected to the V iThe content transmission delay of is small, and its local model should also occupy a larger weight for aggregation to reduce the content transmission delay. The weight of asynchronous aggregation χ i Calculated as:
[0158]
[0159] Where μ1 and μ2 are the coefficients of location weight and transmission weight respectively (i.e. μ1+μ2=1). The values of μ1 and μ2 can be set according to the location and transmission in actual application. s is the size of the request content. Therefore, from the local RSU to V i The content transmission delay is affected by the local RSU and V i The effect of the transmission rate between and R R,i Further calculation of χ i ,Right now:
[0160]
[0161] Since the local RSU knows each vehicle V at the beginning of the asynchronous federation i dis(S,V i ) and P i , so the local RSU can calculate R R,i , further calculate χ i So far, the rth round of asynchronous federated learning has been completed, and the updated global model ω has been obtained. r .
[0162] S6. The local RSU uses the current global cache update model to update the cache content of the local RSU based on the request content corresponding to the cache requests of all user vehicles in its service area.
[0163] To verify the technical effect of the present invention, the following experimental hardware environment uses the Windows 10 operating system, the CPU is Intel Core i7-8750k (2.20GHz), the memory is 8GB, and it is developed and verified on the PyCharm platform. Multiple roadside units RSU are distributed on a road, and the local roadside unit, the adjacent roadside unit, the macro base station and the vehicle perform collaborative caching. The overall architecture of the experiment in this embodiment is as follows Figure 2 As shown, the red car is a mobile vehicle. The vehicle moves in the city, the local roadside unit communicates with the vehicles within its coverage area, the neighboring roadside units communicate with the local roadside unit, the macro base station communicates with the vehicle, and the vehicles communicate with each other.
[0164] Figure 3This is the asynchronous federated DQN algorithm training used in this embodiment of the present invention. Red represents vehicles selected for training, while gray represents unselected vehicles. The local roadside unit (RSU) performs DQN training locally to obtain initial global parameters. Selected vehicles then download these global parameters for local training. To account for vehicle mobility, the first trained vehicle uploads its model. Finally, the model is aggregated based on each vehicle's asynchronous aggregation weights, and the global model is uploaded.
[0165] The cache hit rate under different local RSU cache capacities is as follows: Figure 4 As shown. The cache hit rate of the RSU is controlled within the range of 50 to 400, and the cache hit rate of the present invention and the other three baseline schemes is compared. It can be seen that as the capacity increases, the cache hit rate of all schemes increases. The cache hit rate of the present invention is improved by an average of about 45%, 5%, 11% and 2% respectively compared with the other strategies. This is because the local RSU caches more content with a larger capacity. Under the limited cache capacity of the RSU, the RSU will naturally generate more cache updates. Therefore, it is more likely to obtain the vehicle's requested content from the local RSU. In addition, the random scheme provides the worst cache hit rate because the scheme only randomly selects content. In addition, the performance of the present invention, c-∈-greedy and CAFR is better than that of the random scheme and Thompson sampling. This is because the random and Thompson sampling schemes cannot predict the cache content through learning, while the present invention, c-∈-greedy and CAFR schemes determine the cache content by observing the historical request content. In addition, the present invention is slightly better than c-∈-greedy and CAFR.
[0166] The request delay under different local RSU cache capacities is as follows Figure 5 As shown. As the cache capacity increases, the content transmission delay of each scheme is decreasing. The request delay of the present invention is reduced by an average of about 15%, 7%, and 11% compared with the random, c-∈-greedy, and Thompson sampling schemes, respectively. This is because as the cache capacity increases, each RSU will cache more content, and each vehicle is more likely to obtain content from the local RSU, thereby reducing the delay in content transmission. The content transmission delay of the present invention is always in a better position. This is because the cache hit rate of the present invention is higher than that of other schemes, and more vehicles can obtain content directly from the local RSU, thereby reducing the delay in content transmission.
[0167] The cache hit rate under different content amounts is as follows Figure 6As shown in the figure, by varying the number of contents (from 200 to 900), the cache hit rates of the five algorithms all decrease with increasing content size, given the limited RSU cache capacity. However, the proposed method can better observe the historical requests and rewards for cached content, so its cache hit rate decreases less than that of the other algorithms.
[0168] Request delays under different content amounts are as follows Figure 7 Figure 2 shows the content transmission delays of different schemes for different content amounts. As the amount of content increases, the transmission delays of all five algorithms increase slightly. This is because the cache hit rate decreases with the increasing amount of content, given the limited cache capacity of the RSU. However, the increase in request latency of the proposed algorithm is smaller than that of other algorithms, making it more stable overall.
[0169] The average training time per round for asynchronous federation and federation is as follows Figure 8 As shown in the figure, the training time per round of the present invention is within 2.5s and 3s, while the training time per round of the FedAVG solution is within 31s and 33s. This shows that the training time of the present invention is much shorter than that of the FedAVG solution. This is because the FedAVG solution needs to aggregate the local models of all vehicles for each round of global model update, while the asynchronous federation aggregates the local models of the vehicles as they are received in each round.
[0170] The above embodiments are preferred embodiments of the present application. Ordinary technicians in this field can also make various changes or improvements on this basis. Without departing from the overall concept of the present application, these changes or improvements should fall within the scope of protection required by the present application.
Claims
1. An intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles, characterized by: include: The user vehicle sends a cache request to the local RSU; After receiving the cache request, the local RSU determines whether there is a request content corresponding to the cache request in its own cache content. If so, the request content is sent to the user vehicle. If not, a nearby user vehicle, a neighboring RSU or a macro base station with the request content is selected according to the preset priority to respond to the cache request of the user vehicle, and the request content is also sent to the local RSU. Each user vehicle within the coverage of the local RSU: After receiving the request content corresponding to its own cache request, it uses the vehicle's local cache update model to update the vehicle's local cache content, and updates and trains the vehicle's local cache update model based on the vehicle's current local experience pool data; After the local cache update model of each selected vehicle is updated, it is uploaded to the local RSU; each selected vehicle is selected by its local RSU based on the vehicle's residence time within its coverage area; When the local RSU receives any local cache update model, it updates the global cache update model based on the corresponding vehicle weight; The local RSU uses the current global cache update model to update the cache content of the local RSU based on the request content corresponding to the cache requests of all user vehicles in its service area; The local cache update model of the vehicle is referred to as the local model, and the global cache update model of the RSU is referred to as the global model. The initial local model of each vehicle is obtained by downloading the global model of the local RSU. The global model and the local model use a deep reinforcement learning algorithm to decide whether to update their respective cache contents. The global model uses a deep reinforcement learning algorithm to decide whether to update the cache content. The state includes the cache status and request status of multiple preset cache contents. The action is to replace the preset cache contents with the request content corresponding to the cache request. The reward is to minimize the cache request delay of each user vehicle. The local model of the i-th user vehicle uses a deep reinforcement learning algorithm to decide whether to update the cache content. In the definition, the state includes the cache state and request state of multiple preset cache contents. The action is: whether to replace the preset multiple cache contents with the request content corresponding to the cache request. The reward is: minimizing the cache request delay of the i-th user vehicle; The reward functions of the global model and each local model are: Where D t represents the delayed reward at time t; λ1, λ2, λ3, and λ4 represent the learning rates in the reward function of nearby vehicles, local RSU, nearby RSU, and macro base station response cache requests, respectively. V,i d R,i 、 d M,i They represent the transmission delays of nearby vehicles, local RSU, nearby RSU, and macro base station in response to the cache request of the i-th vehicle.
2. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 1 is characterized in that: The preset priorities, from high to low, are: nearby user vehicles, adjacent RSUs, and macro base stations. The methods for sending the request content are: Nearby user vehicles use social perception between vehicles to send the request content to the user vehicle; The neighboring RSU first sends the request content to the local RSU, which then sends it to the user vehicle; The macro base station receives the cache request from the user vehicle and then directly sends the request content to the user vehicle.
3. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 1 is characterized in that: The global model uses a deep reinforcement learning algorithm to decide whether to update the cache content and is defined as: Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the global model of the nth RSU in time slot t, It represents the cache status of K preset cache contents corresponding to the global model of the nth RSU in time slot t, represents the cache status of the global model of the nth RSU for the kth preset cache content in time slot t, Represents cached. Represents not cached; represents the request status of the i-th user vehicle for K preset cache contents in time slot t, represents the request status of the i-th vehicle for the k-th preset cache content in time slot t, Send a request on behalf of Indicates that the request was not sent; P k Indicates the cache location of the kth preset cache content; is the set of user vehicle numbers within the coverage area of the nth RSU; The action is: whether to replace multiple cache contents with the request content corresponding to the cache request, expressed as in, represents the action of the global model of the nth RSU in time slot t, Indicates whether the global model of the nth RSU replaces the kth preset cache content with the request content corresponding to the cache request currently received from the i-th user vehicle in time slot t, Indicates replacement, Indicates no replacement; The reward is: minimize the cache request latency of each user's vehicle; The local model of the i-th user vehicle uses a deep reinforcement learning algorithm to decide whether to update the cache content and is defined as: Status: includes cache status and request status of preset multiple cache contents, expressed as in, represents the state of the current local model at time slot t, Indicates the cache status of the current local model for K preset cache contents in time slot t, Indicates the cache status of the current local model for the kth preset cache content in time slot t, Represents cached. Represents not cached; The action is: whether to replace multiple cache contents with the request content corresponding to the cache request, expressed as in, represents the action of the current local model in time slot t, Indicates whether the current local model replaces the k-th preset cache content with the requested content obtained from the cache request in time slot t. Indicates replacement, Indicates no replacement; The reward is: minimize the cache request latency of the i-th user vehicle.
4. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 1 is characterized in that: The transmission delay d of the nearby vehicles responding to the cache request of the i-th vehicle V,i The calculation formula is: Where s represents the size of the request content in response to the cache request; R V,i Represents the distance between nearby vehicles V and the i-th vehicle V i The transmission rate between nearby vehicles V; α represents the contact rate between nearby vehicles V and vehicle V i Probability of establishing communication; N represents the total number of vehicles with requested content cached in time slot t within the coverage area of the local RSU. t represents the total number of vehicles within the coverage area of the local RSU in time slot t; B is the available bandwidth; p V Indicates the transmit power level used by nearby vehicles, dis(V,V i ) indicates the proximity of vehicles V and V i The distance between i (dis(V,V i )) indicates vehicle V i The channel gain with the nearby vehicle V, is the noise power; The transmission delay d of the local RSU responding to the cache request of the i-th vehicle R,i The calculation formula is: Where R R,i Represents the local RSU and the user vehicle V i The transmission rate between B Indicates the transmit power level used by the local RSU; dis(S,V i ) represents the local RSU and vehicle V i The distance between i (dis(S,V i )) represents the local RSU and vehicle V i The channel gain between Transmission delay of the nearby RSU in responding to the cache request of the i-th vehicle The calculation formula is: Where R R-R Indicates the transmission rate between adjacent RSUs; The transmission delay d of the macro base station in responding to the cache request of the i-th vehicle M,i The calculation formula is: Where R M,i Represents the macro base station and the user vehicle V i The transmission rate between M Indicates the transmit power level used by the macro base station; dis(M,V i ) represents the macro base station and vehicle V i The distance between i (dis(M,V i )) represents the macro base station and vehicle V i The channel gain between .
5. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 4 is characterized in that: Let x = V, M, S represent nearby vehicles, macro base stations, and local RSUs, respectively. Then the relationship between nearby vehicles, local RSUs, macro base stations, and vehicles V i The channel gain calculation formula between is: h i (dis(x,V i ))=α i (dis(x,V i ))g i Where, α i (dis(x,V i )) is the large-scale fading effect including path loss and shadowing, where the path loss is calculated as 128.1+37.6log 10 (dis(x,V i ), the shadow follows a log-normal distribution; g i It is a small-scale fading effect.
6. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 1 is characterized in that: The method by which the local RSU selects a vehicle based on the vehicle's residence time within its coverage area is: Calculate V for any vehicle i The duration of stay within the local RSU coverage area Where, L s is the coverage of the local RSU, P i For user vehicle V i The distance that the current round travels within the local RSU coverage area, S i For user vehicle V i speed; If the vehicle V i The duration of stay within the local RSU coverage area Greater than the average training time T T and inference time T I The sum of vehicle V i Participate in asynchronous federated training as a selected vehicle.
7. The intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles according to claim 1 is characterized in that: The vehicle weight used to update the global cache update model is calculated as: Where, χ i For user vehicle V i The weight of L; μ1 and μ2 are the coefficients of position weight and transmission weight respectively, μ1+μ2=1; s is the coverage of the local RSU, P i For user vehicle V i The distance that the current round travels within the local RSU coverage area, R R,i Represents the local RSU and the user vehicle V i The transmission rate between R,i' Represents the local RSU and any vehicle V within its coverage area i' The transmission rate between them, V' is the number set of selected vehicles within the coverage of the local RSU.
8. An intelligent cooperative caching system based on asynchronous federated learning and social perception in the Internet of Vehicles, characterized by: It includes a macro base station, several RSUs and several user vehicles. There are several user vehicles within the coverage range of each RSU. The user vehicle sends a cache request to the RSU in its coverage range, and the RSU that receives the cache request is the local RSU, and the remaining RSUs are nearby RSUs; the intelligent cooperative caching system based on asynchronous federated learning and social perception in the Internet of Vehicles is used to implement the intelligent cooperative caching strategy based on asynchronous federated learning and social perception in the Internet of Vehicles described in any one of claims 1-7.
Citation Information
Patent Citations
Cooperative edge caching method based on asynchronous federation and deep reinforcement learning
CN115297170A