Improved RSU popular file distribution method based on deep reinforcement learning
By introducing a deep reinforcement learning method with dual Q learning and experience-first playback between RSUs, the problem of overestimation of Q value and low sample utilization in popular file allocation between RSUs is solved, and higher cache hit rate and lower file transfer delay are achieved.
Patent Information
- Application Number
- CN202510091991.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, popular file allocation between RSUs has problems with duplicate sampling or missed sampling caused by overestimation of Q value and random sampling, resulting in misallocation of files, reducing cache hit rate and increasing file transfer delay.
Adopt the popular RSU file allocation method based on deep reinforcement learning, introduce dual Q learning and experience priority playback, optimize file allocation decisions between RSUs, improve the hit rate of cached files and reduce file transfer delay.
Through dual Q learning, reduce Q value overestimation, priority experience playback improves sample utilization, significantly improves cache hit rate and reduces file transfer delay, and solves the problem of poor file allocation in the existing technology.
Smart Images

Figure CN120050626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle edge computing, and in particular to an improved RSU popular file allocation method based on deep reinforcement learning. Background Art
[0002] The Internet of Vehicles can enable communication between vehicles and roadside units, as well as between vehicles. With the popularization of autonomous driving and in-vehicle multimedia services, the amount of data in computationally intensive and delay-sensitive applications in the Internet of Vehicles has grown rapidly. Therefore, researchers have proposed Vehicle Edge Computing (VEC), which deploys computing or storage devices at the network edge (such as roadside units, base stations, or idle vehicles), effectively reducing backbone network congestion and data transmission latency.
[0003] Vehicle edge computing can be divided into vehicle offloading computing tasks and vehicle requesting data files. Task offloading mainly studies how vehicles optimize offloading decisions. File requests mainly study how various servers respond to requests. Taking in-vehicle multimedia applications as an example, vehicles may frequently make file requests. To shorten the transmission delay and reduce the load on the cloud server, edge caching technology can improve file response performance by caching data files at edge nodes (base stations or idle vehicles, etc.). However, there are also some problems with edge caching. Usually, the storage capacity of a single edge node is limited and cannot cache all the content required by vehicles. If the caching node does not have the file currently requested by the vehicle, on the one hand, it reduces the cache file hit rate, and on the other hand, it also needs to send a forwarding request to the cloud server, increasing the transmission delay.
[0004] In the Internet of Vehicles environment, due to storage space limitations, isolated caching nodes can only provide the file content cached by themselves and cannot meet various file requests made by numerous vehicles. Through file transfer between caching nodes to achieve cooperative caching, the present invention performs file allocation on two RSUs. When the local RSU does not cache a certain popular file, it can query and obtain the required content from the neighboring RSU. The above RSU - to - RSU file allocation mechanism can optimize the utilization rate of the caching space of edge nodes, increase the probability of file acquisition, and reduce the content acquisition time. Therefore, it is crucial to design an efficient RSU - to - RSU file allocation scheme.
[0005] Existing researchers have proposed a popular file allocation scheme based on asynchronous federation, "Mobility-Aware Cooperative Caching in Vehicular Edge Computing Based on Asynchronous Federation and Deep Reinforcement Learning". After the RSU determines the popular files, the popular files are cached in the local RSU and adjacent RSUs respectively. However, the above file allocation research also has the following problems. First, the problem of overestimating the Q value of the Dueling-DQN algorithm. Second, the problem of repeated sampling or missing sampling caused by the Dueling-DQN randomly sampling empirical data. The above technical problems will lead to incorrect allocation of popular files, so this is the subject to be solved by the present invention. Summary of the Invention
[0006] The purpose of the present invention is to: aiming at the problem of poor file allocation in the existing scheme, to implement an improved RSU popular file allocation method based on deep reinforcement learning in the vehicle network, so as to achieve the goal of improving the cache hit rate and reducing the file transmission delay.
[0007] The inventive concept of the present invention is: The present invention proposes an improved RSU popular file allocation method based on deep reinforcement learning. The present invention designs and introduces double Q learning and experience prioritized replay, and proposes an improved RSU popular file allocation method based on deep reinforcement learning. Accordingly, the local RSU and neighbor RSUs can optimize the cache file allocation.
[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is specifically: an improved RSU popular file allocation method based on deep reinforcement learning, including the following steps:
[0009] S1: In each round, based on the file requests of all vehicles within the communication range, two RSUs determine F c popular files to be cached;
[0010] S2: The RSU initializes the main Q network, the target Q network and the experience replay buffer;
[0011] S3: The RSU trains the target Q network and updates the main Q network;
[0012] S4: Based on the results of deep reinforcement learning, the local RSU and neighbor RSUs allocate popular files to respond to vehicle file requests;
[0013] S1 specifically includes the following steps:
[0014] S101. For urban roads, the scenario of the present invention includes a macro base station (MBS) with a cloud center attached, a roadside unit (RSU) with an edge server, and vehicles with storage capabilities. For the cloud center connected to the MBS, it is assumed that the cloud server has sufficient cache capacity and computing power to cache all available files and provide file services to vehicles. The RSUs deployed on both sides of the road are responsible for forwarding the data exchanged between the MBS and vehicle users. To speed up file response, the RSUs are configured with roadside servers, but the storage capacity of the RSUs is limited and they can only cache some files. The content requested by the vehicle is first sent to the local RSU through "vehicle-to-roadside unit" (V2R) communication. If the local RSU has cached the file, it responds to the vehicle request. If the local RSU has not cached it, the vehicle needs to request from the cloud server connected to the MBS. To simplify the problem, only two roadside units (also called base stations hereinafter), namely the local RSU and the neighboring RSU, are considered in this scenario.
[0015] S102. Based on the file requests of all vehicles within the communication range of the base stations, the two RSUs determine F c popular files to be cached. Specifically, each vehicle submits to its base station a user preference information (vehicle user ID, score values of each file) containing all files. After the RSU receives the above information submitted by each vehicle, it summarizes it into a file preference score table. The element M x,y in the x-th row and y-th column of the table represents the preference value of vehicle user x for file y, and its value range is in the interval [0, 1]. 0 represents that the user's interest level in this file is 0, and 1 represents that the user's interest level in this file is 1. Based on this file preference score table, by summing up for file y the popularity of the file (i.e., the total score value of file y) can be obtained. At this point, the RSU can obtain the popularity of all files and the set N of popular files. The storage capacity of a single RSU is limited and it can only cache some files. Therefore, the present invention considers distributing popular files between the local RSU and the neighboring RSU. This is to improve the file cache hit rate and reduce the file transmission delay.
[0016] S103. In a round r, the algorithm executes Ns time slots t. At the beginning of the r-th round, the local RSU randomly selects c files from the set N of popular files for caching, and the neighboring RSU randomly caches c files from N that have not been selected by the local RSU.
[0017] S2 specifically includes the following steps:
[0018] S201. Initialize the main Q network:
[0019] Take the file requests of all vehicles within the local RSU as the input of the main Q-network ω. Define the state s(t) and the action a(t), which represent the file set and the file update decision respectively. a(t)=1 means that the local RSU needs to randomly select a part of the files from N to replace the local files. a(t)=0 means that the local RSU does not need to update the local cached files. Using the main Q-network, calculate the Q of different file allocation decisions a (a ∈ set {0, 1}) of the RSU at time slot t present The value function is
[0020]
[0021] where V(s(t); ω) is the state value function only related to s(t) in ω, A(s(t), a; ω) is the value function of the decision a(t) in ω, and E[A(s(t), a; ω)] is the mean value of all decision value functions.
[0022] S202. Initialize the target Q-network:
[0023] The target Q-network ω - has the same network architecture as the main Q-network ω, but the target Q-network is used to calculate the impact after the RSU takes the decision a(t). Therefore, calculate the Q next value at s(t + 1), and its expression can be written as
[0024]
[0025] where V(s(t + 1); ω - ) is the state value function only related to s(t + 1) in ω - , A(s(t + 1), a(t); ω - ) is the value function of the decision a(t) in ω - , and E[A(s(t + 1), a; ω - )] is the mean value of all decision value functions.
[0026] S203. Initialize the experience replay buffer:
[0027] The maximum capacity of the experience replay buffer is G, which is used to store tuples.
[0028] S3 specifically includes the following steps:
[0029] S301. Determine and execute the best decision:
[0030] At time slot t, the algorithm takes the current cached file set of the RSU (i.e., the state s(t)), combined with the input vehicle file requests, as the input of the main Q-network ω. According to the greedy algorithm, select the maximum Q presentMake a decision a (a ∈ set {0, 1}) at the value as the best decision a(t) at the current time slot t. The decision selection function can be written as
[0031]
[0032] Execute the decision a(t) to obtain the updated state s(t + 1).
[0033] S302. Increase the tuple priority and calculate the sampling probability:
[0034] When vehicle V i r (the i-th vehicle in the r-th round) sends a file request, due to the limited cache capacity of the local RSU, it cannot cache all popular files. Therefore, during the file acquisition process, the file may be in the cache space of neighboring RSUs or the MBS, resulting in different transmission delays.
[0035] Case 1: If the local RSU has the cached file q requested, then the local RSU transmits the content l to V i r , and the transmission delay calculation formula can be written as
[0036]
[0037] where z is the size of the file q, and TR r RSU,i, is the transmission rate between the local RSU and vehicle V i r .
[0038] Case 2: If the requested file q is cached in a neighboring RSU. Then during the transmission process, the neighboring RSU needs to first transmit the requested content to the local RSU through a wired link, and then the local RSU transmits the content to the requesting vehicle. At this time, the transmission delay calculation formula can be written as
[0039]
[0040] where, TR r RSU,RSU is the transmission rate of the wired link between the local RSU and the adjacent RSU.
[0041] Case 3: If the file q is neither cached in the local RSU nor in the adjacent RSU, then the MBS needs to deliver the content q to V i r , and at this time the transmission delay calculation formula can be written as
[0042]
[0043] where, TR rMBS,i is the transmission rate between the MBS and the vehicle V i r The transmission delay formula based on the above three cases, for better distinction, defines the reward function for the vehicle V
[0044] to receive the requested file q, which can be written as i r respectively
[0045]
[0046] where s n (t) is the set of files cached by the neighbor RSU. λ 1 +λ 2 +λ 3 = 1 and λ 1 <<λ 2 <λ 3 .
[0047] Therefore, the total reward function formula R(t) can be written as
[0048]
[0049] where rq i r is the number of files requested from the vehicle V i r , and n r is the total number of vehicles
[0050] The existing algorithm "Mobility-Aware Cooperative Caching in Vehicular Edge Computing Based on Asynchronous Federated and Deep Reinforcement Learning" represents the update process from s(t) to s(t + 1) as the initial tuple (s(t), a(t), R(t), s(t + 1)). However, due to the data composition limitation of the traditional DQN tuple, randomly selecting the experience data in the buffer during the training process of the target Q network may lead to repeated sampling of experience data and omission of important data. Therefore, the present invention considers using prioritized experience replay. The specific method is to add a parameter pe(s(t)) for comparing priorities to each initial tuple. The improved tuple format is (s(t), a(t), R(t), s(t + 1), pe(s(t))), where the calculation formula of pe(s(t)) can be written as
[0051] pe(s(t)) = |R(t)+γ D (Q next (s(t + 1), a(t + 1); ω - ) - Q present (s(t), a(t); ω))| + p(9)
[0052] where γD is the discount factor, and p is a very small positive constant to prevent the initial value of pe(s(t)) from being 0.
[0053] Convert the sampled preliminary tuple into (s u , a u , R u , s u ', pe(s u ))), where (s u , a u , R u , s u ', pe(s u )) represents the u-th (u = (1, 2, 3...)) tuple in the mini-batch.
[0054] Store the improved tuple in the experience replay buffer. The present invention defines the tuple priority probability, and the formula can be written as
[0055]
[0056] where ξ is the priority exponent, I is the size of the mini-batch, and when ξ = 0, it is uniform experience replay.
[0057] When the number of tuples stored in the experience replay buffer is greater than I (I is a natural number), the local RSU samples I tuples in descending order according to the tuple priority probability and defines them as a mini-batch.
[0058] S303. Calculate the target Q value and the loss function:
[0059] First, calculate the target Q value (Target Q-Value) of all tuples. Then, the target Q value TQ of tuple u u can be written as
[0060] TQ u = R u + γ D Q next (s u+1 , a u+1 ; ω - ) (11)
[0061] In order to replay experiences at different frequencies according to the importance of the experiences, thereby improving the learning efficiency, the present invention defines the priority of the experiences and samples according to the above priorities, so that important experiences have a higher sampling probability, thereby accelerating the convergence speed of the model. Therefore, the present invention introduces prioritized experience replay (PER) for replaying samples (i.e., tuples). DQN calculates the error between the target Q value and the current Q value. The larger the error value, the higher the priority of the tuple. According to the order of the priorities from high to low, the tuples in the replay buffer are extracted, and the experiences with higher update value for the main Q network can be preferentially selected, thereby accelerating the learning process. However, blindly learning high-error experiences will cause the main Q network to have the risk of overfitting to high-error samples. Therefore, it is necessary to correct the loss function of high-error samples based on the tuple priority probability. Therefore, when calculating the loss function, the present invention considers adding the importance sampling weight factor. Define the importance sampling weight coefficient as
[0062]
[0063] where τ is the compensation coefficient of the non-uniform probability, which gradually increases to 1 as the training process progresses, and G is the buffer size.
[0064] Secondly, calculate the average loss function of all tuples. Based on the squared sum of the differences between the TQ u of all tuples u and the existing Q u present The average loss function of all tuples can be written as
[0065]
[0066] S304. Update the main Q network ω:
[0067] At the end of time slot t, define the gradient of the average loss function of all tuples as According to the loss function gradient, update the main Q network ω, and its expression can be written as
[0068]
[0069] where η ω is the learning rate of the main Q network for the gradient.
[0070] The algorithm completes the iteration of time slot t and continues to enter the next time slot t + 1. Every b time slots (b is a positive integer), assign the current main Q network ω parameters to the target network ω - .
[0071] So far, input s(t) and s n (t) into the trained main Q network ω respectively to obtain s(t + 1) and sn (t + 1), which are the sets of popular files that need to be cached by the local RSU and its neighboring RSUs respectively.
[0072] S4 specifically includes the following steps:
[0073] When the algorithm completes the r-th round of model training, the current cache file allocation scheme can be obtained, based on which the local RSU and its neighboring RSUs complete the file allocation and deployment for this round.
[0074] Vehicles request files from the local RSU. If the local RSU has cached the file, the local RSU directly responds to the vehicle's file request. If the local RSU has not cached the file, it obtains the file from the neighboring RSU via a wired link and then responds to the vehicle's file request. If neither the local RSU nor its neighboring RSUs have cached the file, the macro base station responds to the vehicle's file request. After each vehicle obtains the file, the current round r ends and the next round begins.
[0075] Compared with the prior art, the beneficial effects of the present invention are as follows: (Please combine the innovative points of the present invention, and expand in detail the technical effects or technical functions generated by each innovative point in a separate paragraph; the current technical effects content is too simple and concise)
[0076] (1) The main Q-network is responsible for both selecting decisions and evaluating the Q-values of the decisions. Therefore, once the local RSU overestimates the Q-value of a certain decision a(t) at time slot t, it will wrongly select that decision. And the update of the subsequent state s(t + 1) depends on a(t), so the subsequent decision a(t + 1) may continue to overestimate the Q-value, resulting in the accumulation of Q-value errors and ultimately reducing the learning efficiency of the main Q-network. To address the Q-value overestimation problem of the Dueling-DQN algorithm, the DDDQP of the present invention uses double Q-learning (the main Q-network and the target Q-network) to separate decision selection and evaluation. Through the above separation, the algorithm of the present invention can avoid directly using the same main Q-network to select and evaluate decisions, thereby reducing the probability of Q-value overestimation, alleviating the Q-value overestimation problem, and ultimately improving the learning efficiency and accelerating the algorithm convergence rate.
[0077] (2) Traditional DQN randomly samples samples during training, resulting in problems of reusing samples and missing important samples, which reduces the generalization ability of neural network training. To address this problem, the present invention utilizes prioritized experience replay (PER). Specifically, the algorithm assigns a priority to each sample to quantify the importance of the sample. During the training process, the sample sampling probability is adjusted according to the priority level, and important samples are more likely to be sampled, which will help improve the generalization ability of the neural network. Description of the Drawings
[0078] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0079] Figure 1 Schematic diagram of the overall process of the improved RSU popular file allocation method based on deep reinforcement learning designed for the present invention.
[0080] Figure 2 Schematic diagram of the comparison results of the cache hit rate between the present invention and the existing method at different cache capacities.
[0081] Figure 3 Schematic diagram of the comparison results of the file transmission delay between the present invention and the existing method at different cache capacities. Detailed implementation manners
[0082] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0083] Embodiment 1
[0084] See Figure 1 , the technical solution provided in this embodiment is: an improved RSU popular file allocation method based on deep reinforcement learning, including the following steps:
[0085] S1 specifically includes the following steps:
[0086] S101. For urban roads, the scenario in this embodiment includes a macro base station (MBS) attached to a cloud center, a road side unit (RSU) with an edge server, and vehicles with storage capabilities. For the cloud center connected to the MBS, it is assumed that the cloud server has sufficient cache capacity and computing power, can cache all available files, and can provide file services to vehicles. The RSUs deployed on both sides of the road are responsible for forwarding the data exchanged between the MBS and vehicle users. To speed up file response, the RSU is configured with a roadside server, but the storage capacity of the RSU is limited and can only cache some files. The content requested by the vehicle is first sent to the local RSU through "vehicle-to-roadside unit" (V2R) communication. If the local RSU caches the file, it responds to the vehicle request. If the local RSU does not cache it, the vehicle needs to request from the cloud server connected to the MBS. To simplify the problem, only two roadside units (also called base stations hereinafter), namely the local RSU and its neighboring RSU, are considered in this scenario.
[0087] S102. Based on the file requests of all vehicles within the communication range of the base stations, the two RSUs determine Fc A popular file to be cached. Specifically, each vehicle submits to the local base station a user preference information (vehicle user ID, score value of each file) containing all files. After the RSU receives the above information submitted by each vehicle, it summarizes it into a file preference score table, and the element M at the x-th row and y-th column in the table x,y represents the preference value of vehicle user x for file y, and its value range is in the interval [0, 1]. 0 represents that the user's interest level in this file is 0, and 1 represents that the user's interest level in this file is 1. Based on this file preference score table, sum up for file y to obtain the popularity of the file (i.e., the total score value of file y). At this point, the RSU can obtain the popularity of all files and the set N of popular files. The storage capacity of a single RSU is limited and can only cache some files. Therefore, this embodiment considers distributing popular files between the local RSU and neighboring RSUs. This can improve the file cache hit rate and reduce the file transmission delay.
[0088] S103. In a round r, the algorithm executes Ns time slots t. At the beginning of the r-th round, the local RSU randomly selects c files from the set N of popular files for caching, and the neighboring RSU randomly caches c files from N that are not selected by the local RSU.
[0089] S2 specifically includes the following steps:
[0090] S201. Initialize the main Q network:
[0091] Take the file requests of all vehicles within the range of the local RSU as the input of the main Q network ω. Define the state s(t) and the action a(t), which represent the file set and the file update decision respectively. a(t) = 1 means that the local RSU needs to randomly select a part of the files from N to replace the local files. a(t) = 0 means that the local RSU does not need to update the local cached files. Using the main Q network, calculate the Q present value function for different file allocation decisions a (a ∈ set {0, 1}) of the RSU at time slot t as
[0092]
[0093] where V(s(t); ω) is the state value function in ω that is only related to s(t), A(s(t), a; ω) is the value function of the decision a(t) in ω, and E[A(s(t), a; ω)] is the mean value of all decision value functions.
[0094] S202. Initialize the target Q network:
[0095] The target Q network ω -It has the same network architecture as the main Q-network ω, but the target Q-network is used to calculate the impact after the RSU takes the decision a(t). Therefore, the Q value under s(t+1) is calculated, and its expression can be written as next as
[0096]
[0097] where V(s(t+1); ω - ) is the state value function in ω that is only related to s(t+1), A(s(t+1), a(t); ω - ) is the value function of the decision a(t) in ω - , and E[A(s(t+1), a; ω - )] is the mean value of all decision value functions. -
[0098] S203. Initialize the experience replay buffer:
[0099] The maximum capacity of the experience replay buffer is G, which is used to store tuples.
[0100] S3 specifically includes the following steps:
[0101] S301. Determine and execute the optimal decision:
[0102] At time slot t, the algorithm takes the current cached file set of the RSU (i.e., state s(t)), combined with the input vehicle file requests, as the input of the main Q-network ω. According to the greedy algorithm, the decision a (a ∈ set {0, 1}) with the maximum Q value is selected as the optimal decision a(t) at the current time slot t. The decision selection function can be written as present and the decision a(t) is executed to obtain the updated state s(t+1).
[0103]
[0104]
[0105] S302. Increase the tuple priority and calculate the sampling probability:
[0106] When vehicle V i r (the i-th vehicle in the r-th round) sends a file request, due to the limited local RSU cache capacity, it is impossible to cache all popular files. Therefore, during the file acquisition process, the file may be in the cache space of neighboring RSUs or the MBS, resulting in different transmission delays.
[0107] Case 1: If the local RSU has the cached file q requested, the local RSU transmits the content l to V i r , and the transmission delay calculation formula can be written as
[0108]
[0109] where z is the size of file q, and TR r RSU,i, is the transmission rate between the local RSU and vehicle V i r during the transmission process.
[0110] Case 2: If the requested file q is cached in a neighboring RSU. Then, during the transmission process, the neighboring RSU needs to first transmit the requested content to the local RSU through a wired link, and then the local RSU transmits the content to the requesting vehicle. At this time, the transmission delay calculation formula can be written as
[0111]
[0112] where, TR r RSU,RSU is the transmission rate of the wired link between the local RSU and the neighboring RSU.
[0113] Case 3: If file q is neither cached in the local RSU nor in the neighboring RSU, then the MBS needs to deliver content q to V i r , and at this time, the transmission delay calculation formula can be written as
[0114]
[0115] where, TR r MBS,i is the transmission rate between the MBS and vehicle V i r during the transmission process.
[0116] Based on the transmission delay formulas for the above three cases, for better distinction, define the reward function for vehicle V i r to receive the requested file q, which can be written as
[0117]
[0118] where, define s n (t) as the set of files cached in the neighboring RSU. λ 1 +λ 2 +λ 3 = 1 and λ 1 <<λ 2 <λ 3 .
[0119] Therefore, the total reward function formula R(t) can be written as
[0120]
[0121] where rq i r is the number of request files from vehicle V i r and n r is the total number of vehicles.
[0122] The existing algorithm "Mobility-Aware Cooperative Caching in Vehicular Edge Computing Based on Asynchronous Federated and Deep Reinforcement Learning" represents the update process from s(t) to s(t + 1) as the initial tuple (s(t), a(t), R(t), s(t + 1)). However, due to the data composition limitation of the traditional DQN tuple, randomly selecting the experience data in the buffer during the training process of the target Q network may lead to repeated sampling of experience data and missing important data. Therefore, this embodiment considers using prioritized experience replay. The specific method is to add a parameter pe(s(t)) for comparing priorities to each initial tuple, and the improved tuple format is (s(t), a(t), R(t), s(t + 1), pe(s(t))), where the calculation formula of pe(s(t)) can be written as
[0123] pe(s(t)) = |R(t) + γ D (Q next (s(t + 1), a(t + 1); ω - ) - Q present (s(t), a(t); ω))| + p(9)
[0124] where γ D is the discount factor, and p is a very small positive constant to prevent the initial value of pe(s(t)) from being 0.
[0125] Convert the sampled preliminary tuple to (s u , a u , R u , s u ', pe(s u ))), where (s u , a u , R u , s u ', pe(s u )) represents the u-th (u = (1, 2, 3...)) tuple.
[0126] Store the improved tuple in the experience replay buffer. This embodiment defines the tuple priority probability, and the formula can be written as
[0127]
[0128] where ξ is the priority exponent, I is the size of the mini-batch, and when ξ = 0, it is uniform experience replay.
[0129] When the number of tuples stored in the experience replay buffer is greater than I (I is a natural number), the local RSU samples I tuples in descending order according to the tuple priority probability and defines them as a mini-batch.
[0130] S303. Calculate the target Q-value and the loss function:
[0131] First, calculate the target Q-value (Target Q-Value) of all tuples. Then, the target Q-value TQ of tuple u u can be written as
[0132] TQ u = R u + γ D Q next (s u+1 , a u+1 ; ω - ) (11)
[0133] To replay experiences at different frequencies according to the importance of experiences, thereby improving the learning efficiency, in this embodiment, by defining the priority of experiences and sampling according to the above priorities, important experiences obtain a higher sampling probability, thus accelerating the convergence speed of the model. Therefore, this embodiment introduces prioritized experience replay (PER) for replaying samples (i.e., tuples). DQN calculates the error between the target Q-value and the current Q-value. The larger the error value, the higher the tuple priority. According to the order from high to low priority, tuples in the replay buffer are extracted, and experiences with higher update value for the main Q network can be preferentially selected, thus accelerating the learning process. However, blindly learning high-error experiences will cause the main Q network to have a risk of overfitting to high-error samples. Therefore, it is necessary to correct the loss function of high-error samples based on the tuple priority probability. Therefore, when calculating the loss function, this embodiment considers adding the importance sampling weight factor. Define the importance sampling weight coefficient as
[0134]
[0135] where τ is the compensation coefficient for non-uniform probability, which gradually increases to 1 as the training process progresses, and G is the buffer size.
[0136] Secondly, calculate the average loss function of all tuples. Based on the sum of squares of the differences between the TQ u after of all tuples u and the existing Q u present the average loss function of all tuples can be written as
[0137]
[0138] S304. Update the main Q-network ω:
[0139] At the end of time slot t, define the gradient of the average loss function for all tuples as According to the gradient of the loss function, update the main Q-network ω, and its expression can be written as
[0140]
[0141] where η ω is the learning rate of the main Q-network for the gradient.
[0142] The algorithm completes the iteration of time slot t and continues to the next time slot t + 1. Every b time slots (b is a positive integer), assign the parameters of the current main Q-network ω to the target network ω - .
[0143] So far, input s(t) and s n (t) into the trained main Q-network ω respectively to obtain s(t + 1) and s n (t + 1), which are the sets of popular files to be cached by the local RSU and neighbor RSUs respectively.
[0144] S4 specifically includes the following steps:
[0145] When the algorithm completes the r-th round of model training, the current cache file allocation scheme can be obtained, and based on this, the local RSU and neighbor RSUs complete the file allocation and deployment for this round.
[0146] Vehicles request files from the local RSU. If the local RSU has cached the file, the local RSU directly responds to the vehicle's file request. If the local RSU has not cached the file, it obtains the file from the neighbor RSU through a wired link and then responds to the vehicle's file request. If neither the local RSU nor the neighbor RSU has cached the file, the macro base station responds to the vehicle's file request. After each vehicle obtains the file, the current round r ends and the next round begins.
[0147] Embodiment 2
[0148] Simulation Scenario and Parameter Settings
[0149] This embodiment is for the urban road scenario in the vehicle-to-everything network. The experiment uses python3.8 as the simulation tool to conduct a comparative simulation of the improved RSU popular file allocation method based on deep reinforcement learning and the existing RSU popular file allocation method based on Dueling-DQN in this embodiment.
[0150] Improved RSU Popular File Allocation Method Based on Deep Reinforcement Learning: Aiming at the technical problem of mis-caching popular files in the existing popular file allocation schemes, an improved RSU popular file allocation method based on deep reinforcement learning is implemented in the vehicle-to-everything (V2X) network. In this embodiment, double Q-learning and experience prioritized replay are introduced, and an improved RSU popular file allocation method based on deep reinforcement learning is proposed. Accordingly, the RSU allocates and caches files, achieving the goal of improving the cache hit rate and reducing the file transmission delay.
[0151] RSU Popular File Allocation Method Based on Dueling-DQN: Based on the Dueling-DQN deep reinforcement learning algorithm, the RSU aims to maximize the cache hit rate and minimize the transmission delay, allocates popular files in the local RSU and neighbor RSU cache spaces, and responds to the file requests of vehicles within the region.
[0152] In the experimental scenario, this embodiment involves three communication modes: vehicle-to-vehicle (V2V), vehicle-to-roadside unit (V2R), and vehicle-to-macro base station (V2M). Therefore, the V2X communication system adopts a cellular V2X (C-V2X) architecture based on the 3rd Generation Partnership Project (3GPP). The specific parameter configurations of the experiment are listed in Table 1 in detail. The MovieLens 1M dataset is randomly distributed to each vehicle as its local dataset. In terms of data processing, each vehicle randomly selects 90% of its local dataset as training data, and the remaining 10% is used as test data. In each round of the experiment, the vehicles randomly select some files from the test dataset as the target files they request. Table 1 shows the experimental parameter values.
[0153] Table 1 Parameters of the RSU Popular File Allocation Method
[0154]
[0155]
[0156] The experiment uses the cache hit rate and content transmission delay as performance metrics.
[0157] Define the cache hit rate: If the content requested by a vehicle is cached in the local RSU, the vehicle can directly obtain the content from the local RSU, and this situation is called a cache hit; otherwise, it is called a cache miss. Define the cache hit rate as the ratio of the number of times a vehicle successfully obtains the requested file from the local RSU to the total number of requests.
[0158] The definition of file transfer delay is the ratio of the total transfer delay of all vehicles to obtain the requested file to the number of requests. Since the local RSU cache capacity is limited and cannot cache all popular files. Therefore, during the process of obtaining files, the files may exist in the cache spaces of adjacent RSUs or MBSs, resulting in different transfer delays.
[0159] Figure 2 It shows the comparison of cache hit ratios of the two methods under different cache capacities... As can be seen from the figure, the cache hit ratios of both schemes increase with the increase of cache capacity. This is because after the cache capacities of the local RSU and adjacent RSUs are increased, the number of popular files that can be cached increases, and the content requests of vehicles in the area are more easily satisfied. In addition, due to the problem of overestimation of Q values in the RSU popular file allocation method based on Dueling-DQN, unnecessary files are cached, while the Q value of the embodiment of this algorithm can be maintained at a reasonable level, and files can be cached accurately to ensure the RSU to respond to file requests.
[0160] Figure 3 It shows the comparison of file transfer delays of the two methods under different cache capacities. As can be seen from the figure, the file transfer delays of both schemes gradually decrease with the increase of cache capacity. This is because after the cache capacities of the local RSU and adjacent RSUs are increased, the number of popular files that can be cached increases, and the file requests of vehicles in the area are more easily satisfied. In addition, since the RSU popular file allocation method based on Dueling-DQN randomly samples empirical data, problems of repeated sampling or missed sampling are caused. While the embodiment of this algorithm introduces Prioritized Experience Replay (PER) and samples according to the importance of empirical data, which can make more effective use of key empirical data, thereby improving the generalization ability. The present invention can optimize the deployment of popular files among RSUs, making it easier for vehicles to obtain requested files from local RSUs. Therefore, the file transfer delay of this method is lower than that of the benchmark method.
[0161] Comprehensive analysis shows that from Figure 2 and Figure 3 it can be seen that compared with the RSU popular file allocation method based on Dueling-DQN, the present invention effectively enhances the cache hit ratio and shortens the file transfer delay, bringing a better solution for improving the performance of file requests in the vehicle-to-everything network.
[0162] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An improved RSU popular file allocation method based on deep reinforcement learning, characterized in that: The following steps are involved: S1: In each round, based on the file requests of all vehicles within the communication range, the two RSUs determine F c popular files to be cached; S2: RSU initializes the main Q network, target Q network and experience replay buffer; S3: RSU trains the target Q network and updates the main Q network; S4: Based on the results of deep reinforcement learning, the local RSU distributes popular files to neighbor RSUs to respond to vehicle file requests.
2. According to claim 1, an improved RSU popular file allocation method based on deep reinforcement learning is characterized in that: The step S1 specifically includes the following: S101. For urban roads, the scenario includes a macro base station MBS with a cloud center, a roadside unit RSU with an edge server, and a vehicle with storage capacity. For the cloud center connected to the MBS, it is assumed that the cloud server has sufficient cache capacity and computing power, caches all available files and can provide file services to vehicles. The RSUs deployed on both sides of the road are responsible for forwarding the data interacted between the MBS and vehicle users. The content requested by the vehicle is first sent to the local RSU through the vehicle-to-roadside unit V2R communication. If the local RSU has cached the file, it responds to the vehicle request. If the local RSU has no cache, the vehicle needs to request the cloud server connected to the MBS. In this scenario, only two roadside units, the local RSU and the neighboring RSU, are considered, also known as base stations; S102: Based on the file requests of all vehicles within the communication range of the base station, the two RSUs determine F c popular files to be cached, each vehicle submits a user preference information containing all files to the base station, that is, the vehicle user ID and the score of each file. After receiving the above information submitted by each vehicle, the RSU summarizes it into a file preference score table. The element M in the x row and y column of the table x,y Represents the preference value of vehicle user x for file y, and its value range is [0,1], where 0 represents the user's interest in the file is 0, and 1 represents the user's interest in the file is 1. Based on the file preference score table, the sum of file y is summarized. That is, the popularity of the file is obtained, that is, the total score of file y. At this point, RSU obtains the popularity of all files and the popular file set N. The storage capacity of a single RSU is limited and can only cache some files. Consider distributing popular files between the local RSU and the neighboring RSU to improve the file cache hit rate and reduce the file transmission delay; S103. In a round r, the algorithm executes Ns time slots t. At the beginning of the rth round, the local RSU randomly selects c files from the popular file set N to cache, and the neighbor RSU randomly caches c files from N that are not selected by the local RSU.
3. According to claim 2, an improved RSU popular file allocation method based on deep reinforcement learning is characterized in that: The step S2 specifically includes the following: S201, initialize the main Q network: The file requests of all vehicles within the local RSU range are used as the input of the main Q network ω. The state s(t) and action a(t) are defined to represent the file set and file update decision, respectively. a(t) = 1 means that the local RSU needs to randomly select a part of the files from N to replace the local files. a(t) = 0 means that the local RSU does not need to update the local cache files. Using the main Q network, calculate the Q of different RSU file allocation decisions a (a∈set {0,1}) in time slot t present The value function is: Where V(s(t); ω) is the state value function related only to s(t) in ω, A(s(t), a; ω) is the value function of decision a(t) in ω, and E[A(s(t), a; ω)] is the mean of all decision value functions. S202, initializing the target Q network: Target Q network ω - The target Q network is used to calculate the impact of the RSU decision a(t) after the network architecture is the same as the main Q network ω. Therefore, the target Q network is used to calculate the Q at s(t+1) next The value is expressed as: Among them, V(s(t+1);ω - ) is ω - The state value function only related to s(t+1) is A(s(t+1),a(t);ω - ) is ω - The value function of decision a(t) in , E[A(s(t+1),a;ω - )] is the mean of all decision value functions; S203, initializing the experience playback buffer: The maximum capacity of the experience replay buffer is G, which is used to store tuples.
4. According to the improved RSU popular file allocation method based on deep reinforcement learning in claim 3, it is characterized in that: The step S3 is specifically as follows: S301. Determine and implement the best decision: At time slot t, the algorithm takes the current RSU cache file set, i.e., state s(t), and the input vehicle file request as the input of the main Q network ω, and selects the maximum Q according to the greedy algorithm. present The decision a(a∈set {0,1}) under the value is the best decision a(t) under the current time slot t. The decision selection function is Execute decision a(t) to get the updated state s(t+1); S302, increase the tuple priority and calculate the sampling probability: When the vehicle V i r , i.e., after the i-th car in the r-th round sends a file request, due to the limited cache capacity of the local RSU, it is unable to cache all popular files. In the process of obtaining files, there are files in the cache space of neighboring RSUs or MBSs, resulting in different transmission delays; Case 1: If the local RSU has the requested file q cached, the local RSU transmits content l to V i r , the transmission delay calculation formula is written as Where z is the size of file q, TR r RSU,i , is the local RSU and vehicle V i r The transmission rate between Case 2: If the requested file q is cached in the neighbor RSU, then during the transmission process, the neighbor RSU needs to first transmit the requested content to the local RSU through a wired link, and the local RSU then transmits the content to the requesting vehicle. In this case, the transmission delay calculation formula is Among them, TR r RSU,RSU is the transmission rate of the wired link between the local RSU and the adjacent RSU; Case 3: If file q is neither cached in the local RSU nor in the neighboring RSU, the MBS needs to deliver the content q to V i r , the transmission delay calculation formula is: Among them, TR r MBS,i MBS and vehicle V i r The transmission rate between Based on the transmission delay formulas of the above three cases, in order to better distinguish, define vehicle V i r The reward functions for receiving the requested file q are Among them, the definition s n (t) is the set of files cached by neighboring RSUs, λ1+λ2+λ3=1 and λ1<<λ2<λ3; Therefore, the total reward function formula R(t) is where rq i r Is from vehicle V i r The number of requested files, n r is the total number of vehicles; The update process from s(t) to s(t+1) is represented as the initial tuple (s(t), a(t), R(t), s(t+1)). Due to the data composition limitation of the traditional DQN tuple, the experience data in the buffer is randomly selected during the training of the target Q network, resulting in repeated sampling of the experience data and missing important data. Considering the use of priority experience replay, a parameter pe(s(t)) for comparing priorities is added to each initial tuple. The improved tuple format is (s(t), a(t), R(t), s(t+1), pe(s(t))), where the calculation formula of pe(s(t)) is on(s(t))=|R(t)+γ D (Q next (s(t+1),a(t+1);ω - )-Q present (s(t),a(t);ω))|+p(9) where γ D is the discount factor, p is a very small positive constant, which prevents the initial value of pe(s(t)) from being 0; Convert the sampled preliminary tuple into (s u ,a u ,R u ,s u ',pe(s u )), where (s u ,a u ,R u ,s u ',pe(s u )) represents the uth (u=(1,2,3…))th tuple in the mini-batch; The improved tuple is stored in the experience replay buffer, and the tuple priority probability is defined as follows: Where ξ is the priority index, I is the size of the mini-batch, and ξ = 0 is uniform experience replay; When the number of tuples stored in the experience replay buffer is greater than I, where I is a natural number, the local RSU samples I tuples in descending order according to the tuple priority probability and defines it as a mini-batch; S303, calculate the target Q value and loss function: First, calculate the target Q-value of all tuples, then the target Q-value TQ of tuple u u for TQ u =R u +γ D Q next (s u+1 ,a u+1 ;ω - )(11) Consider adding the importance sampling weight factor and define the importance sampling weight coefficient φ u for Among them, τ is the compensation coefficient of non-uniform probability, which gradually increases to 1 as the training progresses, and G is the buffer size; Secondly, the average loss function of all tuples is calculated, based on the TQ of all tuples u. u After and existing Q u present The sum of squared differences of all tuples is S304, update the main Q network ω: At the end of time slot t, the average loss function gradient of all tuples is defined as According to the loss function gradient, update the main Q network ω, which is expressed as Among them, η ω is the learning rate of the main Q network for the gradient; The algorithm completes the iteration of time slot t and continues to the next time slot t+1. After every b time slots, b is a positive integer, the current main Q network ω parameter is assigned to the target network ω - ; Substitute s(t) with s n (t) are respectively input into the trained main Q network ω to obtain s(t+1) and s n (t+1), which are the popular file sets that the local RSU and the neighboring RSU need to cache respectively.
5. According to the improved RSU popular file allocation method based on deep reinforcement learning according to claim 4, it is characterized in that: The step S4 is specifically as follows: When the algorithm completes the rth round of model training, it obtains the current cache file allocation plan, based on which the local RSU and neighboring RSU complete the file allocation deployment for this round; The vehicle requests a file from the local RSU. If the local RSU has cached the file, the local RSU directly responds to the vehicle's file request. If the local RSU has not cached the file, it obtains the file from the neighboring RSU through a wired link and then responds to the vehicle's file request. If neither the local RSU nor the neighboring RSU has cached the file, the macro base station responds to the vehicle's file request. After each vehicle obtains the file, the current round r ends and the next round begins.