A segmented caching method that combines user preferences and objective popularity
By combining user preferences and objective popularity into a segmented caching approach, and by optimizing the caching strategy using collaborative filtering and adaptive particle swarm optimization algorithms, the problem of limited device cache space is solved, resulting in a more efficient cache hit rate and cellular traffic offloading.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTH CHINA NORMAL UNIV
- Filing Date
- 2022-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Given limited device cache space, how can we more efficiently offload cellular traffic and improve the system's average cache hit rate, taking into account users' unique preferences and diverse content needs?
A segmented caching approach combining user preferences and objective popularity is adopted. The collaborative filtering algorithm predicts the user's preference for files that have not been rated. The file topic classification and historical request records are combined to build a probability model of user requests for files. The adaptive particle swarm optimization algorithm is used to optimize the cache hit rate. The user cache space is segmented to cache popular and personalized files.
It improves the system's cache hit rate, effectively utilizes the device's cache space, meets users' personalized needs, and reduces the pressure on base stations during peak network periods.
Smart Images

Figure CN115904249B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile communication technology, and more specifically, to a segmented caching method that combines user preferences and objective popularity. Background Technology
[0002] With the rapid development of modern technology, mobile data is growing exponentially, putting enormous pressure on base stations (BS). Therefore, to provide users with higher quality services, we need to use alternative methods to handle the ever-increasing data traffic and reduce the pressure on cellular networks. Device-to-device (D2D) communication is a highly popular emerging communication technology in recent years. It is defined as direct, short-range communication between mobile users without the need for a base station. When a device requests content, if a nearby device has the requested content cached, the data can be directly transmitted between the devices via D2D communication, effectively offloading cellular traffic and significantly reducing peak data traffic demands. Furthermore, D2D technology offers advantages such as high spatial efficiency, high power efficiency, and high system performance.
[0003] However, in reality, a device's cache space is limited, while the demand for requested content is diverse. Therefore, a device cannot indiscriminately cache all content; what content should be pre-cached to more efficiently offload cellular traffic becomes a question worth considering. Summary of the Invention
[0004] The purpose of this invention is to improve the average cache hit rate of the system and more efficiently offload cellular traffic when the device's cache space is limited. To achieve this purpose, the following technical solution is adopted:
[0005] A segmented caching method that combines user preferences and objective popularity, the segmented caching method comprising the following steps:
[0006] Obtain user ratings and topic preference ratings for each file in the base station file library;
[0007] The probability of a user requesting each file in the base station's file library is obtained from the rating and topic preference ratings, and the request probabilities are ranked first. These files form the first file library;
[0008] The popularity of each file is determined based on the request probability of all users at the base station for each file in the base station's file library. All files are then ranked by popularity, and the top-ranked files are... The files form a second file library, and files that are duplicates of the first file library are removed.
[0009] The user's cache space is divided into a first space that caches only files from the first file library and a second space that caches only files from the second file library.
[0010] A further improvement is that obtaining the user's rating for each file in the base station's file library includes using a collaborative filtering algorithm to obtain the similarity between users and then predicting the rating for files that the user has not rated.
[0011] A further improvement lies in the specific method of using collaborative filtering algorithms to obtain the similarity between users and then predicting the rating of documents that users have not yet rated:
[0012] Calculate target users Other users Similarity between Find the target user User sets with similar behaviors in document rating and classification First, based on the Pearson similarity calculation formula, it is initially determined that... ,Right now:
[0013]
[0014] in, On behalf of users For the file The rating and rating On behalf of users For the file The rating and rating On behalf of users Average score for all files On behalf of users Average score for all files The number of files in the library. Represents the total number of users;
[0015] Based on formula (1), another scoring parameter is proposed. :
[0016]
[0017]
[0018] in, Indicates user and users The number of publicly rated documents and Representing users respectively Minimum and maximum values for the number of publicly rated files shared with other users;
[0019] Set a threshold To filter out similarity Users whose threshold is higher than this threshold, when ≤ , indicating user and users They have similar scoring and rating behaviors, namely And expressed as:
[0020]
[0021] Based on the obtained user set For target users Documents that have not been rated will be assigned a rating or score, and the predicted rating will be represented as follows:
[0022]
[0023] in, {0,1,2,3,4...Z}, The higher the value, the better the user experience. For the file The more you like it, =0 indicates user No file Rating and rating.
[0024] A further improvement is that the method for obtaining the user's topic preference rating for each file in the base station file library includes obtaining the user's topic preference rating for each file in the base station file library based on the file topic classification and the user's historical request records.
[0025] A further improvement lies in obtaining the user's topic preference rating for each file in the base station file library based on file topic classification and the user's historical request records, including:
[0026] Assuming the entire file repository contains a total of T topics, use sets... The theme is indicated. and documents The attribute function is Its value depends on the file. Is it included in the topic? Li, that is:
[0027]
[0028] Assumption Indicates user For the subject of the document The preference function, which is expressed by the mutual information formula, assumes... By user Historical requests to record decisions, The definition is as follows:
[0029]
[0030] in, It is mutual information. For includes A collection of all files, For users Historical request records, For users The historical request record contains The probability of topic files, To exist throughout the entire file library The probability of a topic file;
[0031] user For the file Interest level parameters Defined as:
[0032]
[0033] Received ∈[−1,1], The smaller the value, the better for the user. For the file The less interested, according to Get users For the file Theme Preference Rating The definition is as follows:
[0034] .
[0035] A further improvement is that the method for obtaining the probability of a user requesting each file in the base station file library from the rating and topic preference rating includes:
[0036] Assumption , Representing users respectively The lowest and highest ratings for files in the file library, parameters Indicates user The difference in scoring and rating different documents is defined as:
[0037]
[0038] Therefore, users For the file Request probability The expression is as follows:
[0039]
[0040] Used to adjust for differences in file preferences among users who give the same file different scores;
[0041] Assumption , Representing users respectively The parameters are the lowest and highest topic preference ratings for files in the file library. Indicates user The difference in the rating of different document topic preferences is defined as:
[0042]
[0043] Therefore, users For the file Request probability The expression is as follows:
[0044]
[0045] We get users For the file Total request probability as follows:
[0046]
[0047]
[0048]
[0049] in, Indicates user For the file Request probability The weight.
[0050] A further improvement lies in, based on all users' opinions on the file The request probability is used to obtain the file. The expression for popularity is:
[0051] .
[0052] A further improvement is that the segmented caching method also includes:
[0053] Construct an average cache hit rate optimization problem; solve the average cache hit rate optimization problem using an adaptive particle swarm optimization algorithm to obtain the optimized average cache hit rate.
[0054] A further improvement is that the method for constructing the average cache hit rate optimization problem includes: assuming a cache hit rate Defined as user Request file At that time, the probability of finding and obtaining the target file within its own cache space or within the range of successfully linked D2D links in the surrounding area, assuming that when a user Request file At that time, the user had already cached the file. Then, the requested file is retrieved from its own cache space, and the probability of the event occurring is... Suppose a user Request file , Not in its own cache When that happens, a request is sent to the user within the D2D communication range; if a cache exists... Users then obtain information through the D2D link. The probability of an event occurring And users Request file It is also a probabilistic event, let's say... That is, the probability of its occurrence; the cache hit rate is as follows:
[0055]
[0056] in, Indicates user The activity level, i.e., the user The probability of issuing a file request indicates the user's... The activity level is defined as:
[0057]
[0058] If user For the file If there is active scoring and rating behavior, then otherwise ;
[0059] set up On behalf of users Cached files The probability of the user Found in its own cache space probability That is equivalent to user Cached files probability ,Right now:
[0060]
[0061] There must be at least one user within the area where the D2D link can be successfully established. = Cached users Requested document The probability is:
[0062]
[0063] in, represent Cached files The probability, This represents the user within a time segment. and users The probability of successfully performing a D2D link;
[0064] So users Acquired via D2D link The probability is used Represented as:
[0065]
[0066] in, This represents the probability of a successful D2D link. The system's average cache hit rate is:
[0067]
[0068] Assuming user The sizes of the first space and the second space are respectively , Represents the first document library, Representing the second file library, the number of duplicate files removed is ;
[0069] When the file At that time, the user Cached files probability Represented as:
[0070]
[0071] When the file At that time, the user Cached files probability Represented as:
[0072]
[0073] Therefore, users cached files probability for:
[0074]
[0075] Therefore, to maximize the average cache hit rate, it is necessary to find the optimal segmentation point of both the base station file library and the user cache space. Hence, the average cache hit rate optimization problem is defined as follows:
[0076]
[0077] in,
[0078] At the same time, both decision variables are discrete variables due to the influence of the file library segmentation point and the user cache space segmentation point. Under the influence of these factors, the problem of formula (25) becomes a non-convex nonlinear constraint optimization problem.
[0079] A further improvement is that the method for solving the average cache hit rate optimization problem using the adaptive particle swarm optimization algorithm includes: dynamically updating the inertia weight between the minimum and maximum inertia values during each search iteration of the particle swarm optimization algorithm. The dynamic update formula for the inertia weight is as follows:
[0080]
[0081]
[0082] in, It is the first iteration , Where N is the maximum number of iterations, and N is the population size. It is the optimal value of the current fitness function. This is the current location. The fitness function value, It is the minimum value of the current inertia weight. It is the maximum value of the current inertia weight;
[0083] Formula (26) calculates the average difference between the fitness function value of all particles in the current iter-th iteration and the historical best fitness function value, which is used to represent the current search status of the particle swarm. Formula (27) uses... Parameters to dynamically adjust inertia weights , exist[ , The search range is updated within a certain range and varies with the particle's fitness. Therefore, at the beginning of the iteration, the particle swarm search range is very large and needs to be dynamically increased. The value facilitates global search, but needs to be dynamically reduced when the particle swarm locks onto the target area. This value facilitates local search, thus enabling a faster and more accurate discovery of the optimal solution for the target.
[0084] The beneficial effects of this invention are as follows: Based on the collaborative filtering algorithm, this invention predicts the rating of unrated files based on existing ratings, then obtains user preferences for different topics through historical request records, and finally combines the ratings and topic preferences to model each user's preferences for each file. This invention considers the unique preferences of each user from multiple aspects, and also takes into account the problem of inconsistency between the rated files and the browsed files.
[0085] Based on segmented caching, this invention innovatively combines user preferences with objective popularity. While ensuring that the caching of user terminals takes into account popular files, it also takes into account the unique preferences and needs of each user, thereby achieving a better system cache hit rate.
[0086] The optimization problem of the average cache hit rate of the system under the caching method of the present invention is a non-convex nonlinear constrained optimization problem. The present invention is based on the classic particle swarm optimization (PSO) algorithm and combines the actual caching method with an adaptive particle swarm optimization algorithm PSO-MPhit that dynamically updates the inertia factor in the standard PSO. In this algorithm, the inertia factor actively adjusts its value according to the particle swarm search progress, which greatly improves the convergence speed of the algorithm. Attached Figure Description
[0087] Figure 1 This is a flowchart of the segmented caching method of the present invention;
[0088] Figure 2 Fixed segmentation points After =2, Segmentation points along with the file library The change graph;
[0089] Figure 3 This is a graph showing how the cache hit rate changes as the number of users changes;
[0090] Figure 4 A graph showing the change in the ratio of the runtime required to simulate a user personalization preference strategy under different file counts and user counts to the runtime of the PSO-MPhit cache of this invention.
[0091] Figure 5 A graph showing how the cache hit rate changes as the probability of a successful D2D link changes;
[0092] Figure 6 This graph shows the change in cache hit rate as the cache size increases.
[0093] Figure 7 This graph shows how the cache hit rate changes as the number of files increases.
[0094] Figure 8 This graph shows the change in cache hit rate as the contact frequency parameter increases. Detailed Implementation
[0095] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0096] The cellular D2D communication network model is as follows: each fixed-size cell contains one base station (BS) and U uniformly distributed mobile devices. In dedicated mode, users transmit content via the BS and via D2D links using non-overlapping orthogonal radio resources. The two links do not interfere with each other, and there is no co-channel interference between D2D links. We assume that each mobile device has a fixed buffer space of the same size, M, the BS's file library is F, and the probability of a successful D2D link between two mobile devices is .
[0097] For two users to engage in D2D communication, the physical distance between them needs to be kept within a fixed value. Based on existing research, it's known that the physical distance between two users is related to their contact time; the longer the contact time, the shorter the physical distance. The probability of two users establishing a D2D link can be simulated using link time and link interval as parameters. Link time is the time it takes for two users to establish a D2D link, and the link interval is the time between two consecutive D2D links. We assume that both connection time and interval time follow an exponential distribution. These two parameters directly simulate the contact opportunity for two users to establish a D2D link, rather than simulating their geographical location, and can be directly analyzed from a large number of real historical contact trajectories.
[0098] We assume users respectively and users Connection time parameters and interval time parameters between these two users Indicates in Did these two users connect before the specified time? Based on Markov chains, we can determine whether the users... and users Probability of successful D2D linking for:
[0099] .
[0100] like Figure 1As shown in the figure, this invention proposes a segmented caching method that combines user preferences and objective popularity to improve the system cache hit rate. The segmented caching method includes the following steps: Step S1, obtaining the user's rating and topic preference rating for each file in the base station file library; Step S2, obtaining the user's request probability for each file in the base station file library from the rating and topic preference rating, and ranking the request probabilities of the top-ranked files. The first file library is formed from these files; Step S3: Based on the request probability of each file in the base station file library by all users of the base station, the popularity of each file is obtained, and all files are ranked by popularity. The files form a second file library, and files that are duplicates of the first file library are removed; in step S4, the user's cache space is divided into a first space that caches only files from the first file library and a second space that caches only files from the second file library.
[0101] Specifically, we exported a rating matrix for each user's file based on the Movielens dataset. The matrix values It is used to represent users For the file The degree of liking {0,1,2,3,4,5}. The higher the value, the better the user experience. For the file The more you like it. Among them, when =5 indicates the user One of my favorite files is ;when = 0 indicates user No file Rating and rating. In real life, there are a vast number of files, and users cannot possibly download and view them all. Therefore, it's impossible for users to rate every single file. This means... Typically, the data is relatively sparse. We utilize a collaborative filtering algorithm to obtain the similarity between users and then predict the ratings of unrated documents, thereby predicting the user's level of liking for unrated documents. The specific method for using collaborative filtering to obtain the similarity between users and then predicting the ratings of unrated documents is as follows:
[0102] Collaborative filtering is a neighborhood-based algorithm, so we need to calculate the target user. Other users Similarity between Find the target user User sets with similar behaviors in document rating and classification First, based on the Pearson similarity calculation formula, it is initially determined that... ,Right now:
[0103]
[0104] in, On behalf of users For the file The rating and rating On behalf of users For the file The rating and rating On behalf of users Average score for all files On behalf of users Average score for all files The number of files in the library. Represents the total number of users.
[0105] Formula (1) above does not consider the number of rated documents. Therefore, when the number of publicly rated documents is small, the similarity between the two users calculated by formula (1) will be very high, resulting in low reliability of the user similarity. For example, it is unreasonable to conclude that two users with only one publicly rated document are highly similar by formula (1). This invention proposes a rating parameter based on formula (1). :
[0106]
[0107]
[0108] in, Indicates user and users The number of publicly rated documents and Representing users respectively Minimum and maximum values for the number of publicly rated files shared with other users;
[0109] Set threshold To filter out similarity Users whose threshold is higher than this threshold, when ≤ , indicating user and users They have similar scoring and rating behaviors, namely And expressed as:
[0110]
[0111] Finally, based on the obtained user set For target users Documents that have not been rated will be assigned a rating or score, and the predicted rating will be represented as follows:
[0112]
[0113] in, {0,1,2,3,4,5}, The higher the value, the better the user experience. For the file The more you like it, =0 indicates user No file Rating and rating.
[0114] Since a user's rating of a file does not necessarily mean that the user has viewed the file, and the rating score does not fully reflect the user's level of liking for the file, we should also consider the user's actual request history to more accurately predict user preferences. User preferences for files are inextricably linked to the file type.
[0115] Therefore, the method for obtaining a user's topic preference rating for each file in the base station file library according to the present invention includes obtaining the user's topic preference rating for each file in the base station file library based on file topic classification and the user's historical request records. Specifically, obtaining the user's topic preference rating for each file in the base station file library based on file topic classification and the user's historical request records includes: assuming that all files in the file library contain a total of T topics, using a set... The theme is indicated. and documents The attribute function is Its value depends on the file. Is it included in the topic? Li, that is:
[0116]
[0117] Assumption Indicates user For the subject of the document The preference function, which is expressed by the mutual information formula, assumes... By user Historical requests to record decisions, The definition is as follows:
[0118]
[0119] in, It is mutual information. For includes A collection of all files, For users Historical request records, For users The historical request record contains The probability of topic files, To exist throughout the entire file library The probability of a topic file.
[0120] user For the file Interest level parameters Defined as:
[0121]
[0122] Based on the above formula, we can obtain user information by classifying file topics and analyzing user historical request records. For the file Interest level parameters ∈[−1,1], The smaller the value, the better for the user. For the file The less interested one is, the easier it is to adjust parameters. In the application of predictive rating mechanisms, we based on Get users For the file Theme Preference Rating The definition is as follows:
[0123] .
[0124] Define users by combining a rating matrix and historical request records. For the file Request probability A user's historical request records, such as browsing time, likes, and favorites, can reflect user preferences to some extent. Specifically, the method for obtaining the probability of a user requesting each file in the base station file library from the rating and topic preference ratings includes: assuming... , Representing users respectively The lowest and highest ratings for files in the file library, parameters Indicates user The difference in scoring and rating different documents is defined as:
[0125]
[0126] Users are evaluated through document rating. For the file Request probability The expression is as follows:
[0127]
[0128] This is used to adjust for differences in file preferences between users who give the same file different scores.
[0129] Similarly, assuming , Representing users respectively The parameters are the lowest and highest topic preference ratings for files in the file library. Indicates user The difference in the rating of different document topic preferences is defined as:
[0130]
[0131] Therefore, users For the file Request probability The expression is as follows:
[0132]
[0133] From the above formula, we can obtain the user's... For the file Total request probability as follows:
[0134]
[0135]
[0136]
[0137] in, Indicates user For the file Request probability The weight of each file based on all users' opinions. Based on the request probability, we obtain the file. The expression for popularity is:
[0138] .
[0139] The purpose of using D2D caching is to reduce the pressure on base stations during peak network periods by obtaining the target file from the user's own cache space or through D2D communication, instead of through the base station. Therefore, we can use the system's average cache hit rate over any given time segment. As a performance indicator of the present invention.
[0140] The segmented caching method further includes constructing an average cache hit rate optimization problem, solving the average cache hit rate optimization problem using an adaptive particle swarm optimization algorithm, and obtaining an optimized average cache hit rate. The method for constructing the average cache hit rate optimization problem includes:
[0141] Assuming cache hit rate Defined as user Request file At that time, the probability of finding and obtaining the target file within its own cache space or within the range of successfully linked D2D links in the surrounding area, assuming that when a user Request file At that time, the user had already cached the file. Then, the requested file is retrieved from its own cache space, and the probability of the event occurring is... Suppose a user Request file , Not in its own cache When that happens, a request is sent to the user within the D2D communication range; if a cache exists... Users then obtain information through the D2D link. The probability of an event occurring And users Request file It is also a probabilistic event, let's say... That is, the probability of its occurrence. This probability can be obtained by the method described above, which uses scoring and topic preference ratings to obtain the probability of a user requesting each file in the base station file library. Specifically, the cache hit rate is defined as follows:
[0142]
[0143] in, Indicates user The activity level, i.e., the user The probability of issuing a file request indicates the user's... The activity level is defined as:
[0144]
[0145] If user For the file If there is active scoring and rating behavior, then otherwise .
[0146] set up On behalf of users Cached files The probability of the user Found in its own cache space probability That is equivalent to user Cached files probability ,Right now:
[0147]
[0148] There must be at least one user within the area where the D2D link can be successfully established. = Cached users Requested document The probability is:
[0149]
[0150] in, represent Cached files The probability, This represents the user within a time segment. and users The probability of successfully performing a D2D link;
[0151] So users Acquired via D2D link The probability is used Represented as:
[0152]
[0153] in, This represents the probability of a successful D2D link.
[0154] The system average cache hit rate is:
[0155]
[0156] single user terminal The cache space M will be temporarily divided into two segments. That is, the first space and the second space, whose sizes are respectively Then, all files in file library F are sorted by objective popularity, with the most popular files listed first. The first file library is created by copying the files. Next, we will access the file library. All files according to user The degree of preference for the document Sort the files from highest to lowest, and then sort the first few... File copying creates a second file library If a file exists at this time... At the same time belong to and Then put the file from Remove from the list, and record the total number of items removed from the list. Number of files removed The number of files is
[0157] Let users First Space Cache only from The file is cached randomly with equal probability. At that time, the user Cached files probability Represented as:
[0158]
[0159] in, A file library composed of files ranked by objective popularity, which users use for... The preferences are highly similar, with most users showing some interest in the files in this library. Compared to the original file library F, The number of files is greatly reduced, and retaining only popular files ensures more efficient use of user cache space. Popular files have a certain probability of being cached by each user, and even if a user does not cache popular files, they can still obtain files from surrounding users through the D2D link. Conversely, even if a user is not interested in popular files, popular files cached in their local space can still be used by other users in the D2D communication network, effectively utilizing the advantages of D2D collaborative caching.
[0160] When the file At that time, the user This file will definitely be cached. From this, we can conclude that: when the file At that time, the user Cached files probability Represented as:
[0161]
[0162] The user's cache space is divided into two segments. , By leveraging the homogeneity of content preferences, cached files can be shared by multiple users in the vicinity via D2D collaborative communication. It places greater emphasis on the needs of the majority of users in the surrounding area. However, considering the heterogeneity of user preferences, in order to satisfy the unique preferences of individual users, This is a cache from This is a file for users. A personalized file library that only caches files. Users prioritize files based on their preferences, ensuring that users can benefit from the local cache.
[0163] Therefore, users cached files probability for:
[0164]
[0165] Therefore, to maximize the average cache hit rate, it is necessary to find the optimal segmentation point of both the base station file library and the user cache space. Hence, the average cache hit rate optimization problem is defined as follows:
[0166]
[0167] in, .
[0168] Simultaneously influenced by the file library segmentation point and the user cache space segmentation point, both decision variables are discrete variables, making the problem in formula (25) a non-convex nonlinear constrained optimization problem. A local optimum is not equivalent to a global optimum. Because the proof of the function's properties is complex, and it is only used to find the maximum... To demonstrate the advantages of the caching strategy proposed in this paper, we briefly prove the non-convex property of (25), thereby explaining the necessity of using the PSO algorithm to improve computation speed and reduce time complexity in the following sections. The proof is as follows:
[0169] Lemma 1: It is convex if and only if , For any The t is convex.
[0170] The above lemma shows that when a high-dimensional function is convex, it can be equivalent to the superposition of an infinite number of one-dimensional convex functions. This function remains convex in its domain if and only if it is restricted to an arbitrary straight line. Let's assume that the problem in equation (25) is a convex function.
[0171] Assumption 1: Formula (25) A variable is the file library segment point. Segmentation points of user cache space A bivariate convex function. Then when When fixed to a constant, in the domain In this context, the problem in formula (25) becomes a univariate convex function.
[0172] Proof 1: For a univariate convex function, taking two points x and y on the convex domain, the value of its convex combination should be less than or equal to the value of the convex combination itself.
[0173]
[0174] Because the formulas in this invention have complex constraints and the variables take discrete values, for ease of explanation, we will directly use fixed... The graph of the values of the subsequent formula (25) is given, as follows: Figure 2 As shown, in formula (25) the variables After fixing, within the domain... The curve formed by uniform sampling is obviously a non-convex function graph, which contradicts the assumption in hypothesis 1 that formula (25) is a multidimensional convex function. Reviewing the properties of multidimensional convex functions in Lemma 1, we can conclude that formula (25) is a high-dimensional non-convex function.
[0175] To solve this problem, i.e., to find the globally optimal solution to formula (25), we need to consider that the segmentation points of the total file library and the segmentation points of the user terminal cache space are simultaneously optimal. Finding the optimal solution to the above formula is quite challenging because the invented segmented caching strategy uses collaborative filtering algorithms and a large real dataset, and has a large number of user ratings and an uncertain total number of files and user terminal capacity, all of which greatly increase the complexity of the algorithm.
[0176] To simplify the problem of finding the global optimum, we propose an adaptive optimization algorithm based on PSO to maximize cache hit rate, which we call Adaptive Particle Swarm Optimization (PSO-MPhit). As a global optimization algorithm, PSO-MPhit simulates the foraging behavior of bird flocks and dynamically updates the inertia factor. Compared with the traditional PSO algorithm, it is faster and has stronger global search capabilities; it is particularly suitable for solving nonlinear and multimodal problems.
[0177] Inertia weight is an important evolutionary parameter that determines the degree to which a particle's previous velocity influences its current velocity. A smaller inertia weight results in stronger local search capability but weaker global search capability; conversely, a larger inertia weight results in weaker local search capability but stronger global search capability. Therefore, to optimize the algorithm's search capability and performance, and to balance the particle's local and global search capabilities, the inertia weight value is adjusted.
[0178] In this invention, the inertia weight is dynamically updated between the minimum and maximum inertia values during each search iteration of the particle swarm optimization algorithm. The dynamic update formula for the inertia weight is as follows:
[0179]
[0180]
[0181] in, It is the first iteration , Where N is the maximum number of iterations, and N is the population size. It is the optimal value of the current fitness function. This is the current location. The fitness function value, It is the minimum value of the current inertia weight. It is the maximum value of the current inertia weight.
[0182] Formula (26) calculates the average difference between the fitness function value of all particles in the current iter-th iteration and the historical best fitness function value, which is used to represent the current search status of the particle swarm. Formula (27) uses... Parameters to dynamically adjust inertia weights , exist[ , The search range is updated within a certain range and varies with the particle's fitness. Therefore, at the beginning of the iteration, the particle swarm search range is very large and needs to be dynamically increased. The value facilitates global search, but needs to be dynamically reduced when the particle swarm locks onto the target area. This value facilitates local search, thus enabling a faster and more accurate discovery of the optimal solution for the target.
[0183] This embodiment performs numerical simulation of the caching strategy proposed in this invention and verifies the performance of our algorithm by comparing the numerical results with other caching strategies under different parameter variations. First, to verify the system performance of our strategy under a single parameter variation, we need to pre-set the relevant parameters, which generally do not change unless otherwise specified: In collaborative filtering algorithms, the threshold for determining user similarity... Set to 0.6, this parameter is used to adjust the difference in user preference for files with different ratings. Let's set it to 5, as for The value is 0.3.
[0184] Figure 3A comparison was made of the changes in the hit probability of the five caching strategies as the number of users changed. It is clear that the cache hit rate increases with the number of users, since the number of files and terminal cache capacity are fixed. If more mobile terminals are available around the user to provide D2D services, this will undoubtedly improve the hit rate. In the most popular caching scheme (MPCStrategy), everyone caches the same files, and users cannot help each other. Therefore, regardless of changes in U, its hit probability remains essentially unchanged and is lower than the PSO-MPhit caching strategy of this invention. The equal probability random caching algorithm (EPRCStrategy) caches all files indiscriminately. Since most files in the file library have relatively low popularity, caching a large number of these files only wastes user cache space, hence its consistently low hit probability. Although the SBRC strategy improves upon the EPRC strategy by caching files with different popularity ranges with varying probabilities, it does not consider user preferences. Whether to cache a file is based on general public preference, so the probability of a user finding a file they like in their own cache space is lower than the strategy of this invention. In this invention, when the number of users is small, it tends to selectively cache files that are more favored by users to increase the probability of self-caching. At the same time, this invention also caches some highly popular files with equal probability, ensuring that all highly popular files have a chance to be cached, and the caching probability is much higher than that of the EPRC strategy. In other words, it incorporates the advantages of the SBRC strategy, so we can... Figure 3 The performance of the EPRC and SBRC strategies is lower than that of the PSO-MPhit strategy. The Greedy Algorithm Strategy for user-personalized preferences searches for the element with the highest hit rate gain through multiple iterations, considering the impact of each user's cached files on their own and surrounding users' hit rate gains. However, its computational complexity is extremely high, and the computation time is long. The PSO-MPhit strategy of this invention has a lower computation time, while achieving a hit rate close to that of the Greedy Algorithm Strategy for user-personalized preferences. This means that, with different caching methods, the algorithm complexity is lower, and this invention achieves near-optimal performance.
[0185] To fully illustrate the advantages and disadvantages of the user personalization preference caching strategy and the PSO-MPhit strategy of this invention in terms of computational complexity, this invention simulates the ratio of the runtime required for the user personalization preference strategy simulation to the runtime of the PSO-MPhit caching under different numbers of files and users. Figure 4It can be seen that this ratio is always much greater than 1, indicating that the runtime required for caching simulation of user personalized preferences is indeed much greater than the runtime of PSO-MPhit caching. As the number of users increases, the slope of the graph showing the ratio of the runtime of these two strategies to the number of files also increases. This is because the greedy algorithm needs to find the optimal setting of the cache space for each user. An increase in the number of users means an increase in the number of iterations, leading to greater computational complexity. In contrast, the number of iterations of the PSO-MPhit strategy is not affected by the number of users. Therefore, the PSO-MPhit strategy is more suitable for environments with a high number of users.
[0186] Figure 5 This paper compares the changes in cache hit probability as the probability of successful D2D links changes. The strategy of this invention allocates a portion of the cache space to cache files based on users' personalized preferences, ensuring a good cache hit probability even when the probability of successful D2D links is low. In contrast, the SBRC and EPRC strategies indiscriminately cache all or part of the file library. Users obtain their preferred files primarily through the D2D link, resulting in a low cache hit probability when the probability of successful D2D links is low. The user-personalized preference caching strategy determines the optimal cache location based on the highest cache hit rate gain, meaning it can adjust user file caching according to probability changes, thus maintaining a consistently high cache hit rate. Additionally, this invention also caches files from the file library with equal probability in another portion of the cache space. The performance gain from caching this portion of files increases with increasing probability, but since it is necessarily smaller than the cache space of other caching strategies, its performance gain is smaller than that of the SBRC and EPRC strategies. In the MPC strategy, everyone caches the same files, making it unnecessary to request files from other users through the D2D link. Therefore, regardless of the changes, its hit probability remains essentially unchanged.
[0187] Figure 6 This demonstrates that a larger cache capacity means more files can be cached by users, leading to an increased cache hit rate. The strategy of this invention exhibits the largest increase compared to other strategies. A larger cache capacity requires more rational use. Different users have different activity levels; obviously, more active users need more cache space, while excessive cache space is wasteful for inactive users. In the strategy of this invention, cache space partitioning is a crucial step. By partitioning the cache space, both personalized user preferences and general public preferences are considered. This ensures that users can cache as many files as they like, while also allowing each user to cache some popular files. This provides D2D services to users around them who are making requests, leveraging the advantages of D2D collaborative communication and fully utilizing user cache space. Simulation results also demonstrate that the strategy of this invention outperforms other strategies.
[0188] Figure 7 The results show that as the number of files increases, the cache hit probability decreases. The PSO-MPhi strategy of this invention outperforms other strategies except for the user-personalized preference caching strategy, and its performance remains comparable to the user-personalized preference caching strategy. When the number of files increases, user choices become more diverse, and the objective popularity of files becomes more dispersed. At this point, the MPC strategy, limited by cache capacity, cannot accurately cache files with a higher probability of user requests, resulting in the lowest cache hit rate. This invention, because it has a cache space dedicated to caching user-preferred files, and to address the problem of insufficient cache space as the number of files increases, also incorporates the advantages of the SBRC strategy, caching a portion of the file library with equal probability, caching as many diverse files as possible to facilitate D2D communication. By optimizing the segmentation of the file library, users can cache files they like but are not commonly used, and because files in the high-popularity range are cached with equal probability, even if users cannot find the required high-popularity files in their own cache space, they can more easily find them through D2D. The user-personalized preference caching strategy achieves a similar effect through traversal, and... Figure 4 We know that the user-personalized preference caching strategy is more computationally complex and takes longer to compute than the PSO-MPhi strategy. In practical applications, the number of files is enormous, making the PSO-MPhi strategy of this invention clearly more suitable for real-world applications.
[0189] The probability of contact between users depends on the contact time parameter and the contact frequency parameter. A higher contact frequency parameter increases the probability of contact between two users, meaning a greater probability that both users are simultaneously within D2D communication range. This undoubtedly increases the probability that a user will receive assistance from nearby D2D users. Figure 8 As we know, as the contact frequency parameter increases, the system's cache hit probability also gradually increases, with an effect similar to the success probability of a D2D connection. The figure shows that the performance of the caching strategy of this invention is greater than other caching strategies. When the contact frequency parameter increases to a certain value, users can basically perform D2D communication. At this point, even if a user cannot find their preferred file within their own cache, they have a fairly high probability of finding the required file through D2D communication. Therefore, caching more files with equal probability obviously brings greater system gains. Thus, the cache hit rates of the SBRC and EPRC strategies are very close. The user-personalized preference caching strategy and the PSO-MPhi strategy consider user personalization preferences, ensuring that users can find their favorite files in their own cache space, while files of moderate interest can also be found nearby. Therefore, their cache hit rates are still higher than other strategies.
[0190] Simulation results show that the segmented caching method and optimization algorithm of the present invention have good performance. After comprehensively considering the objective popularity and users' personal preferences, the system's average cache hit rate has a certain advantage compared with other popular caching strategies and optimization algorithms.
[0191] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A segmented caching method combining user preferences and objective popularity, characterized in that, The segmented caching method includes the following steps: Obtain user ratings and topic preference ratings for each file in the base station file library; The probability of a user requesting each file in the base station's file library is obtained from the rating and topic preference ratings, and the request probabilities are ranked first. These files form the first file library; The popularity of each file is determined based on the request probability of all users at the base station for each file in the base station's file library. All files are then ranked by popularity, and the top-ranked files are... The files form a second file library, and files that are duplicates of the first file library are removed. The user's cache space is divided into a first space that caches only files from the first file library and a second space that caches only files from the second file library. The method for obtaining the probability of a user requesting each file in the base station file library from the scoring rating and topic preference rating includes: Assumption , Representing users respectively The lowest and highest ratings for files in the file library, parameters Indicates user The difference in scoring and rating different documents is defined as: ; Therefore, users For the file Request probability The expression is as follows: ; Used to adjust for differences in file preferences among users who give the same file different scores; Assumption , Representing users respectively The parameters are the lowest and highest topic preference ratings for files in the file library. Indicates user The difference in the rating of different document topic preferences is defined as: ; Therefore, users For the file Request probability The expression is as follows: ; We get users For the file Total request probability as follows: ; ; ; in, Indicates user For the file Request probability The weights; Based on all users' views of the file The request probability is used to obtain the file. The expression for popularity is: 。 2. The segmented caching method combining user preferences and objective popularity as described in claim 1, characterized in that, The process of obtaining user ratings for each file in the base station's file library includes using a collaborative filtering algorithm to obtain similarity between users and then predicting ratings for files that users have not yet rated.
3. The segmented caching method combining user preferences and objective popularity as described in claim 2, characterized in that, The specific method for using collaborative filtering algorithms to obtain the similarity between users and then predicting the rating of documents that users have not yet rated is as follows: Calculate target users Other users Similarity between Find the target user User sets with similar behaviors in document rating and classification First, based on the Pearson similarity calculation formula, it is initially determined that... ,Right now: ; in, On behalf of users For the file The rating and rating On behalf of users For the file The rating and rating On behalf of users Average score for all files On behalf of users Average score for all files The number of files in the library. Represents the total number of users; Based on formula (1), another scoring parameter is proposed. : ; ; in, Indicates user and users The number of publicly rated documents and Representing users respectively Minimum and maximum values for the number of publicly rated files shared with other users; Set a threshold To filter out similarity Users whose threshold is higher than this threshold, when ≤ , indicating user and users They have similar scoring and rating behaviors, namely And expressed as: ; Based on the obtained user set For target users Documents that have not been rated will be assigned a rating or score, and the predicted rating will be represented as follows: ; in, {0,1,2,3,4...Z}, The higher the value, the better the user experience. For the file The more you like it, =0 indicates user No file Rating and rating.
4. The segmented caching method combining user preferences and objective popularity as described in claim 3, characterized in that, The method for obtaining the user's topic preference rating for each file in the base station file library includes obtaining the user's topic preference rating for each file in the base station file library based on the file topic classification and the user's historical request records.
5. The segmented caching method combining user preferences and objective popularity according to claim 4, characterized in that, Based on the file topic classification and the user's historical request records, the user's topic preference rating for each file in the base station file library includes: Assuming the entire file repository contains a total of T topics, use sets... The theme is indicated. and documents The attribute function is Its value depends on the file. Is it included in the topic? Li, that is: ; Assumption Indicates user For the subject of the document The preference function, which is expressed by the mutual information formula, assumes... By user Historical requests to record decisions, The definition is as follows: ; in, It is mutual information. For includes A collection of all files, For users Historical request records, For users The historical request record contains The probability of topic files, To exist throughout the entire file library The probability of a topic file; user For the file Interest level parameters Defined as: ; Received ∈[−1,1], The smaller the value, the better for the user. For the file The less interested, according to Get users For the file Theme Preference Rating The definition is as follows: 。 6. The segmented caching method combining user preferences and objective popularity according to claim 5, characterized in that, The segmented caching method further includes: Construct a problem to optimize the average cache hit rate; The average cache hit rate optimization problem is solved by using an adaptive particle swarm optimization algorithm, resulting in an optimized average cache hit rate.
7. A segmented caching method combining user preferences and objective popularity as described in claim 6, characterized in that, The methods for constructing the average cache hit rate optimization problem include: Assuming cache hit rate Defined as user Request file At that time, the probability of finding and obtaining the target file within its own cache space or within the range of successfully linked D2D links in the surrounding area, assuming that when a user Request file At that time, the user had already cached the file. Then, the requested file is retrieved from its own cache space, and the probability of the event occurring is... Suppose a user Request file , Not in its own cache When that happens, a request is sent to the user within the D2D communication range; if a cache exists... Users then obtain information through the D2D link. The probability of an event occurring And users Request file It is also a probabilistic event, let's say... That is, the probability of its occurrence. The cache hit rate is defined as follows: ; in, Indicates user The activity level, i.e., the user The probability of issuing a file request indicates the user's... The activity level is defined as: ; If user For the file If there is active scoring and rating behavior, then otherwise ; set up On behalf of users Cached files The probability of the user Found in its own cache space probability That is equivalent to user Cached files probability ,Right now: ; There must be at least one user within the area where the D2D link can be successfully established. = Cached users Requested document The probability is: ; in, represent Cached files The probability, This represents the user within a time segment. and users The probability of successfully performing a D2D link; So users Acquired via D2D link The probability is used Represented as: ; in, The probability of a successful D2D link; The system's average cache hit rate is: ; Assuming user The sizes of the first space and the second space are respectively , Represents the first document library, Representing the second file library, the number of duplicate files removed is ; When the file At that time, the user Cached files probability Represented as: ; When the file At that time, the user Cached files probability Represented as: ; Therefore, users cached files probability for: ; Therefore, to maximize the average cache hit rate, it is necessary to find the optimal segmentation point of both the base station file library and the user cache space. Hence, the average cache hit rate optimization problem is defined as follows: ; in, At the same time, both decision variables are discrete variables due to the influence of the file library segmentation point and the user cache space segmentation point. Under the influence of these factors, the problem of formula (25) becomes a non-convex nonlinear constraint optimization problem.
8. The segmented caching method combining user preferences and objective popularity according to claim 7, characterized in that, The method of solving the average cache hit rate optimization problem using the adaptive particle swarm optimization algorithm includes: In each search iteration of the particle swarm optimization algorithm, the inertia weight is dynamically updated between the minimum and maximum inertia values. The formula for dynamically updating the inertia weight is as follows: ; ; in, It is the first iteration , Where N is the maximum number of iterations, and N is the population size. It is the optimal value of the current fitness function. This is the current location. The fitness function value, It is the minimum value of the current inertia weight. It is the maximum value of the current inertia weight; Formula (26) calculates the average difference between the fitness function value of all particles in the current iter-th iteration and the historical best fitness function value, which is used to represent the current search status of the particle swarm. Formula (27) uses... Parameters to dynamically adjust inertia weights , exist[ , The search range is updated within a certain range and varies with the particle's fitness. Therefore, at the beginning of the iteration, the particle swarm search range is very large and needs to be dynamically increased. The value facilitates global search, but needs to be dynamically reduced when the particle swarm locks onto the target area. This value facilitates local search, thus enabling a faster and more accurate discovery of the optimal solution for the target.
Citation Information
Patent Citations
Base station and content caching method based on local popularity
CN109639844A
Heterogeneous network cache decision-making method based on user preference prediction
CN111860595A