Reliability-driven duplicated data deletion storage optimization method

By building a user-content two-part graph in a mobile edge network, combining request distribution and life cycle prediction replica counts, and using file association graphs and multi-objective optimization strategies for replica deployment and elimination, the problem of predicting and deploying replica counts in popular content storage is solved, and user experience and storage resource utilization is improved.

CN119960670APending Publication Date: 2025-05-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411926302.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In mobile edge networks, it is difficult for the prior art to effectively predict the number of replicas of popular content and conduct reasonable replica deployment and elimination when edge storage resources are limited, resulting in a decline in user experience and low storage resource utilization.

Method used

By building a user-content two-part graph, popular content is selected; combining user request distribution and future life cycle, dynamic LSTM model is used to predict the life cycle of content, and the number of copies to be stored in the content is calculated; building a file access correlation graph, select the copy to be deleted, and determine the replica deployment location through a multi-objective optimization strategy.

Benefits of technology

Improves the accuracy and efficiency of popular content storage, reduces computing overhead, guarantees service reliability, alleviates performance degradation caused by frequent changes in the number of replicas, and improves the hit rate of user requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960670A_ABST
    Figure CN119960670A_ABST
Patent Text Reader

Abstract

The invention relates to a reliability-driven duplicated data deletion storage optimization method, which belongs to the field of mobile edge network storage, and specifically comprises the following steps that: a cloud service provider or a user issues contents to an edge server, and the edge server counts user access records and stores the user access records; and determining the life cycle and the copy number of the popular content in combination with access statistics and aging characteristics. And based on the determined number of the copies, establishing a content association graph according to the access relationship, screening out repeated copies to be deleted, modeling a copy deployment problem into an optimization problem, and selecting an optimal decision for copy deployment. According to the method, the mobility of the user and the content change are considered, the replica deletion method and the replica deployment model are designed for improving the service experience of the user in the mobile edge network, the influence on other content access after the replica is deleted is reduced, the optimal deployment decision is found, and the user experience is improved. And the utilization rate of edge storage is improved while the requirements of different users are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of popular content storage in a mobile edge network (content is presented in the form of files in this article, and files and content refer to the same thing and have the same meaning), and in particular to a method for predicting popular content copies, eliminating cached content, and deploying copies. Background Art

[0002] Mobile Edge Computing provides users with close-range storage and computing resources. Service providers can store content such as videos and streaming media on the edge side to provide users with low-latency, high-quality services. However, the capacity of edge servers is limited, and the needs of mobile users are also diverse. Edge servers can fully utilize their storage space by storing popular content to meet the requests of most users. As users move or popular content changes, edge servers need to dynamically determine the content to be stored. Therefore, designing a popular content storage method in a mobile edge network that can predict the popular content to be stored and formulate corresponding storage and deletion strategies is the key to improving mobile user experience and improving edge server utilization.

[0003] In order to improve the utilization rate of edge storage resources and meet the access needs of most users, it is necessary to predict the popularity and access volume of content and store popular content at the edge. At present, there are two main methods for predicting content popularity. One is to train machine learning models through data sets to predict popularity and access volume, and the other is to establish relevant mathematical models for special access patterns through statistical and probabilistic methods to solve the relevant popularity and access volume. The former requires deep learning model training, which has the problem of more model parameters and longer training time, which is not suitable for delay-sensitive scenarios on the edge side. The mathematical model derived by the latter is often targeted at specific data and scenarios, without high universality, and different data may have unstable prediction accuracy. In addition, it is also necessary to consider the impact of changes in popular content and user mobility on the number of copies of popular content. When deploying copies, on the one hand, the above factors are also limited by edge storage resources, that is, duplicate deletion of copies eliminated in the previous round; on the other hand, it is also necessary to consider the differences in content requests from different users. In summary, how to combine user mobility and needs to determine the current popular content and its number of copies, as well as how to deduplicate and deploy copies when storage resources are limited has become an urgent problem to be solved.

[0004] Therefore, the present invention proposes a reliability-driven deduplication storage optimization strategy.

[0005] After searching, the application publication number CN113965937B belongs to the technical field of predicting content popularity. The method includes providing a content popularity prediction method based on clustering federated learning in a fog wireless access network, and the steps are as follows: constructing initial features of local users and content based on local user information and content information collected by fog access points; establishing a prediction model of the probability of local users requesting content for each fog access point based on the initial features and historical request records; using clustering federated learning to perform distributed training on the prediction models of each fog access point and realize the specialization of model parameters; according to the content information, taking the content request probability of mobile users as the prediction target, establishing a preference model of mobile users; integrating the prediction results of local popularity and mobile popularity to obtain the final prediction result of content popularity; enabling fog access points to accurately predict and dynamically update content popularity, and through model specialization, adaptively distinguish regional differences in content popularity, while reducing communication costs.

[0006] The difference from the above method is that the present invention considers content storage in mobile edge scenarios. For mobile edge scenarios, clustering federated learning methods and inference user preference models may require long iterations and high computing resource requirements, which are not universal for mobile edge scenarios that are sensitive to latency and have limited computing resources. In addition, in view of the changes in popular content over time and the needs of mobile users, it is necessary to consider the redeployment of edge storage content, including the deduplication and copy deployment of edge content. The above patent provides a method for predicting the popularity of content, which does not take into account the particularity of mobile edge scenarios and the elimination strategy of popular content. This patent proposes that through the proposed popular content copy prediction method and copy deployment method, popular content can be effectively stored and the access needs of mobile users can be met under the constraints of limited edge storage and user access latency. Summary of the invention

[0007] The present invention aims to solve the above problems of the prior art. A reliability-driven deduplication storage optimization method is proposed. The technical solution of the present invention is as follows:

[0008] A reliability-driven deduplication storage optimization method comprises the following steps:

[0009] S1. The edge server builds a user-content bipartite graph and filters out popular content based on the degree of content nodes in the bipartite graph;

[0010] S2: The edge server calculates the user request distribution of popular content and sends it to the edge controller, which calculates the final request distribution result;

[0011] S3, the edge controller trains a dynamic LSTM model to predict the future life cycle of popular content;

[0012] S4. The edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content.

[0013] S5. The edge server builds a file access association graph, and the edge controller selects the replica to be deleted;

[0014] S6. Establish optimization goals and calculate where the replicas need to be deployed.

[0015] Furthermore, in step S1, the edge server constructs a user-content bipartite graph, and filters out popular content based on the degree of content nodes in the bipartite graph, specifically including:

[0016] S11. The edge controller obtains access records from each edge server;

[0017] S12. Each edge server builds a user-content bipartite graph based on the access log;

[0018] S13. The edge controller selects K contents with the most access users as popular contents.

[0019] Furthermore, in step S2, the edge server calculates the user request distribution of popular content and sends it to the edge controller, and the controller calculates the final request distribution result, which specifically includes:

[0020] S21. Each edge server calculates the user request distribution of popular content on the server based on the user-content bipartite graph and transmits it to the edge controller;

[0021] S22. The edge controller summarizes the user request distribution of each edge server regarding popular content and calculates the final result of the user request distribution.

[0022] Furthermore, in step S3, the edge controller trains a dynamic LSTM model to predict the future life cycle of popular content, specifically including:

[0023] S31. Train a multi-step prediction LSTM model based on the data set and obtain a high-accuracy model by iterative parameter combination;

[0024] S32. Taking the number of visits to popular content in a certain time window as input, obtaining a prediction result;

[0025] S33. Calculate the future life cycle of the content based on the prediction results and the historical visits of the content.

[0026] Furthermore, in step S4, the edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content, specifically including:

[0027] S41. Determine the weight factor based on the content publishing time and access data;

[0028] S42. Select a weight factor and calculate the number of copies of the content based on the user request distribution and the future life cycle.

[0029] Furthermore, in step S5, the edge server constructs a file access association graph, and the edge controller selects the copy to be deleted, which specifically includes:

[0030] S51. The edge controller divides the number of content copies into two categories: copies that need to be deleted and copies that need to be added;

[0031] S52. The edge server establishes a file access association graph based on the user-content bipartite graph and transmits the corresponding results to the controller;

[0032] S53. The edge controller filters the copies to be deleted according to the file association graph and notifies the relevant edge servers.

[0033] Furthermore, in step S6, the optimization goal is established and the location where the replica needs to be deployed is calculated, which specifically includes:

[0034] S61. The edge server calculates the relevant indicators for each deployment decision;

[0035] S62. Establish optimization goals, calculate target values ​​for all deployment decisions, and select the optimal deployment decision.

[0036] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the reliability-driven deduplication storage optimization method as described in any one of the items is implemented.

[0037] A non-transitory computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements any reliability-driven deduplication storage optimization method as described in any one of the items.

[0038] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the reliability-driven deduplication storage optimization method as described in any one of the items is implemented.

[0039] The advantages and beneficial effects of the present invention are as follows:

[0040] 1. Aiming at the unreliable edge network and user mobility problems in popular content storage, and analyzing the existing popularity evaluation model and content copy calculation method, the present invention combines the user request distribution and the future life cycle to jointly determine the popular content and the number of its copies in step S4. Compared with other machine learning methods, the proposed method reduces the computational overhead; it can alleviate the performance degradation caused by the frequent changes in the number of copies while ensuring service reliability.

[0041] 2. In response to the contradiction between the content replica deployment problem and the limited edge storage resources, first, in step S5 of the present invention, a file access association graph is constructed to delete the replicas that need to be eliminated, which solves the problems caused by limited edge storage resources and changes in popular content, and can reduce the impact of users accessing other related content after deleting the replicas. Compared with the memory-first elimination method, the proposed method improves the hit rate of user requests. Secondly, in step S6 of the present invention, a multi-objective optimized replica deployment strategy is proposed. The problem is converted into an optimization indicator to select the optimal decision. This decision can meet the requests and access latency requirements of most users for popular content, and can also meet the preference needs of some users. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flow chart of the overall preferred embodiment provided by the present invention;

[0043] Figure 2 This is the system architecture diagram;

[0044] Figure 3 To construct an example diagram of a file association graph;

[0045] Figure 4 The following is an example diagram of edge server network topology. DETAILED DESCRIPTION

[0046] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.

[0047] The technical solution of the present invention to solve the above technical problems is:

[0048] 1. In order to solve the problem of inaccuracy and inefficiency in screening popular content and determining the number of its copies, this paper proposes a calculation scheme based on user request distribution and future life cycle. It selects Top K content according to the access relationship, calculates the content popularity through request distribution, and combines the LSTM model to predict the number of visits in multiple time periods in the future to jointly determine the number of copies of the content.

[0049] 2. In view of the contradiction between content replica deployment and limited edge storage resources, this paper first proposes a deduplication strategy based on file association graph. Secondly, this paper transforms the replica deployment problem into an optimization problem and solves the optimal solution for replica deployment.

[0050] like Figure 1 As shown, the present invention provides a reliability-driven deduplication storage optimization strategy, characterized in that it includes the following steps:

[0051] S1. The edge server builds a user-content bipartite graph and filters out popular content based on the degree of content nodes in the bipartite graph;

[0052] S2. The edge server calculates the user request distribution of popular content and sends it to the edge controller, which calculates the final request distribution result;

[0053] S3. The edge controller trains a dynamic LSTM model to predict the future life cycle of popular content;

[0054] S4. The edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content.

[0055] S5. The edge server builds a file access association graph, and the edge controller selects the replica to be deleted;

[0056] S6. Establish optimization goals and calculate where the replicas need to be deployed.

[0057] In this embodiment, in step S1, the edge server constructs a user-content bipartite graph, and filters out popular content based on the degree of content nodes in the bipartite graph, which specifically includes the following steps:

[0058] In this embodiment, in step S1, the edge server constructs a user-content bipartite graph, and filters out popular content based on the degree of content nodes in the bipartite graph, which specifically includes the following steps:

[0059] S11. The edge controller obtains access records from each edge server

[0060] The edge controller requests access records of the past 24 hours from all edge servers and processes the data.

[0061] S12. Each edge server builds a user-content bipartite graph based on access logs

[0062] This paper defines popular content as the content with the most users. In order to better formalize the description, this paper uses a graph to describe the access relationship between users and content. Given an undirected graph G<

[0063] V,E>, V=U∪F. For the node set V, it consists of the user set U and the content set F; for the edge e∈E, w is the weight of e, which means that user u has accessed file f a total of w times in the past 24 hours. In this way, the degree of each content node f reflects the number of users who have accessed the content.

[0064] S13. The edge controller selects K contents with the most access users as popular contents.

[0065] The edge controller can filter out the K content nodes with the largest degree from the access records as popular content, and can also count the number of visits to each popular content in the past 24 hours. Filtering out the K nodes with the largest degree can be achieved through the maximum heap method. After determining the popular content, the edge controller broadcasts it to the edge server.

[0066] In this embodiment, in step S2, the edge server calculates the user request distribution of popular content and then sends it to the edge controller, and the controller calculates the final request distribution result, which specifically includes the following steps:

[0067] S21. Each edge server calculates the user request distribution of popular content on the server based on the user-content bipartite graph and transmits it to the edge controller

[0068] Since the number of requests for the same content may not be evenly distributed among different users, this paper introduces user request distribution, which represents the proportion of different user requests to the total number of visits to the content. The request distribution of each user for file f can be calculated by the following formula:

[0069]

[0070] For file f, its average request distribution can be expressed by the following formula:

[0071]

[0072] Given that cnt1 has an initial value of 0, for each if Then cnt1 needs to be increased by 1. Finally, the request distribution of file f can be calculated by the following formula:

[0073]

[0074] For the same content, when the number of visits of more than half of the users is less than or equal to the average number of visits, this article believes that the request distribution for the file has reached the average standard.

[0075] S22. The edge controller summarizes the user request distribution of each edge server for popular content and calculates the final result of the user request distribution

[0076] Given cnt2 initial value 0, for each, if Then cnt2 needs to be increased by 1. Finally, the request distribution of file f can be calculated by the following formula:

[0077]

[0078] The above formula shows that when the user distribution is relatively even, the controller will select the most even one. When the user distribution is quite different, the controller will select the one from fp i,f The smallest value greater than or equal to 0.5 is selected as the final result.

[0079] The following is an actual example to illustrate the process of the controller calculating the request distribution. Assume that the access records of file f on three edge servers are as follows, where count represents the number of visits by each user:

[0080]

[0081]

[0082] The final result received by the edge controller is shown in the following table, where total represents the total number of visits on each edge server:

[0083]

[0084] After receiving the results calculated by each edge server, the controller needs to summarize and calculate the final request distribution. In the above example, the final calculated result is 0.20571429.

[0085] In this embodiment, in step S3, the edge controller trains a dynamic LSTM model to predict the future life cycle of popular content, which specifically includes the following steps:

[0086] S31. Train a multi-step prediction LSTM model based on the data set and obtain a high-accuracy model by iterative parameter combination

[0087] This article uses the popular music dataset provided by Alibaba Cloud to train the LSTM model. This dataset records the user's access records to content, including user id, content id, timestamp and other fields. In order to facilitate the application of different scenarios and datasets, the input set req of the model is the number of visits to the content in the past X moments, and the output result res of the model represents the number of visits to the content in the future y moments. Among them, X and y are both configurable. In this dataset, X = 12 and y = 5. Time is in hours.

[0088] This article detects the input data in the dataset and discards blank and abnormal historical data. Then, the parameters of the network structure are dynamically determined, including the number of desnse layers, the number of LSTM network layers, and the number of neurons as the objects to be optimized. There are k candidate values for the number of desnse layers, n candidate values for the number of LSTM layers, and u candidate values for the number of neurons. Finally, a model with relatively high accuracy is selected. On this dataset, k = 5, n = 2, and u = 32.

[0089] S32. Use the access volume of popular content within a certain time window as the input to obtain the prediction result

[0090] When the size of the input content access volume set req meets X, the prediction result set res can be directly given by the model.

[0091] S33. Calculate the future life cycle of the content based on the prediction result and the historical access volume of the content

[0092] First, calculate the average access volume of the content in the past 12 hours Secondly, calculate the ttl of the content through the following formula, indicating that the content will be deleted after ttl time rounds:

[0093]

[0094] Among them, the introduced step function h(x) is defined as follows:

[0095]

[0096] Specifically, when the statistically obtained content access volume does not meet the input requirement X of the model, the life cycle of the content needs to be calculated separately. When |req| ≤ y, ttl = 3; when y < |req| < X:

[0097]

[0098] where req now represents the access volume at the current moment f, and req now-1 represents the access volume at the previous moment f.

[0099] After obtaining the ttl of the file, the future life cycle of the file at the current moment can be calculated:

[0100]

[0101] For the convenience of understanding, the following will give an example of calculating the life cycle of a file. Assume that the access volume of file f in the past 12 hours is req = {121, 20, 236, 70, 201, 102, 33, 100, 0, 80, 100, 200}, and the model output res = {142, 122, 121, 98, 50}. Calculate is 105.25, then ttl=3,l f =0.6.

[0102] Count the number of file visits per hour in hours. Assuming that the file visit data is less than 12 hours, if the file visit collection size is less than or equal to the output result set size, the default ttl is 3; when it can be counted that the file visit collection size is greater than the output result set size but less than 12 hours, it is necessary to compare the visit volume of the previous moment. If there is no decrease, keep the ttl unchanged. If the visit volume decreases. Then the ttl is the ttl of the previous moment minus 1. In particular, when the ttl is reduced to 1, it will no longer decrease. For example: file f at time t0 req = {100}, at this time, ttl = 3, l f =0.6.

[0103] Until time t5, req = {123,23,331,45,234,124,221}. At this time, since the file access volume has not decreased, ttl remains unchanged; at time t6, req = {123,23,331,45,234,124,221,99}. At this time, ttl = 2,l f =0.4. Assuming that the number of visits continues to decrease at subsequent times, then from time t7 to t 10 time, ttl=1,l f =0.2.

[0104] In this embodiment, in step S4, the edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content, which specifically includes the following steps:

[0105] S41. Determine the weight factor based on the content publishing time and access data;

[0106] According to the above steps S2 and S3, the number of copies of the user can be calculated by the following formula:

[0107]

[0108] In the formula, a represents the weight factor, and N represents the number of edge servers. By default, a = 0.5. In particular, when predicting the future life cycle, the number of visits to the content may not meet the size of X mentioned in step S3. Since there is not enough data, the weights of different factors need to be adjusted. In this case, a = 0.7 is set.

[0109] S42. Select a weight factor and calculate the number of copies of the content based on the user request distribution and the future life cycle.

[0110] Calculate the copies of each popular content through the formula fThe output results of ttl are the number of copies of each content and its ttl. The reason for this design is that when the statistical access volume is insufficient, a file can become a popular file only if there are a large number of users accessing it in a short period of time. Therefore, it is necessary to increase the weight factor of the user request distribution. Secondly, due to insufficient data volume, the calculated life cycle may have a large error, so it is necessary to reduce the weight factor of the future life cycle. Third, since the access to new content is uncertain, if the change of its life cycle leads to the elimination of content and subsequent redeployment, such frequent operations may affect system performance.

[0111] In this embodiment, in step S5, the edge server constructs a file access association graph, and the edge controller selects the copy to be deleted, which specifically includes the following steps:

[0112] S51. The edge controller divides the number of content copies into two categories: copies that need to be deleted and copies that need to be added.

[0113] First, the controller compares the results of the new round of calculations with the results of the previous round, and selects the replicas that need to be deleted. The remaining results are the replicas that need to be newly deployed. For example, the controller calculates the number of content replicas in the previous round as shown in the following table:

[0114] <![CDATA[f1]]> <![CDATA[f2]]> <![CDATA[f3]]> <![CDATA[f4]]> <![CDATA[f5]]> <![CDATA[copies f ]]> 8 6 5 4 7 ttl 5 3 2 2 4

[0115] The new round of results calculated are shown in the following table:

[0116] <![CDATA[f1]]> <![CDATA[f2]]> <![CDATA[f3]]> <![CDATA[f4]]> <![CDATA[f6]]> <![CDATA[copies f ]]> 4 7 4 8 5 ttl 3 3 2 2 4

[0117] The controller can classify the relevant results into: replicas that need to be deleted: {f1:4,f3:1,f5:7}; replicas that need to be newly deployed: {f2∶1,f4∶4,f6:5}. In terms of execution order, this article gives priority to replicas that need to be deleted. When deploying new replicas, they are deployed in strict accordance with the order of the filtered popular content.

[0118] S52. The edge server establishes a file access association graph based on the user-content bipartite graph and transmits the corresponding results to the controller

[0119] After each edge server receives the copy that needs to be deleted from the controller, it builds a file access association graph based on the user-content bipartite graph mentioned in S1. If there is duplicate content to be deleted in the file access association graph, the edge server needs to send the content id and its file association to the controller. The process of building the file association graph is given below.

[0120] According to the bipartite graph G<V,E> ,V=U∪F construct graph G R <X,D> It is a file association diagram.

[0121] (1) If u in G i At the same time with f p ,f q If f p ,f q Add X and associate an initial edge with a weight of 1;

[0122] (2) andu j ≠u i , if u in G j At the same time with f p ,f q Associate. p ,f q The weight of the edge is increased by 1.

[0123] Enumerate all the relationships between U and F in G, and build a file association graph based on the above two items. Figure 3 The construction from a bipartite graph to a file association graph is shown.

[0124] S53. The edge controller filters the copies to be deleted according to the file association graph and notifies the relevant edge servers.

[0125] After receiving the result sent by the edge server, the controller selects the file with the smallest association degree in turn until the requirement is met, and then sends the information of the corresponding copy to be deleted to the edge server. For example, for file f3, the controller receives the following data:

[0126] <![CDATA[s1]]> <![CDATA[s2]]> <![CDATA[s3]]> <![CDATA[s4]]> <![CDATA[s5]]> degree 8 3 7 6 5

[0127] Based on the above results, the controller will notify the S2 edge server to delete file f3. In particular, when there are multiple identical results, the controller will randomly select a server to delete. For file f1, multiple copies need to be deleted. The controller will adopt a greedy strategy and select until all copies are deleted. For file f5, all edge servers need to be notified to delete the copies of f5.

[0128] When updating ttl, if there is ttl in the result, the latest ttl is used as the main one. If not, it means that the file is not a popular file in this round. It needs to be swapped out of the memory and the ttl value is reduced by 1. As time goes by, the ttl is reduced to 0. When the edge server storage resources reach the threshold, the files stored on the disk are deduplicated. The idle bandwidth is used to upload the content to the cloud server to ensure that the data is not lost.

[0129] In this embodiment, in step S6, establishing the optimization goal and calculating the location where the replica needs to be deployed specifically includes the following steps:

[0130] S61. The edge server calculates the relevant indicators for each deployment decision;

[0131] Given the number of copies of a content, each edge server needs to calculate,the storage load of the edge server, the latency of accessing the content, the distribution of user requests for the content,and the diversity of the content stored by the edge server.

[0132]

[0133] M i,j represents the remaining memory space after edge server j stores file i. j,used Indicates the memory used by the edge server, size i Indicates the size of file i, mem j Indicates the memory size of the edge server.

[0134]

[0135] In this paper, the access delay is quantified by the number of hops. In the network topology, if a node can cover more edges, then the accessible range will be larger and the delay will be within the specified hop count range. i,j ) represents the set of files i that can be found on server j. |N| represents the number of nodes in the network topology. When the network topology is determined, the L of edge servers i,j In this article, the default delay hop limit is 2 hops.

[0136]

[0137] C i,j Indicates the diversity of files stored on the server, where files j represents the amount of content stored in the server memory, |F| represents the number of new popular content given by the controller, and for fp j,i represents the distribution of user requests for file i on server j.

[0138] Taking f6 as an example, the data on 10 edge servers are given as shown in the following table:

[0139]

[0140]

[0141] Among them, the network topology between edge servers is as follows Figure 4 shown.

[0142] S62. Establish optimization goals, calculate target values ​​for all deployment decisions, and select the optimal deployment decision.

[0143] For each content, each edge server needs to calculate the above indicators, and then calculate the desn value. The controller decides on which edge servers to deploy copies.

[0144] decision=M i,j +L i,j +C i,j -fp j,i #(13)

[0145] For the data given in S61, the decision value of each edge server is calculated as follows:

[0146] <![CDATA[s1]]> <![CDATA[s2]]> <![CDATA[s3]]> <![CDATA[s4]]> <![CDATA[s5]]> decision 1.218544 1.814255 0.959836 1.069984 1.437987 <![CDATA[s6]]> <![CDATA[s7]]> <![CDATA[s8]]> <![CDATA[s9]]> <![CDATA[s 10 ]]> decision 1.926714 1.85768 2.010124 1.564763 1.588891

[0147] Since f6 needs to select 5 replicas, the controller selects the 5 edge servers with the highest decision values ​​in turn, which are {s8,s6,s7,s2,s 10}.

[0148] The present invention aims at the problem that the screening of popular content and the determination of the number of its copies in the mobile edge network are not accurate and efficient enough, and proposes a calculation scheme based on user request distribution and future life cycle, selects Top K content according to the access relationship, calculates the content popularity through the request distribution, and combines the LSTM model to predict the access volume of multiple time periods in the future to jointly determine the number of copies of the content. Compared with other methods, this paper considers that the number of requesting users rather than the access volume can better reflect the characteristics of popular content and mobile users, and the LSTM model with dynamically determined parameters can adapt to more scenarios compared with other methods. Compared with other models, the computational overhead of this paper is smaller and more suitable for mobile edge networks. In addition, this paper proposes a content copy storage method. In terms of optimizing storage resources, selecting and eliminating content through the file association graph can effectively reduce the impact of deleting content on the access of other content, and at the same time, combining the deduplication technology to optimize the hard disk storage of the edge server. On the other side of the copy deployment, it is necessary to calculate the storage load of the edge server, the delay of accessing the content, the user request distribution of the content, and the content diversity stored by the edge server. Construct an optimization index to select the most suitable deployment location for the current copy.

[0149] The systems, devices, modules or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. Computer-readable media include permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carriers.

[0150] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0151] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A reliability-driven deduplication storage optimization method, characterized in that: The following steps are involved: S1. The edge server builds a user-content bipartite graph and filters out popular content based on the degree of content nodes in the bipartite graph; S2: The edge server calculates the user request distribution of popular content and sends it to the edge controller, which calculates the final request distribution result; S3, the edge controller trains a dynamic LSTM model to predict the future life cycle of popular content; S4. The edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content. S5. The edge server builds a file access association graph, and the edge controller selects the replica to be deleted; S6. Establish optimization goals and calculate where the replicas need to be deployed.

2. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S1, the edge server constructs a user-content bipartite graph and filters out popular content based on the degree of content nodes in the bipartite graph, specifically including: S11. The edge controller obtains access records from each edge server; S12. Each edge server builds a user-content bipartite graph based on the access log; S13. The edge controller selects K contents with the most access users as popular contents.

3. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S2, the edge server calculates the user request distribution of popular content and sends it to the edge controller, which calculates the final request distribution result, specifically including: S21. Each edge server calculates the user request distribution of popular content on the server based on the user-content bipartite graph and transmits it to the edge controller; S22. The edge controller summarizes the user request distribution of each edge server regarding popular content and calculates the final result of the user request distribution.

4. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S3, the edge controller trains a dynamic LSTM model to predict the future life cycle of popular content, specifically including: S31. Train a multi-step prediction LSTM model based on the data set and obtain a high-accuracy model by iterative parameter combination; S32. Taking the number of visits to popular content in a certain time window as input, obtaining a prediction result; S33. Calculate the future life cycle of the content based on the prediction results and the historical visits of the content.

5. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S4, the edge controller calculates the number of copies of the content that need to be stored based on the user request distribution and future life cycle of the content, specifically including: S41. Determine the weight factor based on the content publishing time and access data; S42. Select a weight factor and calculate the number of copies of the content based on the user request distribution and the future life cycle.

6. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S5, the edge server constructs a file access association graph, and the edge controller selects the replica to be deleted, which specifically includes: S51. The edge controller divides the number of content copies into two categories: copies that need to be deleted and copies that need to be added; S52. The edge server establishes a file access association graph based on the user-content bipartite graph and transmits the corresponding results to the controller; S53. The edge controller filters the copies to be deleted according to the file association graph and notifies the relevant edge servers.

7. The reliability-driven deduplication storage optimization method according to claim 1, characterized in that: In step S6, the optimization goal is established and the location where the replica needs to be deployed is calculated, which specifically includes: S61. The edge server calculates the relevant indicators for each deployment decision; S62. Establish optimization goals, calculate target values ​​for all deployment decisions, and select the optimal deployment decision.

8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the reliability-driven deduplication storage optimization method according to any one of claims 1 to 7 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the reliability-driven deduplication storage optimization method according to any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the reliability-driven deduplication storage optimization method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A content popularity prediction method based on clustering federated learning in fog wireless access network

    CN113965937B