Mobile application cache resource optimization method and device
By deleting cache data in ascending order through the policy value network and cache value, combined with server-side optimization, the efficiency problem of mobile application cache resource management in dynamic environments is solved, and efficient cache resource optimization and dynamic adaptation are achieved.
Patent Information
- Application Number
- CN202510950860.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-14
AI Technical Summary
Existing mobile application cache resource management methods are difficult to adapt to dynamically changing environments, resulting in a decrease in cache hit rate and resource waste. Rule-based methods cannot respond to changes in user access patterns in a timely manner, while Q-Learning-based methods have high computational overhead and are not suitable for efficient operation on mobile devices.
The strategy value network is used to process cache processing data, obtain the first value parameter, delete cache data in ascending order according to the cache value, and update the value parameter by monitoring the cache overhead and hit rate, and optimize cache resource management in combination with the server-side strategy network.
Efficiently optimize cache space on mobile devices, adapt to dynamic environmental changes, improve cache hit rate, reduce resource waste, and lower computing overhead.
Smart Images

Figure CN120780482A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cache optimization, in particular to a mobile application cache resource optimization method and device. BACKGROUND
[0002] With the rapid development of mobile devices and the increasing complexity of user needs, the amount of resources that mobile applications installed thereon need to obtain from the network during user use has greatly increased. Therefore, efficient cache resource management has become a key to improving the performance of mobile applications and the user experience.
[0003] Currently, the management methods of cache resources for mobile applications mainly include rule-based methods and machine learning driven methods.
[0004] Rule-based methods, such as LRU, LFU, ARC, and the like, rely on user historical access data and cannot adapt to dynamic changes in the environment.
[0005] Taking a scheme of managing cache resources based on Q-Learning as an example, machine learning driven methods can learn access patterns, but have large computational overhead and are difficult to efficiently run on mobile devices. SUMMARY
[0006] In order to realize cache resource optimization that can adapt to dynamic changes in the environment and efficiently run on mobile devices, the present application discloses the following technical solutions:
[0007] The first aspect of the present application provides a mobile application cache resource optimization method applied to a mobile device, and the method comprises:
[0008] obtaining a first value parameter, wherein the first value parameter is obtained by a server from cache processing data of a plurality of mobile devices according to a policy value network;
[0009] performing value calculation on key data indicators corresponding to cache data of the mobile device according to the first value parameter to obtain cache values of the cache data, wherein the key data indicators are part of a plurality of data indicators;
[0010] sequentially deleting at least one item of cache data in ascending order of cache values, so that the total data amount of cache data in the mobile device matches the cache space capacity.
[0011] Optionally, after sequentially deleting at least one item of cache data in ascending order of cache values, the method further comprises:
[0012] when it is detected that the deleted cache data causes cache miss problems, restoring the cache data that causes the cache miss problems.
[0013] Optionally, after the at least one cache data is sequentially deleted in the ascending order of the cache values, the method further comprises:
[0014] monitoring the cache overhead indicator and the cache hit rate indicator;
[0015] updating the first value parameter according to the data indicators corresponding to each cache data in a preset time period when the cache overhead indicator and the cache hit rate indicator meet a target update condition.
[0016] Optionally, the updating the first value parameter according to the data indicators corresponding to each cache data in a preset time period comprises at least one of the following:
[0017] if the average value of the storage cost of cache miss data is greater than a storage cost threshold, adjusting a weight parameter corresponding to the storage cost in the first value parameter downward;
[0018] if the average value of the access frequency of cache miss data is greater than an access frequency threshold, adjusting a weight parameter corresponding to the access frequency in the first value parameter upward;
[0019] if the average value of the data freshness of cache miss data is greater than a data freshness threshold, adjusting a weight parameter corresponding to the data freshness in the first value parameter upward;
[0020] if the average value of the predicted access probability of cache miss data is greater than a predicted probability threshold, adjusting a weight parameter corresponding to the predicted access probability in the first value parameter upward.
[0021] Optionally, the cache overhead indicator and the cache hit rate indicator meet a target update condition, comprising:
[0022] in a preset time period, a descending amplitude of the cache overhead indicator is greater than a first threshold, and a descending amplitude of the cache hit rate is greater than a second threshold;
[0023] or, in a preset time period, an ascending amplitude of the cache overhead indicator is greater than a third threshold, and an ascending amplitude of the cache hit rate is greater than a fourth threshold.
[0024] Optionally, the plurality of data indicators comprise:
[0025] access frequency, data freshness, storage cost, predicted access probability, network delay, cache occupancy rate, recent access pattern, historical cache hit rate and cache data update times;
[0026] the key data indicators comprise:
[0027] access frequency, data freshness, storage cost and predicted access probability.
[0028] The second aspect of the application provides a mobile application cache resource optimization method, applied to a server, and the method comprises the following steps:
[0029] obtaining cache processing data of a plurality of mobile devices, wherein the cache processing data at least comprises a plurality of data indicators of cache data of the mobile devices and a cache miss record table of the mobile devices, and the cache miss record table is used for recording cache miss data of the mobile devices;
[0030] processing the cache processing data according to a policy network to obtain a value parameter corresponding to the cache processing data;
[0031] processing the cache processing data and expected cache processing data according to a value network to obtain a current data reward and an expected data reward, wherein the expected cache processing data is cache processing data obtained after the mobile devices perform cache resource optimization based on the value parameter, and the value network and the policy network constitute a policy value network;
[0032] updating network parameters of the policy network and network parameters of the value network according to the current data reward, the expected data reward and an actual reward value to complete training of the policy network and the value network, and the actual reward value is determined according to a cache occupancy rate and a cache hit rate in the cache processing data;
[0033] processing cache processing data of a mobile device according to the trained policy network and sending a first value parameter obtained by processing to the mobile device, so that the mobile device performs cache resource optimization based on the first value parameter.
[0034] Optionally, the processing of the cache processing data according to the policy network to obtain the value parameter corresponding to the cache processing data comprises the following steps:
[0035] composing a plurality of data indicators contained in the cache processing data into an input state, processing the input state according to the policy network to obtain the value parameter corresponding to the cache processing data.
[0036] The third aspect of the application provides a mobile application cache resource optimization device, applied to a mobile device, and the device comprises the following steps:
[0037] an obtaining unit, configured to obtain a first value parameter, wherein the first value parameter is obtained by a server processing cache processing data of a plurality of mobile devices according to a policy value network;
[0038] a calculation unit, configured to perform value calculation on a key data indicator corresponding to cache data of the mobile device according to the first value parameter to obtain a cache value of the cache data, wherein the key data indicator is a part of a plurality of data indicators;
[0039] The deleting unit deletes at least one item of cache data in turn in ascending order of cache value, so that the total data amount of cache data in the mobile device matches the cache space capacity.
[0040] The fourth aspect of the present application provides a mobile application cache resource optimization device applied to a server, and the device comprises:
[0041] The obtaining unit is configured to obtain cache processing data of a plurality of mobile devices, wherein the cache processing data at least comprises a plurality of data indicators of cache data of the mobile devices and a cache miss record table of the mobile devices, and the cache miss record table is used to record cache miss data of the mobile devices.
[0042] The parameter unit is configured to process the cache processing data according to a policy network to obtain a value parameter corresponding to the cache processing data.
[0043] The reward unit is configured to process the cache processing data and expected cache processing data according to a value network to obtain a current data reward and an expected data reward, wherein the expected cache processing data is cache processing data obtained after the mobile device optimizes cache resources based on the value parameter, and the value network and the policy network constitute a policy value network.
[0044] The updating unit is configured to update network parameters of the policy network and network parameters of the value network according to the current data reward, the expected data reward and an actual reward value to complete training of the policy network and the value network, wherein the actual reward value is determined according to a cache occupancy rate and a cache hit rate in the cache processing data.
[0045] The parameter unit is configured to process cache processing data of the mobile device according to the trained policy network and send a first value parameter obtained by processing to the mobile device, so that the mobile device optimizes cache resources based on the first value parameter.
[0046] The present application has the following beneficial effects:
[0047] On the one hand, the first value parameter is used to efficiently determine cache value on the mobile device, so that unnecessary cache data is accurately deleted based on the cache value, and cache space is optimized, and on the other hand, the server can dynamically determine the first value parameter according to the cache processing data, so as to adapt to a dynamically changing environment. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only constitute a part of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.
[0049] Figure 1 is a flow chart of a mobile application cache resource optimization method provided by an embodiment of the present application;
[0050] Figure 2 is a flow chart of another mobile application cache resource optimization method provided by an embodiment of the present application;
[0051] Figure 3 is a structural schematic diagram of a mobile application cache resource optimization device provided by an embodiment of the present application;
[0052] Figure 4 is a structural schematic diagram of another mobile application cache resource optimization device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only constitute a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0054] In order to facilitate the understanding of the technical solutions of the present application, first, some terms that may be involved are described.
[0055] Mobile device: a user's mobile phone, tablet computer and the like.
[0056] Mobile application: various application programs (i.e. commonly known as App) installed on a user's mobile device.
[0057] Mobile device cache: a mechanism for temporarily storing data of a mobile device, aiming to improve the data access speed and response time of a mobile application. The mobile device cache is stored in the memory or storage medium of the mobile device.
[0058] Response delay: the time delay required from the time when a user operates a mobile application on a mobile device to the time when the application responds and completes the corresponding operation, including system time delay, network time delay and application time delay, etc., wherein in most mobile application usage scenarios, the network time delay is the most important delay factor.
[0059] Optimal Dynamic Measurement Allocation Strategy (ODMAS): It is a strategy that maximizes the posterior probability of correct ranking of schemes in a dynamic environment by optimizing the simulation budget allocation decision process. This strategy is usually applied to the scheme ranking problem of stochastic complex systems, such as production manufacturing, network communication, medical health, and other cyber-physical systems.
[0060] Dynamic simulation budget optimization allocation: In the Bayesian framework, the simulation budget allocation decision process is formulated as a stochastic dynamic programming problem.
[0061] Content Delivery Network (CDN): A layer of intelligent virtual network built on the existing Internet by placing node servers at various locations in the network. It can redirect user requests to the nearest node server based on network traffic, connection status, load, distance to user, response time, and other comprehensive information, improving the response speed of user access to web resources.
[0062] Cache hit rate: When a user accesses a network resource, if the data to be accessed by the user has already been cached in the cache and does not need to be obtained from the source server on the network, it is called cache hit. Conversely, if the data to be accessed by the user does not exist in the cache, it is called cache miss, and the resource needs to be downloaded from the source server on the network. Cache hit rate = hit number / (hit number + miss number). Cache hit rate is one of the important factors to judge the good or bad of network resource request acceleration (i.e. network request optimization).
[0063] Cache overhead: The resource consumption caused by the use of cache technology, in this scheme, mainly refers to the network bandwidth and storage process overhead caused by network redundancy overhead, i.e. the resource consumption of data that does not exist in the cache and needs to be downloaded from the network and then put into the cache for use.
[0064] Cold data: Data that is rarely accessed, such as historical backup data or expired log data.
[0065] Hot data: Data that is frequently accessed or modified, such as real-time transaction data and user activity data.
[0066] Least Recently Used (LRU): A common page replacement algorithm that selects the least recently used page for eviction. This algorithm assigns each page an access field that records the time t that has elapsed since the page was last accessed. When a page must be evicted, the page with the largest t value, i.e., the least recently used page, is selected for eviction. In simple terms, this algorithm, when the available cache allocated by the mobile device to the mobile application running thereon is full, will evict the data with the lowest probability of being accessed in the future (i.e., the data that has not been accessed for the longest time) based on the access records of all data. That is, this algorithm considers that the data that has been accessed recently has the highest probability of being accessed in the future.
[0067] Least Frequently Used (LFU): A common page replacement algorithm that requires that the page with the smallest reference count be evicted when a page is replaced, because frequently used pages should have a larger reference count. However, some pages have a large reference count at the beginning but are not used later, and such pages will remain in memory for a long time. Therefore, the reference count register can be right-shifted by one bit to form an exponentially decaying average reference count. In simple terms, this algorithm, when the available cache allocated by the mobile device to the mobile application running thereon is full, will evict the data with the lowest probability of being accessed in the future (i.e., the data that has been accessed the least number of times in a certain period of time) based on the number of times each data has been accessed in a certain period of time. That is, this algorithm considers that the data that has been accessed frequently in a certain period of time has the highest probability of being accessed in the future.
[0068] Adaptive Replacement Cache (ARC): A page replacement algorithm that combines the advantages of LRU and LFU and uses more levels of replacement strategies to achieve dynamic adjustment by maintaining two lists of LRU and LFU. ARC can better adapt to different cache usage scenarios. Although ARC optimizes the shortcomings of LRU and LFU to some extent, it still only operates based on data access frequency and data access time, and is not intelligent enough in complex mobile network and mobile device scenarios.
[0069] Q-Learning: A value-based reinforcement learning algorithm based on the idea of value iteration, which estimates the value function Q of each state-action pair to guide the selection of the best action in each state. Q is Q(s, a), where s represents the state and a represents the action, and Q(s, a) represents the expected return that can be obtained by taking action a in state s at a certain time.
[0070] As explained in the background, in the rule-based cache resource management manner such as LRU, LFU and ARC, the mobile device cleans up the mobile application cache according to a fixed rule. Specifically, when it is necessary to release the cache, LRU selects the least recently used cache data to be discarded, and LFU selects the cache data with the least access frequency in a period of time to be discarded. LRU only depends on the time of the latest access to a certain cache data, and ignores the frequency of accessing the cache data, the size of the cache data and other dimensions, which may incorrectly delete the cache data with high frequency access but no access in the recent period. LFU only depends on the frequency of accessing a certain cache data in the past, and cannot timely discard the cache data with burst high frequency but no longer used, such as the cache data of hot news which is accessed with high frequency in a short time but has no access demand afterward, but due to the high frequency of access in a period of time, it is still considered to be kept in the cache.
[0071] At the same time, considering that the cache space of the mobile device is limited, therefore the utilization rate of the cache space is also an index to be considered. In terms of the utilization rate of the cache space, LRU and LFU rely on fixed rules to clean up the data, which may cause multiple small files with slightly high frequency of use to be saved, and large files with low frequency of use but long time-consuming to load from the network to be deleted, which not only causes the time-consuming of network loading, but also for the limited cache space, the utilization rate of a single large file is higher than that of multiple small files.
[0072] And in the event of a sudden situation, for example, when the user access pattern changes, LRU and LFU cannot respond intelligently in a timely manner. For example, accidental full cache data traversal operation occurs, as all cache data needs to be operated, the history access record of the cache data will be refreshed to a very short time, no matter how long the previous recent access time is, even if the cache data has not been accessed for a long time, it will also be refreshed to the latest time, causing the history access record of the cache data to be contaminated, and when some cache data needs to be cleaned up in subsequent cache data access, a large amount of data accessed in a very short time will be deleted according to the rule, regardless of whether the "useless" cache data is refreshed due to full traversal or the "useful" cache data is actually frequently accessed. As for LFU, it will have problems when dealing with sudden sparse traffic. Because the cache data that is frequently accessed in the early stage, that is, the cache data with a high access frequency, has a high position in the access frequency record table, the sudden sparse traffic cannot be compared with the access frequency of these cache data that already occupies a high position in the access frequency record table, and it is difficult to be retained by the LFU mode. These past high-frequency access cache data may not be accessed in the future, but they will occupy cache space for a long time due to their historical access frequency. ARC optimizes LRU and LFU in the case of sudden changes in user access patterns, and does not cause a large amount of cache data to be deleted and long-time single high-frequency cache data to occupy space, but it is still optimized based on access frequency and access time, and does not fundamentally consider multiple factors to improve cache data hit rate.
[0073] Therefore, the above-mentioned cache resource management scheme based on specific rules (equivalent to a static cache allocation strategy) is difficult to adapt to dynamically changing user access patterns and resource demands, resulting in a decrease in cache hit rate and a waste of cache resources.
[0074] On the other hand, the Q-learning-based scheme also has the following problems:
[0075] Firstly, the real-time performance is not strong, and the computational overhead is large. The model training of Q-learning relies on a large amount of interaction data, and the training algorithm resource demand is high, and it is usually impossible to deploy training locally on mobile devices, but only the model trained by the cloud can be sent to the mobile device. When making real-time decisions on whether to delete cache data, Q values need to be calculated to determine, but if there are many cache data, the calculation cost of Q values will increase significantly, increasing the operation delay of cache data and affecting the performance in high-concurrency scenarios.
[0076] Secondly, the interpretability is not strong. The Q table of Q-learning is difficult to intuitively explain the strategy logic. If it is found that there is a problem in the large-scale mobile application cache processing, the strategy cannot be quickly adjusted, and usually the model optimization training needs to be performed again. At the same time, due to the black box nature of Q-Learning and other reinforcement learning methods, if unexpected behavior occurs during model training, such as why the model excessively biases towards deleting or retaining a certain type of cache data, it is difficult to debug even if you want to.
[0077] Thirdly, the ability to cope with access mode mutations is relatively weak. At the initial stage of building a mobile application cache cleaning system, due to insufficient amount of collected data, Q-Learning and other methods that simply rely on reinforcement learning will produce high strategy randomness, and a large amount of actual use data needs to be accumulated (i.e. a large number of trial and error) before it can converge, which may cause the mobile application cache hit rate to drop sharply at the initial stage. At the same time, for sudden changes in access mode, such as a large amount of cache data caused by a sudden surge in traffic, Q-Learning relies on exploration mechanisms such as greedy algorithms (such as ε-greedy algorithm) to adapt, but the operation of the exploration mechanism itself also needs adaptation time.
[0078] Fourthly, the stability is insufficient. Q-Learning and other pure reinforcement learning methods will cause the number of system states to increase exponentially when the system size, especially the concurrency, increases. In the context of mobile application cache optimization, if a large amount of fragmented cache data needs to be processed in parallel, this situation may trigger state space explosion, causing the model detection result to fail to converge.
[0079] Fifthly, the scene applicability is insufficient. Q-Learning has a high theoretical upper limit in terms of ability, as it can optimize long-term returns through cache preheating based on predictions of future user operation requests, but in actual use, considering the limited computing power on mobile devices, the computing power available for cache cleaning is very limited. In the case of tight computing power, the scenarios that can be applied by Q-Learning and other pure machine learning driven methods are limited.
[0080] In order to solve the defects of the existing cache resource management scheme, the embodiment provides a mobile application cache resource optimization method, which is applied to a mobile device, please see Figure 1 The flowchart of the method can include the following steps.
[0081] S101, obtaining a first value parameter, the first value parameter being obtained by a server from cache processing data of a plurality of mobile devices according to a strategy value network.
[0082] In step S101, the service end can periodically update the first value parameter according to a certain update period, and automatically issue the latest first value parameter to the mobile device after updating the first value parameter each time. Alternatively, the service end can save the updated first value parameter after updating, and the mobile device periodically requests the first value parameter from the service end, and each time the request is received, the service end sends the latest determined first value parameter to the mobile device in response to the request of the mobile device.
[0083] The policy value network can include two parts of a policy network (also referred to as an actor) and a value network (also referred to as a critic), both of which are deployed on the service end and called by the service end to process the cache processing data reported by each mobile device to obtain the first value parameter.
[0084] In step S102, the cache value of the cache data is obtained by calculating the key data indicators corresponding to the cache data of the mobile device according to the first value parameter, and the key data indicators are part of the multiple data indicators.
[0085] In step S102, the mobile device can store multiple cache data, each of which has corresponding multiple data indicators, part of which have high correlation with the cache value of the cache data, and thus belong to key data indicators.
[0086] After obtaining the first value parameter, for each cache data, the mobile device can calculate the key data indicators of the cache data according to the first value parameter to obtain the cache value corresponding to the cache data.
[0087] Optionally, the multiple data indicators can include:
[0088] Access frequency, data freshness, storage cost, predicted access probability, network delay, cache occupancy rate, recent access pattern, historical cache hit rate and cache data update times;
[0089] The key data indicators can include the access frequency, data freshness, storage cost and predicted access probability in the above multiple data indicators.
[0090] Access frequency (Access Frequency, AF), which represents the access frequency of the cache data in a recent period of time.
[0091] Data freshness (Data Freshness, DF), which represents the existence time of the cache data, mainly for protecting newly added cache data within a certain period of time, because the newly added cache data may not be suitable for the existing cache cleaning mode within a certain period of time because the access frequency and other data have not been updated. Data freshness gradually decreases with time.
[0092] Storage Cost (SC), representing resource consumption of storing data downloaded from the network into the local cache, the larger the file, the higher the storage cost.
[0093] Predicted Access Probability (PAP), which can be calculated by using any algorithm capable of predicting the probability of data being accessed in the prior art, for example, using ARC algorithm based on historical access, combining the access frequency and access time of the cached data to predict the access probability.
[0094] For any piece of cached data, the above four key data indicators of the piece of cached data are calculated using the first value parameter according to the following formula (1) to obtain the cache value of the piece of cached data.
[0095] V (d) = C1*AF (d) + C2*DF (d) - C3*SC (d) + C4*PAP (d), (1).
[0096] Wherein, d represents a piece of cached data stored in the cache module of the mobile device, V (d) represents the cache value of the piece of cached data, AF (d), DF (d), SC (d) and PAP (d) represent the access frequency, data freshness, storage cost and predicted access probability of the piece of cached data in turn; C1, C2, C3 and C4 are four weight parameters contained in the first value parameter, which are weight parameters corresponding to the access frequency, weight parameters corresponding to the data freshness, weight parameters corresponding to the storage cost and weight parameters corresponding to the predicted access probability in turn.
[0097] Optionally, after calculating the cache value of a piece of cached data, the mobile device can record the cache value of the piece of cached data, the piece of cached data itself, and the multiple data indicators corresponding to the piece of cached data into the cache value record table, so as to report the data recorded in the cache value record table as part of the cache processing data to the server at a certain period.
[0098] The cache value record table can be stored in the memory of the mobile device other than the cache module.
[0099] S103, at least one piece of cached data is deleted in turn in ascending order of cache value, so that the total data amount of cached data in the mobile device matches the cache space capacity.
[0100] Step S103 can be executed only when the total data amount of cached data in the mobile device and the cache space capacity do not match, and S103 can be temporarily not executed if the two match.
[0101] The total data amount of the cached data and the cache space capacity do not match, which can mean that the total data amount of the cached data is equal to the upper limit of the cache space capacity, or that the total data amount of the cached data is close to the upper limit of the cache space capacity, i.e., the difference between the total data amount of the cached data and the upper limit of the cache space capacity is less than a preset threshold.
[0102] The total data amount of the cached data and the cache space capacity match, which can mean that the total data amount of the cached data is less than the upper limit of the cache space capacity, or that the total data amount of the cached data is significantly less than the upper limit of the cache space capacity, i.e., the difference between the total data amount of the cached data and the upper limit of the cache space capacity is greater than a preset threshold.
[0103] When performing S103, the mobile device can first delete a cached data with the lowest cache value from the cache module, and then determine whether the total data amount of the remaining cached data matches the cache space capacity. If not, the cached data with the lowest cache value is continuously deleted from the remaining cached data. If not, the next cached data is deleted, and so on, until the total data amount of the remaining cached data matches the cache space capacity after deleting a certain cached data. Then, the deletion is stopped, and the optimization of the cache resource is completed.
[0104] The embodiment has the following beneficial effects:
[0105] On the one hand, the cache value is efficiently determined on the mobile device through the first value parameter, so that the unnecessary cached data is accurately deleted based on the cache value, and the cache space is optimized. On the other hand, the server can dynamically determine the first value parameter according to the cache processing data, so as to adapt to the dynamically changing environment.
[0106] Optionally, after the at least one cached data is deleted in the ascending order of the cache value, the method further includes:
[0107] When it is detected that the deleted cached data causes the cache miss problem, the cached data causing the cache miss problem is restored.
[0108] In the above embodiments, the mobile device can monitor the access requests of each mobile application to the cache module in real time. If the mobile device monitors that a piece of data requested by the mobile application to be accessed does not exist in the cache module, and the piece of data can be found in the cache value record table, it is indicated that the piece of data is the data deleted from the cache module by the mobile device in step S103, and the piece of data causes the cache miss problem. In this case, the mobile device can determine that the piece of data belongs to the cache data of the cache miss problem (referred to as cache miss data), and then record the cache miss data itself, the corresponding cache value and a plurality of data indicators in the cache miss record table, and on the other hand, the cache miss data can be recovered, that is, written into the cache module again, so as to be used by the mobile application.
[0109] If the mobile device monitors that a piece of data requested by the mobile application to be accessed does not exist in the cache module, and the piece of data is not recorded in the cache value record table, it is indicated that the piece of data is not deleted from the cache module by the mobile device, but does not originally exist in the cache module. In this case, the mobile device can obtain the piece of data in a plurality of ways (for example, downloading from the network, reading from the hard disk), write the piece of data into the cache module for use by the mobile application, and calculate the cache value of the piece of data, and record the piece of data itself, the corresponding cache value and a plurality of data indicators in the cache value record table.
[0110] Optionally, after the at least one cache data is deleted in the ascending order of the cache values, the method further comprises:
[0111] monitoring the cache overhead indicator and the cache hit rate indicator;
[0112] when the cache overhead indicator and the cache hit rate indicator meet a target update condition, updating the first value parameter according to the data indicators corresponding to each cache data in a preset time period.
[0113] The cache overhead indicator can be defined as a ratio of the occupied cache space in the cache module to the total cache space capacity.
[0114] The cache hit rate indicator can be defined as a ratio of the number of hit requests in a unit time (such as 1 minute) to the total number of cache requests in the unit time. The hit request refers to an access request to the cache module and successfully finding the required data in the cache module. The total number of cache requests refers to the total number of all access requests to the cache module generated by the mobile application in the unit time.
[0115] The preset time period can be a time period in the latest preset time length as of the current time, for example, can be the latest 5 minutes, the latest 20 minutes or other time lengths as of the current time, without limitation.
[0116] The current time can refer to a time at which the mobile device determines that the cache overhead indicator and the cache hit rate indicator meet the target update condition.
[0117] If the cache overhead indicator and the cache hit rate indicator do not meet the target update condition, the mobile device can continue to monitor the cache overhead indicator and the cache hit rate indicator, and keep the current first value parameter unchanged, unless a new first value parameter is issued by the server.
[0118] The cache overhead indicator and the cache hit rate indicator meeting the target update condition can include both the cache overhead indicator and the cache hit rate indicator significantly increasing, or both the cache overhead indicator and the cache hit rate indicator significantly decreasing.
[0119] In other words, if both indicators significantly increase, it is determined that the target update condition is met, and if both indicators significantly decrease, it is also determined that the target update condition is met; if at least one of the indicators does not significantly increase or decrease, or one significantly increases and the other significantly decreases, it can be determined that the target update condition is not met.
[0120] Specifically, if the decrease amplitude of the cache overhead indicator is greater than a first threshold value within a preset time period, it can be determined that the cache overhead indicator significantly decreases, and if the decrease amplitude of the cache hit rate is greater than a second threshold value within the preset time period, it can be determined that the cache hit rate indicator significantly decreases. Therefore, if the decrease amplitude of the cache overhead indicator is greater than the first threshold value and the decrease amplitude of the cache hit rate is greater than the second threshold value within the preset time period, it can be determined that the cache overhead indicator and the cache hit rate indicator meet the target update condition.
[0121] If the increase amplitude of the cache overhead indicator is greater than a third threshold value within a preset time period, it can be determined that the cache overhead indicator significantly increases, and if the increase amplitude of the cache hit rate is greater than a fourth threshold value within the preset time period, it can be determined that the cache hit rate indicator significantly increases. Therefore, if the increase amplitude of the cache overhead indicator is greater than the third threshold value and the increase amplitude of the cache hit rate is greater than the fourth threshold value within the preset time period, it can be determined that the cache overhead indicator and the cache hit rate indicator meet the target update condition.
[0122] The first threshold value, the second threshold value, the third threshold value and the fourth threshold value can be the same or different, and the above threshold values used by different mobile devices can be the same or different.
[0123] Optionally, the mobile device can update the first value parameter according to the size of the data indicator corresponding to the cached data in any one or more of the following multiple adjustment modes:
[0124] The first adjustment mode is to lower the weight parameter corresponding to the storage cost in the first value parameter if the average of the storage cost of the cache miss data is greater than the storage cost threshold; the lowering amplitude can be a preset first amplitude;
[0125] The second adjustment mode is to increase the weight parameter corresponding to the access frequency in the first value parameter if the average of the access frequency of the cache miss data is greater than the access frequency threshold; the increasing amplitude can be a preset second amplitude;
[0126] The third adjustment mode is to increase the weight parameter corresponding to the data freshness in the first value parameter if the average of the data freshness of the cache miss data is greater than the data freshness threshold; the increasing amplitude can be a preset third amplitude;
[0127] The fourth adjustment mode is to increase the weight parameter corresponding to the predicted access probability in the first value parameter if the average of the predicted access probability of the cache miss data is greater than the predicted probability threshold; the increasing amplitude can be a preset fourth amplitude.
[0128] The first amplitude, the second amplitude, the third amplitude and the fourth amplitude can be the same or different, and the amplitudes can be the same or different between different mobile devices.
[0129] In the first adjustment mode, the mobile device can obtain the storage costs of the plurality of cache miss data recorded in the cache miss record table in a preset time period, calculate the average of the storage costs, and if the average is greater than the preset storage cost threshold, it indicates that the average of the current storage cost is high, which means that the cache data of the file with a large file size is more likely to be mistakenly deleted and needs to be downloaded again, so the weight parameter corresponding to the storage cost in the first value parameter on the mobile device can be appropriately lowered, that is, C3 in formula (1) is lowered, so that the cache data with a large data volume has a higher cache value and is less likely to be deleted.
[0130] In the second adjustment mode, the mobile device can obtain the access frequencies of the plurality of cache miss data recorded in the cache miss record table in a preset time period, calculate the average of the access frequencies, and if the average is greater than the preset access frequency threshold, it indicates that the average of the current access frequency is high, which means that the cache data with a high access frequency in a recent period of time is more likely to be mistakenly deleted and needs to be downloaded again, so the weight parameter corresponding to the access frequency in the first value parameter on the mobile device can be appropriately increased, that is, C1 in formula (1) is increased, so that the cache data with a high access frequency in a recent period of time has a higher cache value and is less likely to be deleted.
[0131] In the third adjustment mode, the mobile device can obtain the data freshness of the plurality of cache miss data recorded in the preset time period from the cache miss record table, calculate the average of the data freshness, and if the average is greater than the preset data freshness threshold, it indicates that the average is large, and in a certain period of time, the newly added cache data is more likely to be mistakenly deleted, resulting in the need to be downloaded again. Therefore, the weight parameter corresponding to the data freshness in the first value parameter, i.e., C2 in the above formula (1), is appropriately increased, so that the newly added cache data has a higher cache value and is less likely to be deleted.
[0132] In the fourth adjustment mode, the mobile device can obtain the predicted access probability of the plurality of cache miss data recorded in the preset time period from the cache miss record table, calculate the average of the predicted access probability, and if the average is greater than the preset predicted probability threshold, it indicates that the average is large, and in a certain period of time, the cache data with a higher predicted access probability calculated based on the ARC algorithm is more likely to be mistakenly deleted, resulting in the need to be downloaded again. Therefore, the weight parameter corresponding to the predicted access probability in the first value parameter, i.e., C4 in the formula (1), is appropriately increased, so that the cache data with a higher predicted access probability calculated based on the ARC algorithm has a higher cache value and is less likely to be deleted.
[0133] The data freshness threshold, the predicted probability threshold, the access frequency threshold, and the storage cost threshold can all be fixed thresholds preset according to actual conditions. The thresholds used by different mobile devices can be the same or different.
[0134] Each mobile device can periodically report the data recorded in the local cache value record table and the cache miss record table to the server as cache processing data of the mobile device at a certain period, so that the server updates the first value parameter and the network parameter of the strategy value network used to determine the first value parameter according to the cache processing data.
[0135] The embodiment also provides a mobile application cache resource optimization method applied to a server, please refer to Figure 2 The method can include the following processes.
[0136] S201, obtaining cache processing data of a plurality of mobile devices, the cache processing data at least including a plurality of data indicators of cache data of the mobile device and a cache miss record table of the mobile device, the cache miss record table being used to record cache miss data of the mobile device.
[0137] In step S201, the server can periodically send a request to each mobile device in communication connection with itself at a certain reporting period to receive the cache processing data fed back by the mobile device in response to the request.
[0138] When the mobile device feeds back the cache processing data each time, it can read out from the local cache value record table each piece of data recorded since the last time of feeding back the cache processing data, the cache value corresponding to the data, and multiple data indicators, and read out from the local cache miss record table each piece of data recorded since the last time of feeding back the cache processing data, the cache value corresponding to the data, and multiple data indicators, and send the read-out data to the server as the cache processing data of the mobile device.
[0139] In S202, the cache processing data is processed according to the policy network to obtain the value parameter corresponding to the cache processing data.
[0140] The policy network (also referred to as an actor) is a part of the policy value network pre-deployed on the server, responsible for generating a policy, aiming to maximize the expected return. The policy network can receive an input state s t and output a corresponding processing policy after processing the input state. In this embodiment, each processing policy can be understood as a set of value parameters. In other words, one processing policy can include four weight parameters, i.e., a weight parameter corresponding to the storage cost, a weight parameter corresponding to the access frequency, a weight parameter corresponding to the data freshness, and a weight parameter corresponding to the predicted access probability. Different processing policies can have different values of the weight parameters.
[0141] In S202, the server can obtain the value parameter corresponding to the cache processing data in the following manner:
[0142] The multiple data indicators included in the cache processing data are combined into an input state, so as to process the input state according to the policy network to obtain the value parameter corresponding to the cache processing data.
[0143] In some embodiments, the server can calculate the average value of each data indicator corresponding to the cache processing data, and then combine the obtained multiple average values into a vector, and input the vector into the policy network as an input state to obtain a processing policy (i.e., a set of value parameters) output by the policy network.
[0144] For example, assuming that 100 pieces of data are recorded in the cache processing data, the server can calculate the average value of the access frequency, the average value of the data freshness, the average value of the storage cost, the average value of the predicted access probability, the average value of the network delay, the average value of the cache occupancy rate, the average value of the recent access pattern, the average value of the historical cache hit rate, and the average value of the cache data update frequency, and combine the nine average values into a 9-dimensional vector, and input the 9-dimensional vector as an input state.
[0145] In some embodiments, the server can also form, for each piece of data recorded in the cache processing data, a state vector of the data corresponding to a plurality of data indicators of the data, thereby obtaining a plurality of state vectors, and input a set of the plurality of state vectors to the policy network as an input state to obtain the value parameter output by the policy network.
[0146] For example, assuming that there are 100 pieces of data recorded in the cache processing data, the server can form a 9-dimensional state vector for each piece of data, with the access frequency, data freshness, storage cost, predicted access probability, network delay, cache occupancy, recent access pattern, historical cache hit rate, and cache data update times of the data as the dimensions of the state vector, thereby obtaining 100 state vectors, and input a set of the 100 state vectors to the policy network as an input state.
[0147] S203, obtaining a current data reward and an expected data reward according to the value network processing the cache processing data and the expected cache processing data, the expected cache processing data being cache processing data obtained after the mobile device optimizes the cache resource based on the value parameter, the value network and the policy network forming a policy value network.
[0148] The value network (also referred to as Critic) is another part of the policy value network pre-deployed on the server, responsible for evaluating the value of the current policy and guiding the policy update.
[0149] In S203, on the one hand, the input state s t can be determined according to the cache processing data and the method of S202, and on the other hand, the expected input state s t+1 can be determined based on the expected cache processing data and the method of S202. t t t+1 t+1 The state value V(s t+1 output by the value network is taken as the current data reward. The expected input state s t+1 is input to the value network to obtain the state value V(s t+1 output by the value network, which is taken as the expected data reward.
[0150] One way to obtain the expected cache processing data is to input the cache processing data and the value parameter output by the policy network to the value network, and the value network processes the cache processing data and the value parameter to predict the expected cache processing data.
[0151] Another way to obtain the expected cache processing data is for the server to wait for a reporting period and take the cache processing data reported by the mobile device next time as the expected cache processing data.
[0152] S204, updating the network parameters of the policy network and the network parameters of the value network according to the current data reward, the expected data reward and the actual reward value, to complete the training of the policy network and the value network, the actual reward value being determined according to the cache hit rate and the cache hit rate in the cache processing data.
[0153] In step S204, the server can calculate the network loss of the policy network by using the current data reward, the expected data reward and the actual reward value according to the following formulas (2) and (3).
[0154]
[0155]
[0156] In formula (2), represents the network loss of the policy network, and θ represents the gradient of the network parameters contained in the policy network, and πθ [] represents a preset loss function calculation formula, the specific expression of which can be referred to existing technical documents in the field of neural network technology, and will not be described herein. log represents the logarithm with base 10. represents the policy output by the policy network currently, that is, the value parameter corresponding to the cache processing data obtained by processing the cache processing data in step S202. δ represents the temporal-difference error (TD error), which can be calculated according to formula (3).
[0157] In formula (3), R t represents the actual reward value, and P1 is a preset discount factor, the specific value of which can be set as needed and is not limited. V(s t ) represents the current data reward, and V(s t+1 ) represents the expected data reward.
[0158] The actual reward value R t can be calculated according to the following formula (4).
[0159] R t = Cache Hit Rate-P2*Storage Overhead, (4).
[0160] Wherein, the Cache Hit Rate represents the cache hit rate of the mobile device, the Storage Overhead represents the cache occupation rate of the mobile device, if the cache hit rate is higher, the reward function is higher, P2 is a preset positive number, used to quantify the cache occupation rate and the cache hit rate, the lower the cache occupation rate is, the higher the reward function is. The establishment of the reward function is to expect to achieve the maximum cache hit rate with the least cache occupation rate. The determination method of the cache hit rate and the cache occupation rate can be referred to related prior art, and will not be described here.
[0161] The server can calculate the network loss of the value network according to the following formula (5).
[0162]
[0163] In formula (5), L (μ) represents the network loss of the value network.
[0164] In the training, t training rounds are performed, that is, In the state,
[0165] In each training, the expected input state s t+1 and the actual reward R t are determined according to the value parameter output by the current policy network, and then the network loss of the policy network and the network loss of the value network are calculated according to the above formulas (2) to (5).
[0166] If both of the above network losses are less than a preset network loss threshold, it can be determined that the training is completed, and step S205 is entered, if at least one of the above network losses is greater than or equal to the preset network loss threshold, the network parameters of the policy network are optimized according to the network loss of the policy network, and the network parameters of the value network are optimized according to the network loss of the value network, then when the cache processing data of the mobile device is received next time, the network loss is calculated according to the above steps to continue the optimization according to the network loss, until both of the above network losses are less than the preset network loss threshold.
[0167] Wherein, when the corresponding network parameters are optimized according to the network loss, the policy gradient (Policy Gradient) algorithm can be used for optimization, or other algorithms can be used for optimization, which is not limited in the embodiment.
[0168] S205, the cache processing data of the mobile device is processed according to the trained policy network, and the first value parameter obtained by processing is sent to the mobile device, so that the mobile device optimizes the cache resources based on the first value parameter.
[0169] In step S205, after the training of the policy network and the value network is completed, the server can periodically collect the cache processing data of the mobile device, call the policy network to process the collected cache processing data according to the method in S202, and then distribute the value parameter output by the policy network as the first value parameter to the mobile device, so that the mobile device deletes the unnecessary cache data according to the first value parameter, and optimizes the cache space. Figure 1 The method shown optimizes the local cache resources.
[0170] The embodiment has the following beneficial effects:
[0171] On the one hand, the cache value is efficiently determined on the mobile device through the first value parameter, so that the unnecessary cache data is accurately deleted based on the cache value, and the cache space is optimized, and on the other hand, the server can dynamically determine the first value parameter according to the cache processing data, so as to adapt to the dynamically changing environment.
[0172] Optionally, since the first value parameter is determined according to the cache processing data of the mobile device, the server can determine the first value parameter specially applicable to each mobile device. In other words, the first value parameter distributed by the server can be different for different mobile devices. In this way, the cache value calculated by the mobile device can be more accurate.
[0173] Alternatively, different mobile devices can also share the same set of first value parameters, which can reduce the load of the server.
[0174] Optionally, when the method of the embodiment is first applied to cache optimization, if the cache processing data cannot be collected, the server can first distribute the initially set first value parameter to the mobile device, and then continuously update the first value parameter according to the received cache processing data during subsequent cache optimization.
[0175] Specifically, compared with the existing cache optimization method, the effect achieved by the present scheme is to realize efficient dynamic cache allocation while ensuring low computational overhead.
[0176] First, compared with rule-based modes such as LRU, LFU and ARC. Under these rule-based modes, the mobile application cache cleaning is performed according to fixed rules. When it is necessary to release the cache, LRU selects the least recently used cache data to be deleted, and LFU selects the cache data with the least access frequency in the past period of time to be deleted. LRU only depends on the time of the most recent access to a certain cache data, and ignores the frequency of accessing the cache data, the size of the cache data and other dimensions, and may incorrectly delete the cache data that is accessed frequently but has not been accessed recently. LFU only depends on the frequency of accessing a certain cache data in the past, and cannot timely delete cache data that is accessed frequently in bursts but is not used thereafter, such as hot news cache data that is accessed frequently in bursts for a short time but is not accessed thereafter, but due to the high access frequency in the past period of time, it is still considered to be needed to be retained in the cache.
[0177] At the same time, considering that the cache space of a mobile device is limited, therefore the utilization rate of the cache space is also an index to be considered. In terms of the utilization rate of the cache space, LRU / LFU relies on fixed rules to clean up data, which may cause multiple small files with slightly high use frequency to be saved, and large files with low use frequency but long time-consuming for loading from the network to be deleted, which not only causes the time-consuming of network loading, but also, for the limited cache space, the utilization rate of a single large file is higher than that of multiple small files.
[0178] And, in the event of a sudden situation, the user access mode changes, LRU / LFU can not be timely and intelligent response. For example, the occurrence of accidental want to traverse the entire cache data, because the need to operate on all cache data, so the history of access records of cache data will be refreshed to a very short time, no matter how long the previous recent access time, even if a long time not accessed cache data will be refreshed to the latest time, resulting in cache data history access records contaminated, may in subsequent cache data access need to clean some cache data, according to the rules will be a large number of deletion of this one short time concentrated access of the entire data, whether the actual is a long time not accessed only because the entire traversal refresh time "useless" cache data, or is itself is often accessed "useful" cache data. And for LFU, when facing a sudden sparse traffic problems will occur. Because the early access, that is, the access frequency of cache data is higher, the access frequency record table has a higher position, the burst sparse traffic due to the access frequency and the access frequency of these cache data already occupying the access frequency record table higher position can not be compared, so it is difficult to be retained by the LFU mode. And these past high-frequency access cache data in the future may not necessarily be accessed, but will be in cache by the history of access frequency in a long period of time to occupy cache space. ARC for LRU / LFU in the event of a sudden user access mode changes have been optimized, will not cause cache data concentrated large batch deletion and long time single high-frequency cache data placeholder, but still only based on access frequency and access time two factors for optimization, and not fundamentally consider multiple factors to improve cache data hit rate.
[0179] And the cache optimization method of the embodiment can consider multiple key data indicators in real time, for example, considering access frequency, data freshness (reflecting time limitation), storage cost (reflecting cache data size and loading cost) and predicted access probability, and dynamically adjusting the cache value of each cache data. For example, using the cache optimization method of the embodiment, it is more likely to save low-frequency use but high cache loading cost large file cache. In LRU or LFU mode, such large files are easily removed, but if you want to load large files from the network, it is time-consuming. And the cache optimization method of the embodiment combines the "new priority" of LRU and the "frequency priority" of LFU, introduces data freshness (DF) and storage cost (SC), adds protection for data cold start (i.e. processing of newly received cache operation record information), balances the cache priority of cold data and hot data, and avoids the extreme problem caused by LRU and LFU based on single index cache optimization.
[0180] Second, compared with machine learning driven patterns such as Q-Learning, the cache optimization scheme of the embodiment has advantages in real-time performance, computational overhead, interpretability, ability to cope with access pattern mutations, stability and scenario applicability.
[0181] First, comparison in real-time performance and computational overhead:
[0182] The model training of Q-learning relies on a large amount of interaction data and has high demand for training computing resources. It is usually impossible to deploy training locally on a mobile device, and only the model trained by the cloud can be downloaded to the mobile device. When making real-time decisions on whether to delete cache data, the Q value needs to be calculated to determine whether to delete the cache data. However, if there is a large amount of cache data, the calculation cost of the Q value will be significantly increased, the operation delay of the cache data will be increased, and the performance in a high-concurrency scenario will be affected.
[0183] The cache optimization method described in the present solution directly performs simple arithmetic operations based on the first value parameter pre-defined by the server. The time required for decision-making in a large number of mobile application cache operation scenarios with high real-time requirements is lower than that of Q-Learning.
[0184] Second, comparison in interpretability:
[0185] The Q table used by Q-learning is difficult to intuitively explain the strategy logic. If a large-scale mobile application cache processing problem is found, the strategy cannot be quickly adjusted, and model optimization training usually needs to be performed again. At the same time, due to the black box nature of Q-Learning and other reinforcement learning methods, if unexpected behavior occurs during model training, such as the model excessively favoring the deletion or retention of a certain type of cache data for unknown reasons, it is difficult to debug even if you want to.
[0186] Although the cache optimization method of the embodiment also uses reinforcement learning to optimize the first value parameter, the final expression can intuitively show what each parameter represents, and dynamic fine-tuning can be performed according to actual work effect feedback. For example, if it is found that there is a deficiency in the retention of large file cache data in actual mobile application cache processing work, and a large number of large file cache data need to be downloaded from the network, the weight parameter corresponding to the storage cost can be directly adjusted to increase the proportion of large files, thereby reducing the situation of incorrect cleaning of large file cache.
[0187] Third, comparison in ability to cope with access pattern mutations:
[0188] The mobile application cache cleaning system is built in the early stage, and the amount of collected data is insufficient. The pure reinforcement learning mode such as Q-Learning has a high strategy randomness, and needs a large amount of actual use data accumulation (that is, a large number of trial and error) to converge, which may cause the mobile application cache hit rate to drop sharply in the early stage. At the same time, for the sudden change of access mode, such as a large amount of cache data caused by a large amount of traffic, Q-Learning relies on the greedy exploration mechanism to adapt, but the operation of the exploration mechanism itself also needs an adaptation time.
[0189] The cache optimization method described in the present solution can rely on real-time monitored indicators such as large file cache hit rate to directly and real-time adjust the weight parameters in the first value parameter in the early stage of system construction, and realize rapid response to changes. The initial parameters can be updated after reinforcement learning after the server obtains a large amount of actual use data.
[0190] Fourth, the comparison of stability:
[0191] The pure reinforcement learning method such as Q-Learning may cause state space explosion and result in model detection results that cannot converge when the system state number increases exponentially when the system size increases, especially the concurrency. In the mobile application cache optimization scenario, if a large amount of fragmented cache data needs to be processed in parallel, this situation may be triggered to cause state space explosion, resulting in model detection results that cannot converge.
[0192] The cache optimization method of the present embodiment is stable in output because it uses a deterministic algorithm without random exploration.
[0193] Fifth, comparison of scene applicability:
[0194] Q-Learning can achieve long-term algorithmic optimization by predicting future user operation requests and preheating cache, and the theoretical upper limit exceeds the theoretical upper limit of the cache optimization method described in the present solution, without considering the computing power and the very complex access mode of the system cache of the AI recommendation system. However, in actual use, considering the limited computing power on mobile devices and the relatively simple access mode of cache on mobile devices.
[0195] The cache optimization method of the present embodiment can achieve a better balance between efficiency and complexity, and can better meet the needs of high real-time processing of a large number of concurrent fragmented cache operations in a resource-limited environment on a mobile device, especially in scenarios that need to respond to a large number of special requests in a short time, such as a large number of large file content pushes in a short time, and simple manual intervention can be used to respond to these scenarios.
[0196] The present embodiment also provides a mobile application cache resource optimization device, please refer toFigure 3 The device can comprise:
[0197] The obtaining unit 301 is configured to obtain a first value parameter, the first value parameter being obtained by a server according to a policy value network processing cache processing data of a plurality of mobile devices;
[0198] The computing unit 302 is configured to perform value calculation on key data indicators corresponding to the cache data of the mobile device according to the first value parameter, to obtain cache values of the cache data, the key data indicators being part of a plurality of data indicators;
[0199] The deleting unit 303 is configured to sequentially delete at least one item of cache data in ascending order of cache values, so that the total data amount of the cache data in the mobile device matches the cache space capacity.
[0200] Optionally, after the deleting unit 303 sequentially deletes at least one item of cache data in ascending order of cache values, the deleting unit 303 is further configured to:
[0201] When it is detected that the deleted cache data causes cache miss problems, the cache data causing the cache miss problems is recovered.
[0202] Optionally, after the deleting unit 303 sequentially deletes at least one item of cache data in ascending order of cache values, the deleting unit 303 is further configured to:
[0203] Monitor cache overhead indicators and cache hit rate indicators;
[0204] When the cache overhead indicators and the cache hit rate indicators meet target update conditions, update the first value parameter according to data indicators corresponding to each cache data in a preset time period.
[0205] Optionally, when the deleting unit 303 updates the first value parameter according to data indicators corresponding to each cache data in a preset time period, the deleting unit 303 is configured to at least one of:
[0206] If an average value of storage costs of cache miss data is greater than a storage cost threshold, down-regulate a weight parameter corresponding to the storage cost in the first value parameter;
[0207] If an average value of access frequencies of the cache miss data is greater than an access frequency threshold, up-regulate a weight parameter corresponding to the access frequency in the first value parameter;
[0208] If an average value of data freshness of the cache miss data is greater than a data freshness threshold, up-regulate a weight parameter corresponding to the data freshness in the first value parameter;
[0209] If an average value of predicted access probabilities of the cache miss data is greater than a predicted probability threshold, up-regulate a weight parameter corresponding to the predicted access probability in the first value parameter.
[0210] Optionally, the deleting unit 303 determines that the cache overhead indicator and the cache hit rate indicator meet the target updating condition, and is configured to:
[0211] a decrease amplitude of the cache overhead indicator is greater than a first threshold value and a decrease amplitude of the cache hit rate is greater than a second threshold value within a preset time length;
[0212] or, an increase amplitude of the cache overhead indicator is greater than a third threshold value and an increase amplitude of the cache hit rate is greater than a fourth threshold value within the preset time length.
[0213] Optionally, the plurality of data indicators include:
[0214] access frequency, data freshness, storage cost, predicted access probability, network delay, cache occupancy rate, recent access pattern, historical cache hit rate and cache data update times;
[0215] The key data indicators include:
[0216] access frequency, data freshness, storage cost and predicted access probability.
[0217] The embodiment further provides a mobile application cache resource optimization device applied to a server, please refer to Figure 4 The device comprises:
[0218] An obtaining unit 401 is configured to obtain cache processing data of a plurality of mobile devices, the cache processing data at least including a plurality of data indicators of cache data of the mobile devices and a cache miss record table of the mobile devices, the cache miss record table being configured to record cache miss data of the mobile devices.
[0219] A parameter unit 402 is configured to process the cache processing data according to a strategy network to obtain a value parameter corresponding to the cache processing data.
[0220] A reward unit 403 is configured to process the cache processing data and expected cache processing data according to a value network to obtain a current data reward and an expected data reward, the expected cache processing data being cache processing data obtained after the mobile devices perform cache resource optimization based on the value parameter, the value network and the strategy network forming a strategy value network.
[0221] An updating unit 404 is configured to update network parameters of the strategy network and network parameters of the value network according to the current data reward, the expected data reward and an actual reward value to complete training of the strategy network and the value network, the actual reward value being determined according to storage cost and historical cache hit rate in the cache processing data.
[0222] The parameter unit 402 is configured to process the cache processing data of the mobile device according to the trained strategy network, and send a first value parameter obtained by processing to the mobile device, so that the mobile device optimizes the cache resource based on the first value parameter.
[0223] Optionally, when the parameter unit 402 processes the cache processing data according to the strategy network and obtains the value parameter corresponding to the cache processing data, the parameter unit 402 is configured to:
[0224] The parameter unit 402 is configured to process the cache processing data of the mobile device according to the trained strategy network, and send a first value parameter obtained by processing to the mobile device, so that the mobile device optimizes the cache resource based on the first value parameter.
[0225] The working principle of the mobile application cache resource optimization device can refer to the related steps of the mobile application cache resource optimization method in the foregoing embodiments, and details are not described herein.
[0226] It can be understood that, before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scenario, etc. of the personal information involved in the present application should be informed to the user and the authorization of the user should be obtained according to relevant laws and regulations.
[0227] For example, when responding to the active request of the user, the user is sent prompt information to explicitly prompt that the operation requested to be executed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium, etc. software or hardware that executes the operation of the technical solutions of the present application according to the prompt information.
[0228] As an optional but not limited implementation manner, the manner of sending prompt information to the user in response to the active request of the user may be, for example, a pop-up window manner, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0229] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present application, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present application.
[0230] It can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solutions should comply with the requirements of the corresponding laws and regulations and relevant provisions.
[0231] It should be noted that each embodiment in the present specification adopts a progressive description manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0232] For ease of description, the above system or apparatus is described in various modules or units respectively in terms of functions. Of course, in the implementation of the present application, the functions of each unit can be implemented in the same or more software and / or hardware.
[0233] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary universal hardware platforms. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0234] Finally, it should be noted that in this document, relational terms such as first and second and third and fourth, and the like can merely be used to distinguish one entity or action from another, without necessarily requiring or implying any actual such relationship or order between or among the entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0235] The above description is only the preferred embodiments of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A mobile application cache resource optimization method, characterized in that: Applied to a mobile device, the method includes: Obtaining a first value parameter, where the first value parameter is obtained by the server side processing cached data of a plurality of mobile devices according to a policy value network; Calculating the value of a key data indicator corresponding to the cached data of the mobile device according to the first value parameter to obtain a cache value of the cached data, where the key data indicator is a part of multiple data indicators; At least one cached data item is deleted in ascending order of cache value, so that the total amount of cached data in the mobile device matches the cache space capacity.
2. The method according to claim 1, characterized in that After deleting at least one cached data item in ascending order of cache value, the method further includes: When it is detected that the deleted cache data causes a cache miss problem, the cache data that caused the cache miss problem is restored.
3. The method according to claim 1, characterized in that After deleting at least one cached data item in ascending order of cache value, the method further includes: Monitor cache overhead metrics and cache hit ratio metrics; When the cache overhead indicator and the cache hit rate indicator meet the target update condition, the first value parameter is updated according to the data indicator corresponding to each cached data within a preset time period.
4. The method according to claim 3, characterized in that The updating of the first value parameter according to the data indicators corresponding to each cached data within a preset time period includes at least one of the following: If the average storage cost of cache miss data is greater than the storage cost threshold, lowering the weight parameter corresponding to the storage cost in the first value parameter; If the average access frequency of cache miss data is greater than the access frequency threshold, the weight parameter corresponding to the access frequency in the first value parameter is increased; If the average value of the data freshness of the cache miss data is greater than the data freshness threshold, the weight parameter corresponding to the data freshness in the first value parameter is increased; If the average value of the predicted access probability of cache miss data is greater than the predicted probability threshold, the weight parameter corresponding to the predicted access probability in the first value parameter is increased.
5. The method according to claim 3, characterized in that The cache overhead indicator and the cache hit rate indicator meet target update conditions, including: Within a preset time period, the decrease in the cache overhead indicator is greater than a first threshold, and the decrease in the cache hit rate is greater than a second threshold; Alternatively, within a preset time period, the increase in the cache overhead indicator is greater than a third threshold, and the increase in the cache hit rate is greater than a fourth threshold.
6. The method according to claim 1, characterized in that The multiple data indicators include: Access frequency, data freshness, storage cost, predicted access probability, network latency, cache occupancy, recent access patterns, historical cache hit rate, and cache data update count; The key data indicators include: Access frequency, data freshness, storage cost, and predicted access probability.
7. A mobile application cache resource optimization method, characterized in that: Applied to a server, the method includes: Obtaining cache processing data of a plurality of mobile devices, the cache processing data including at least a plurality of data indicators of cache data of the mobile devices and a cache miss record table of the mobile devices, the cache miss record table being used to record cache miss data of the mobile devices; Processing the cached processing data according to the policy network to obtain a value parameter corresponding to the cached processing data; Processing the cache processing data and the expected cache processing data according to the value network to obtain a current data reward and an expected data reward, wherein the expected cache processing data is cache processing data obtained after predicting cache resource optimization of the mobile device based on the value parameter, and the value network and the policy network constitute a policy value network; updating network parameters of the policy network and network parameters of the value network according to the current data reward, the expected data reward, and the actual reward value to complete training of the policy network and the value network, wherein the actual reward value is determined according to a cache occupancy rate and a cache hit rate in the cache processing data; The cache processing data of the mobile device is processed according to the trained policy network, and the processed first value parameter is sent to the mobile device, so that the mobile device optimizes the cache resources based on the first value parameter.
8. The method according to claim 7, characterized in that The step of processing the cached data according to the policy network to obtain a value parameter corresponding to the cached data includes: The multiple data indicators contained in the cached processing data are combined into an input state, and the input state is processed according to the strategy network to obtain the value parameters corresponding to the cached processing data.
9. A mobile application cache resource optimization device, characterized in that: Applied to a mobile device, the device comprises: an obtaining unit, configured to obtain a first value parameter, wherein the first value parameter is obtained by the server side processing cache processing data of a plurality of mobile devices according to a policy value network; a calculation unit, configured to perform value calculation on a key data indicator corresponding to the cached data of the mobile device according to the first value parameter to obtain a cache value of the cached data, wherein the key data indicator is a part of multiple data indicators; The deleting unit deletes at least one cached data item in ascending order of cache value, so that the total amount of cached data in the mobile device matches the cache space capacity.
10. A mobile application cache resource optimization device, characterized in that: Applied to a server, the device includes: an obtaining unit, configured to obtain cache processing data of a plurality of mobile devices, the cache processing data comprising at least a plurality of data indicators of cache data of the mobile devices, and a cache miss record table of the mobile devices, the cache miss record table being used to record cache miss data of the mobile devices; a parameter unit, configured to process the cache processing data according to a policy network and obtain a value parameter corresponding to the cache processing data; a reward unit, configured to process the cache processing data and the expected cache processing data according to a value network to obtain a current data reward and an expected data reward, wherein the expected cache processing data is cache processing data obtained after predicting cache resource optimization of the mobile device based on the value parameter, and the value network and the policy network constitute a policy value network; an updating unit, configured to update network parameters of the policy network and network parameters of the value network according to the current data reward, the expected data reward, and an actual reward value, so as to complete training of the policy network and the value network, wherein the actual reward value is determined according to a cache occupancy rate and a cache hit rate in the cache processing data; The parameter unit is used to process cache processing data of the mobile device according to the trained policy network, and send the processed first value parameter to the mobile device, so that the mobile device optimizes cache resources based on the first value parameter.