Cache data optimization method and device, equipment, medium and computer program product
By building a data access relationship diagram and detecting data group centrality indicators, and optimizing the cache allocation strategy, the problems of uneven resource allocation and difficulty in responding to changes in data access patterns in the existing technology are solved, and precise optimization of cached data and system performance improvement are achieved.
Patent Information
- Application Number
- CN202510077948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
AI Technical Summary
The existing cached data optimization technology lacks global optimization methods, resulting in uneven resource allocation and difficulty in responding to changes in data access patterns in a timely manner.
By constructing a data access relationship diagram, detect the centrality indicators of the data group, determine the hotspot potential value of cached data items, and optimize the cache allocation strategy based on these indicators.
It realizes accurate optimization of cached data, improves the overall performance and resource utilization efficiency of the cache system, and can respond to changes in data access mode in a timely manner.
Smart Images

Figure CN119961186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cache data processing technology, and in particular to a cache data optimization method, device, equipment, medium and computer program product. Background Art
[0002] When optimizing data cache, existing technical solutions often focus on the optimization of a single cache data item or a single data group, and lack a global optimization method for the entire cache system. Existing local optimization strategies may lead to uneven resource allocation, and some important data may be overly neglected, thus affecting the overall performance of the cache system. The dynamic cache adjustment mechanism of the prior art is usually based on a preset fixed cycle or threshold trigger. This mechanism lacks a certain degree of flexibility and is difficult to respond to sudden changes in data access patterns in a timely manner. In a highly dynamic environment, user behavior may change instantly due to system events, and fixed-cycle adjustments may miss critical cache data optimization opportunities. Summary of the invention
[0003] The present invention provides a cache data optimization method, device, equipment, medium and computer program product, which are used to solve the defect that the existing cache data optimization technology is not accurate enough and realize the accurate optimization of cache data.
[0004] The present invention provides a cache data optimization method, comprising the following steps: Based on the access timestamp and access source identifier of the cached data item, a data access relationship graph is constructed; Detecting the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; Determining a hotspot potential value of each cached data item based on a centrality index of the data group; Based on the hotspot potential value of each cached data item, the cache allocation strategy of the data group is optimized.
[0005] According to a cache data optimization method provided by the present invention, the construction of a data access relationship graph based on an access timestamp and an access source identifier of a cache data item includes: Determine the co-occurrence frequency of any two cached data items based on the access timestamp of each cached data item; Determine an access source diversity index of each cached data item based on the access source identifier of each cached data item; Determine the edge weights of any two cached data items based on the co-occurrence frequency and access source diversity index of any two cached data items; Construct a data access relationship graph based on the edge weights of any two cached data items.
[0006] According to a cache data optimization method provided by the present invention, the detecting of the data access relationship graph to obtain the centrality index of the data group includes: Determining the modularity of the initial group based on the edge weights; Determining a gain threshold based on a system load and an access pattern change rate; the system load is a load of a cache system corresponding to the cache data item; the access pattern change rate is determined based on the edge weight; Performing an iterative merging operation on the initial group to obtain a modularity gain; Based on the comparison result of the modularity gain and the gain threshold, merging the initial groups to obtain a data group; A centrality index is determined for each of the data groups.
[0007] According to a cache data optimization method provided by the present invention, determining the centrality index of each data group includes: Determining the external connection strength of each of the data groups based on the connection relationship between the data groups; Determining the internal connection density of each of the data groups based on the number of cached data items in each of the data groups; Based on the external connection strength and the internal connection density, a centrality index of each of the data groups is determined.
[0008] According to a cache data optimization method provided by the present invention, the determining of the hotspot potential value of each cache data item based on the centrality index of the data group includes: Determining a centrality influencing factor based on the centrality index of the data group; constructing a decay function of the target data item based on access time information and data association of the target data item; the target data item belongs to the data group; Determine the local hotspot index of the target data item based on the access count information and the average relevance of the target data item; the average relevance is determined based on the edge weight corresponding to the target data item; Based on the centrality influence factor, the attenuation function, the local hotspot index and the access source diversity index, the hotspot potential value of the target data item is determined.
[0009] According to a cache data optimization method provided by the present invention, the optimization of the cache allocation strategy of the data group based on the hotspot potential value of each cache data item includes: When the average hotspot potential value of the cached data items in the data group is greater than the potential threshold, the data group is migrated from the current cache layer to the target cache layer; the processing performance of the target cache layer is higher than that of the current cache layer; When the hotspot potential value of the target data item changes, the cache priority of the target data item in the data group is adjusted.
[0010] The present invention also provides a cache data optimization device, comprising the following modules: A data access relationship graph construction module is used to construct a data access relationship graph based on the access timestamp and access source identifier of the cached data item; A centrality index determination module, used to detect the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; A hotspot potential value determination module, used to determine the hotspot potential value of each cached data item based on the centrality index of the data group; The cache allocation strategy adjustment module is used to optimize the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements any of the cache data optimization methods described above when executing the computer program.
[0012] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the cache data optimization method described in any one of the above methods is implemented.
[0013] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the cache data optimization method described above is implemented.
[0014] The cache data optimization method, device, equipment, medium and computer program product provided by the present invention accurately capture complex data access patterns by constructing a data access relationship graph; based on the analysis and detection of the data access relationship graph, the group centrality index is calculated to provide a more comprehensive metric for the importance of the array; and then combined with a multi-factor fusion method, global and local factors are effectively balanced to provide a more reliable decision-making basis for resource allocation for cache data optimization; finally, a dynamic cache adjustment strategy based on hotspot potential values realizes multi-level optimization from the cache data item level to the data group level, thereby improving the accuracy of cache data optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 This is one of the flow charts of the cache data optimization method provided by the present invention.
[0017] Figure 2 This is the second flow chart of the cache data optimization method provided by the present invention.
[0018] Figure 3 It is a structural schematic diagram of the cache data optimization device provided by the present invention.
[0019] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Combine the following Figure 1-Figure 4 The present invention describes the cache data optimization method, device, equipment, medium and computer program product.
[0022] Figure 1 This is one of the flow charts of the cache data optimization method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 100: construct a data access relationship graph based on the access timestamp and access source identifier of the cached data item; Specifically, a data access relationship graph is constructed through the access timestamp and access source identifier of each cache data item, wherein the nodes in the data access relationship graph are cache data items, and the edges between any two cache data items represent the access association strength.
[0023] Compared with the existing method of only recording the number of cache accesses, the method of recording access timestamps and access source identifiers provides a more fine-grained data access pattern in the distributed cache system, which can capture the temporal relationship of data access and the diversity of access sources. In a large-scale microservice architecture environment, the method of recording access timestamps and access source identifiers enables the distributed cache system to distinguish the access patterns of different services or tenants, thereby providing richer contextual information for subsequent data hotspot detection, which helps to identify the detection patterns of complex hotspots across services or tenants.
[0024] Based on the access timestamps of cache data access records, the co-occurrence frequency of any two cache data items within a predefined time window (for example, 30 minutes) is calculated. Within this window, the number of times any two cache data items are accessed successively can be calculated. By introducing the sliding time window mechanism, the timeliness and accuracy of hotspot detection are significantly improved. Compared with the existing static global statistical methods, this dynamic calculation method can timely reflect changes in data access patterns. In scenarios with highly dynamic loads, this method can quickly identify suddenly appearing combinations of related data items, providing timely decision-making basis for cache preheating and resource allocation.
[0025] Furthermore, the access source diversity index of the cached data items is introduced, and the access source diversity index is obtained by calculating the entropy value of the access source identifier. Finally, the edge weights of the data access relationship graph are adjusted, and the adjusted edge weights are the weighted sum of the co-occurrence frequency and the access source diversity index. Adjusting the edge weights based on diversity can achieve a more comprehensive evaluation of the data association strength. The co-occurrence frequency and the global importance of the cached data items are taken into consideration. Compared with existing methods based only on local statistics, this comprehensive evaluation can better balance local hotspots and global importance. In complex microservice architectures or cloud services deployed across regions, this adjustment can help identify those data item associations that have a significant impact on the performance of the entire system, thereby optimizing cross-service or cross-regional caching strategies and improving the overall response speed and resource utilization efficiency of the cache system.
[0026] Step 200: Detect the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; Specifically, the cache data optimization method provided by the present application uses an improved Louvain algorithm (a community discovery algorithm based on modularity) to calculate modularity. The accuracy of community (group) division is improved by incorporating edge weights into the modularity calculation. Unlike the existing simple graph partitioning that only considers whether a connection exists, the embodiment of the present application makes full use of the constructed fine-grained data access relationship graph. In a large-scale microservice architecture, this improvement enables the distributed cache system to more accurately identify closely related data groups, thereby providing a solid foundation for subsequent data cache strategy optimization. For example, when processing complex transactions involving multi-level associations, the improved group partitioning can better capture the intrinsic connections between these cached data items, which helps the distributed cache system to consider more comprehensive contextual information when making cache decisions.
[0027] Then, the threshold function constructed according to the current system load and access pattern is used to adaptively adjust the group merging conditions. In the actual operating environment, when faced with burst traffic or periodic load fluctuations, the dynamic adjustment of the threshold function can reduce the number of groups by increasing the merging threshold during high load periods, thereby reducing the complexity of cache management. During low load periods, the threshold can be lowered to form a finer-grained group to improve the cache hit rate. This flexible adjustment not only improves the overall performance of the system, but also greatly enhances its adaptability in various complex scenarios.
[0028] Then, the group merging process is iterated until the modularity can no longer be optimized or the preset number of iterations is reached. For each iterative merge, the modularity gain before and after the group merge is calculated, and the modularity gain is compared with the dynamic threshold. The decision on whether to perform the merge is based on the comparison result. For example, when , performs a group merge operation.
[0029] The iterative group merging process combines the modularity gain calculation and dynamic threshold comparison process, providing an accurate and efficient method for group partitioning. Different from the traditional one-time partitioning or fixed number of iterations, this method can adaptively determine the number of iterations and group merging operations according to the complexity of the actual data relationship. In practical applications, for example, when processing social network data with multi-level relationships or complex enterprise resource planning system data, this method can better balance the computational overhead and the quality of group partitioning. By dynamically controlling the group merging process, the system can respond to changes in data access patterns in a short time while avoiding the loss of accuracy caused by excessive merging. This not only improves the real-time performance of the cache system, but also provides a more reliable decision-making basis for resource allocation and load balancing. Finally, based on the internal and external connection relationships of the data group, the centrality index is calculated for each data group.
[0030] Step 300: Determine the hotspot potential value of each cached data item based on the centrality index of the data group; Specifically, the cache data optimization method provided by this application is different from the existing method that only focuses on the characteristics of a single data item. By introducing the community centrality influencing factor , taking into account the overall importance of the group to which the cached data item belongs. By combining the group centrality with the logarithm of the group size, the potential influence of the cached data item in the entire system can be more accurately evaluated, and large-scale microservice architectures or complex distributed database systems can be effectively handled. By considering the three dimensions of time, access frequency, and data relevance at the same time, the dynamic changes in data heat can be more finely characterized. This multi-dimensionality is particularly suitable for processing complex data access patterns. Compared with existing methods that only consider access frequency, this scheme can identify cached data items that are not the most frequently accessed but occupy a key position in the data relationship network.
[0031] The hotspot potential value is a comprehensive and accurate hotspot evaluation indicator determined by organically combining factors such as group influence, local importance, multidimensional attenuation and access source diversity.
[0032] Step 400: Optimize the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
[0033] Finally, according to the calculated hotspot potential value, the cache allocation strategy is dynamically adjusted, including overall migration of data groups with high hotspot potential values, adjustment of cache priorities of cached data items within the group, and optimization of data access paths between groups.
[0034] This embodiment accurately captures complex data access patterns by constructing a data access relationship graph; based on the analysis and detection of the data access relationship graph, the group centrality index is calculated to provide a more comprehensive metric for the importance of the array; and then combined with a multi-factor fusion method, it effectively balances global and local factors to provide a more reliable decision-making basis for resource allocation for cache data optimization; finally, a dynamic cache adjustment strategy based on hotspot potential values realizes multi-level optimization from the cache data item level to the data group level, thereby improving the accuracy of cache data optimization.
[0035] Figure 2 This is the second flow chart of the cache data optimization method provided by the present invention. Figure 2 As shown, the method may also include: Step 110: determining the co-occurrence frequency of any two cached data items based on the access timestamp of each cached data item; Step 120: determining an access source diversity index of each cache data item based on the access source identifier of each cache data item; Step 130: Determine the edge weights of any two cached data items based on the co-occurrence frequency and access source diversity index of any two cached data items; Step 140: construct a data access relationship graph based on the edge weights of any two cached data items.
[0036] Specifically, first, record the access timestamp and access source identifier of each cached data item in the distributed cache system. In the distributed cache system, implement a lightweight logging mechanism. This lightweight logging mechanism can record the key information of each cached data access, which includes the accessed cached data item ID (identifier), access timestamp, and a unique access source identifier. This access source identifier mainly depends on the specific architecture of the distributed cache system, which can be the client IP, user ID, or application instance ID.
[0037] For cached data items and , the calculation of the above co-occurrence frequency is shown in formula 1, where, For cache data items and The co-occurrence frequency of Is to cache data items within a time window and cache data items The number of visits; is the unit length of the time window, for example, one minute.
[0038] ; (1) For each cached data item, count the number of different access sources that access the cached data item and calculate the distribution entropy of these access sources. The higher the distribution entropy, the more dispersed the access sources of the cached data item are, and it may be a more common hot spot. The calculation of the access source diversity index is shown in Formula 2, where, For cache data items The diversity index of access sources; Access to cached data items The total number of different access sources; is the kth access source accessing the cache data item probability. The calculation of is shown in formula 3, where Access cache data item for the kth access source The number of times; For cache data items The total number of visits.
[0039] ; (2) ; (3) The edge weight of the data access relationship graph is determined based on the weighted sum of the co-occurrence frequency and the access source diversity index, as shown in Formula 4, where: For cache data items and cache data items The edge weights between ; and is the weight coefficient, such as ; For cache data items and cache data items The co-occurrence frequency of For cache data items The diversity index of access sources; For cache data items The diversity index of access sources.
[0040] ; (4) Pick and The small value of is to ensure that the edge weight is increased significantly only when both cached data items have high diversity.
[0041] This embodiment constructs a data access relationship graph through two indicators, namely, co-occurrence frequency and access source diversity index, and takes into account the global importance of co-occurrence frequency and cached data items. Compared with the existing method based only on local statistics, it can better balance local hot spots and global importance.
[0042] In one embodiment, the cache data optimization method provided in the embodiment of the present application may further include: Step 210: determining the modularity of the initial group based on the edge weights; Step 220: determining a gain threshold based on a system load and an access pattern change rate; the system load is a load of a cache system corresponding to the cache data item; the access pattern change rate is determined based on the edge weight; Step 230: performing an iterative merging operation on the initial group to obtain a modularity gain; Step 240: Based on the comparison result of the modularity gain and the gain threshold, the initial groups are merged to obtain a data group; Step 250: Determine the centrality index of each of the data groups.
[0043] Specifically, the specific contents of the above step 200 include: The improved Louvain algorithm is applied to the data access relationship graph to perform preliminary group division. Specifically, the edge weights Incorporate the modularity calculation formula to improve the accuracy of group division. The modularity calculation is shown in Formula 5, where: is modularity; Is a cache data item The weighted degree of Is a cache data item The weighted degree of It is the sum of all edge weights in the data access graph; For cache data items Groups you belong to; For cache data items Groups you belong to; Indicates when 1 if the value is 0, otherwise it is 0.
[0044] ; (5) Then, set the threshold function ,in, is the system load factor; is the access pattern change rate. Threshold function The calculation of is shown in Formula 6, where is the basic threshold; is the (distributed cache) system load factor, which represents the ratio of the current system load to the maximum load; The access pattern change rate can be determined by calculating the average value of edge weight changes in consecutive time windows; and is a tuning parameter.
[0045] ; (6) ; (7) The calculation of modularity gain is shown in Formula 7, where: The modularity gain before and after merging groups for each iteration; is the sum of the edge weights within the group; is the sum of the weights of all edges connected to the group; For cache data items The sum of the edge weights connected to the cached data items within the target group. When , the group merging operation is performed. Finally, for each identified data group, the group centrality index is calculated.
[0046] This embodiment provides a more comprehensive metric for evaluating the importance of data groups by calculating the group centrality index.
[0047] In one embodiment, the cache data optimization method provided in the embodiment of the present application may further include: Step 251: Determine the external connection strength of each of the data groups based on the connection relationship between the data groups; Step 252: Determine the internal connection density of each of the data groups based on the number of cached data items in each of the data groups; Step 251: Determine a centrality index of each of the data groups based on the external connection strength and the internal connection density.
[0048] Combining the node degree centrality and the internal connection density of the group, the calculation of the group centrality index is determined as shown in Formula 8, where: is the group centrality index; is the strength of external connections of the group; is the density of connections within the group; is a trade-off factor.
[0049] ; (8) ; (9) ; (10) The calculation of is shown in formula 9, where and Represents different nodes. Representation Node Belong to group , that is, within the specified group. Representation Node Not in group , that is, outside the group. Formula 9 calculates the group The sum of the connection strengths between the nodes in the group and the nodes outside the group; Representation Node and nodes The edge weight between two nodes measures the strength of the connection between them.
[0050] The calculation of is shown in formula 10, where Indicates a specific group of data; Is a data group The number of cached data items in .
[0051] This embodiment adjusts the trade-off factor to flexibly balance the internal cohesion and external influence of the group according to specific needs. This not only helps to allocate cache resources more accurately, but also provides valuable guidance for the horizontal expansion and vertical optimization of the system, thereby improving the cache hit rate while optimizing the overall system architecture.
[0052] In one embodiment, the cache data optimization method provided in the embodiment of the present application may further include: Step 310: determining a centrality influencing factor based on the centrality index of the data group; Step 320: constructing a decay function of the target data item based on the access time information and data association of the target data item; the target data item belongs to the data group; Step 330: Determine the local hotspot index of the target data item based on the access count information and the average relevance of the target data item; the average relevance is determined based on the edge weight corresponding to the target data item; Step 340: Determine the hotspot potential value of the target data item based on the centrality influence factor, the attenuation function, the local hotspot index, and the access source diversity index.
[0053] Specifically, calculate each data group The centrality factor , as shown in formula 11, where For the data community The number of cached data items in; and It is an adjustable parameter used to balance the effects of group centrality and group size; ; (11) ; (12) Next, construct the decay function , as shown in formula 12, where is the time since the last visit; The frequency of recent visits; is the data relevance; is the time decay constant; is the correlation decay constant; is the access frequency threshold, when and Close, or Exceed , the effect on attenuation will be reduced. It is an adjustment parameter used to control the smoothness of frequency attenuation. To adjust the access frequency and The relationship between makes the decay function more sensitive or smoother to the change of access frequency.
[0054] Then, calculate the cache data items Local hotspot index ,in, For cache data items Number of visits; For cache data items The average correlation with other data items; and is the weight coefficient, and ; and are the minimum and maximum access times of all cached data items, respectively; and They are the minimum and maximum values of the average association of all cached data items, respectively.
[0055] ; (13) ; (14) Finally, the cache data items are calculated comprehensively The hotspot potential value of Cache data items , its hotspot potential value The calculation of is shown in formula 14. Where, For cache data items The diversity index of access sources.
[0056] This embodiment can better adapt to complex and changeable distributed environments by integrating multiple factors. It can effectively balance global and local factors when processing large-scale systems across regions and multiple data centers, and provide a more reliable decision-making basis for resource allocation.
[0057] In one embodiment, the cache data optimization method provided in the embodiment of the present application may further include: Step 410: When the average hotspot potential value of the cached data items in the data group is greater than the potential threshold, migrate the data group from the current cache layer to the target cache layer; the processing performance of the target cache layer is higher than that of the current cache layer; Step 420: When the hotspot potential value of the target data item changes, adjust the cache priority of the target data item in the data group.
[0058] Specifically, the embodiments of the present application provide several cache allocation strategy adjustment methods.
[0059] If the average hotspot potential value of a data group exceeds a preset threshold, the data group is migrated as a whole to a cache layer with higher performance.
[0060] Example: Calculate the average hotspot potential value of each data group, compare it with the migration threshold currently set by the system, and perform overall migration operations on the data groups that meet the conditions.
[0061] If the hotspot potential value of a cached data item changes, the cache priority of the cached data item within the data group to which it belongs is adjusted.
[0062] Example: Sort cache data items in a data group according to hotspot potential values, and adjust the cache replacement strategy according to the sorting result. Cache data items with higher hotspot potential values receive higher retention priority.
[0063] If the access frequency between the two data groups exceeds a preset threshold, the data access path between the two data groups is optimized.
[0064] Example: Calculate the access frequency between different data groups. For pairs of data groups that frequently access each other, establish direct data channels in the cache system or adjust their physical storage locations to reduce access latency.
[0065] This embodiment optimizes the cache allocation strategy of the data groups by calculating the hotspot potential value and the access frequency between the data groups.
[0066] The cache data optimization device provided by the present invention is described below. The cache data optimization device described below and the cache data optimization method described above can be referenced to each other.
[0067] Please refer to Figure 3 The present invention also provides a cache data optimization device, comprising: A data access relationship graph construction module 301 is used to construct a data access relationship graph based on the access timestamp and access source identifier of the cached data item; A centrality index determination module 302 is used to detect the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; A hotspot potential value determination module 303, configured to determine a hotspot potential value of each cached data item based on the centrality index of the data group; The cache allocation strategy adjustment module 304 is used to optimize the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
[0068] Optionally, the data access relationship graph construction module includes: a co-occurrence frequency determination unit, configured to determine the co-occurrence frequency of any two cached data items based on the access timestamp of each cached data item; An access source diversity index determining unit, used to determine an access source diversity index of each cache data item based on the access source identifier of each cache data item; An edge weight determination unit, used to determine the edge weights of any two cached data items based on the co-occurrence frequency and access source diversity index of any two cached data items; The data access relationship graph construction unit is used to construct a data access relationship graph based on the edge weights of any two cache data items.
[0069] Optionally, the centrality index determination module includes: A modularity determination unit, configured to determine the modularity of an initial group based on the edge weights; a gain threshold determination unit, configured to determine a gain threshold based on a system load and an access pattern change rate; the system load is a load of a cache system corresponding to the cache data item; the access pattern change rate is determined based on the edge weight; A modularity gain determination unit, configured to perform an iterative merging operation on the initial group to obtain a modularity gain; a data group determination unit, configured to merge the initial groups to obtain a data group based on a comparison result of the modularity gain and the gain threshold; The centrality index determining unit is used to determine the centrality index of each of the data groups.
[0070] Optionally, the centrality index determining unit includes: an external connection strength determination unit, configured to determine the external connection strength of each of the data groups based on the connection relationship between the data groups; an internal connection density determining unit, configured to determine the internal connection density of each of the data groups based on the number of cached data items in each of the data groups; The data group centrality index determining unit is used to determine the centrality index of each of the data groups based on the external connection strength and the internal connection density.
[0071] Optionally, the hotspot potential value determination module includes: a centrality influence factor determination unit, configured to determine a centrality influence factor based on the centrality index of the data group; a decay function construction unit, configured to construct a decay function of a target data item based on access time information and data association of the target data item; the target data item belongs to the data group; A local hotspot index determination unit, configured to determine the local hotspot index of the target data item based on the access count information and the average relevance of the target data item; the average relevance is determined based on the edge weight corresponding to the target data item; A hotspot potential value determination unit is used to determine the hotspot potential value of the target data item based on the centrality influence factor, the attenuation function, the local hotspot index and the access source diversity index.
[0072] Optionally, the cache allocation strategy adjustment module includes: a data group migration unit, configured to migrate the data group from a current cache layer to a target cache layer when an average hotspot potential value of cache data items in the data group is greater than a potential threshold; the processing performance of the target cache layer is higher than that of the current cache layer; The cache priority adjustment unit is used to adjust the cache priority of the target data item in the data group when the hotspot potential value of the target data item changes.
[0073] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the cache data optimization method, which includes: constructing a data access relationship graph based on the access timestamp and access source identifier of the cache data item; detecting the data access relationship graph to obtain the centrality index of the data group; the data group is constructed based on the cache data item; determining the hotspot potential value of each cache data item based on the centrality index of the data group; optimizing the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
[0074] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0075] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cache data optimization method provided by the above-mentioned methods, which includes: constructing a data access relationship graph based on the access timestamp and access source identifier of the cache data item; detecting the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cache data item; based on the centrality index of the data group, determining the hotspot potential value of each cache data item; based on the hotspot potential value of each cache data item, optimizing the cache allocation strategy of the data group.
[0076] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the cache data optimization method provided by the above-mentioned methods, the method comprising: constructing a data access relationship graph based on the access timestamp and access source identifier of the cache data item; detecting the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cache data item; determining a hotspot potential value of each cache data item based on the centrality index of the data group; and optimizing the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
[0077] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0078] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cache data optimization method, characterized in that: include: Based on the access timestamp and access source identifier of the cached data item, a data access relationship graph is constructed; Detecting the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; Determining a hotspot potential value of each cached data item based on a centrality index of the data group; Based on the hotspot potential value of each cached data item, the cache allocation strategy of the data group is optimized.
2. The cache data optimization method according to claim 1, characterized in that: The step of constructing a data access relationship graph based on the access timestamp and access source identifier of the cached data item includes: Determine the co-occurrence frequency of any two cached data items based on the access timestamp of each cached data item; Determine an access source diversity index of each cached data item based on the access source identifier of each cached data item; Determine the edge weights of any two cached data items based on the co-occurrence frequency and access source diversity index of any two cached data items; Construct a data access relationship graph based on the edge weights of any two cached data items.
3. The cache data optimization method according to claim 2, characterized in that: The detecting of the data access relationship graph to obtain the centrality index of the data group includes: Determining the modularity of the initial group based on the edge weights; Determining a gain threshold based on a system load and an access pattern change rate; the system load is a load of a cache system corresponding to the cache data item; the access pattern change rate is determined based on the edge weight; Performing an iterative merging operation on the initial group to obtain a modularity gain; Based on the comparison result of the modularity gain and the gain threshold, merging the initial groups to obtain a data group; A centrality index is determined for each of the data groups.
4. The cache data optimization method according to claim 3, characterized in that: Determining the centrality index of each data group includes: Determining the external connection strength of each of the data groups based on the connection relationship between the data groups; Determining the internal connection density of each of the data groups based on the number of cached data items in each of the data groups; Based on the external connection strength and the internal connection density, a centrality index of each of the data groups is determined.
5. The cache data optimization method according to claim 2, characterized in that: Determining the hotspot potential value of each cached data item based on the centrality index of the data group includes: Determining a centrality impact factor based on the centrality index of the data group; constructing a decay function of the target data item based on access time information and data association of the target data item; the target data item belongs to the data group; Determine the local hotspot index of the target data item based on the access count information and the average relevance of the target data item; the average relevance is determined based on the edge weight corresponding to the target data item; Based on the centrality influence factor, the attenuation function, the local hotspot index and the access source diversity index, the hotspot potential value of the target data item is determined.
6. The cache data optimization method according to claim 5, characterized in that: The optimizing the cache allocation strategy of the data group based on the hotspot potential value of each cache data item includes: When the average hotspot potential value of the cached data items in the data group is greater than the potential threshold, the data group is migrated from the current cache layer to the target cache layer; the processing performance of the target cache layer is higher than that of the current cache layer; When the hotspot potential value of the target data item changes, the cache priority of the target data item in the data group is adjusted.
7. A cache data optimization device, characterized in that: include: A data access relationship graph construction module is used to construct a data access relationship graph based on the access timestamp and access source identifier of the cached data item; A centrality index determination module, used to detect the data access relationship graph to obtain a centrality index of a data group; the data group is constructed based on the cached data item; A hotspot potential value determination module, used to determine the hotspot potential value of each cached data item based on the centrality index of the data group; The cache allocation strategy adjustment module is used to optimize the cache allocation strategy of the data group based on the hotspot potential value of each cache data item.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the cache data optimization method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cache data optimization method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the cache data optimization method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Data caching optimization method in edge computing environment
CN121807920A
A data cache optimization method in an edge computing environment
CN121807920B